logo 胸外文献每日监控
2026年8月13日星期四
← 返回 全部文献

面向 NSCLC 临床试验预筛选的安全感知 AI:基于规则、单智能体与多智能体方法的对比概念验证研究

Safety-aware AI for NSCLC trial pre-screening: a comparative proof-of-concept study of rule-based, single-agent, and multi-agent approaches.

期刊
International Journal of Medical Informatics
PMID
42485959
原文
PubMed ↗
发布日期

作者

  • Tianzuo Yuan — Faculty of Health Sciences, University of Macau, Faculty of Health Sciences Building, E12 Avenida da Universidade, Taipa, Macau, 999078, China. Electronic address: cc31642@um.edu.mo.

作者单位

  • Faculty of Health Sciences, University of Macau, Faculty of Health Sciences Building, E12 Avenida da Universidade, Taipa, Macau, 999078, China. Electronic address: cc31642@um.edu.mo.

摘要

中文

非小细胞肺癌(NSCLC)临床试验常涉及复杂的生物标志物驱动和线数特异性入组标准,使预筛选工作量大且易出错。在这种高风险场景下,假性纳入——即错误地将不符合条件的患者判定为符合条件以进入下游审核——相比其他错误类型具有更大的临床风险。本研究旨在比较用于 NSCLC 试验预筛选的安全感知 AI 配置,重点关注降低假性纳入率并在证据有限时支持不确定性感知弃权。方法:采用干预性 NSCLC 试验入组文本和涵盖分期、生物标志物状态、既往治疗、ECOG 体能状态及排除相关因素变化的结构化合成患者摘要,进行回顾性概念验证评估。共 120 对患者-试验配对经人工审核,分为开发集(n=30)和保留测试集(n=90)。比较五种配置:规则基线、加入保守 Safety Agent 门控的规则法、单智能体 LLM(GPT-4o-mini)以及两种轻量级多智能体变体(M1、M2)。主要安全指标为假性纳入率;次要指标为准确性和不确定率。次要评估包括基于 TCGA 的不完整证据压力测试,以及对来自 PMC Open Access NSCLC 病例报告的 72 对患者-试验配对的真实文本外部验证。结果:Safety Agent 门控将 120 对合并队列的假性纳入率从 11.63% 降至 1.16%(McNemar 检验,p=0.002),保留集从 12.12% 降至 1.52%,同时不确定率上升。单智能体 LLM(GPT-4o-mini)取得最高观察准确率(0.74,95% CI 0.66-0.81),在 86 对金标准不符合条件的配对中无假性纳入(95% CI 0.0-4.3%),不确定率为 7.5%。多智能体 M1 和 M2 在结构化队列上也实现了零假性纳入,不确定率分别为 25.8% 和 1.7%;在 PMC 真实文本验证中,M1 保持零假性纳入(0/44;95% CI 0.0-8.0%),M2 因 Safety Agent 在模糊非结构化文本上的异常标签升级出现 1 例假性纳入。跨五个 LLM 配置、四个供应商的多供应商评估中,五分之四在全队列上实现零假性纳入。在生物标志物和 ECOG 数据缺失的 TCGA 压力测试中,所有评估系统均显示零假性纳入。结论:在受控合成输入条件的概念验证评估中,保守安全门控大幅降低了规则法预筛选的假性纳入率,所有基于 LLM 的配置在结构化合成队列上实现零或近零假性纳入。这些发现支持将显式假性纳入监测、不确定性感知弃权及多维评估作为 AI 辅助肿瘤临床试验预筛选的有用设计原则。然而,结果仅限于结构化合成数据、单一 LLM 家族及小样本人工审核队列;在安全性或实用性可推广之前,仍需在多样化非结构化 EHR 文档中进行前瞻性多中心验证。

English

Non-small cell lung cancer (NSCLC) trials often involve complex biomarker-driven and line-specific eligibility criteria, making pre-screening labor-intensive and error-prone. In this high-stakes setting, false inclusion-incorrectly judging an ineligible patient as eligible for downstream review-poses a greater clinical risk than other error types. To compare safety-aware AI configurations for NSCLC trial pre-screening, with explicit focus on reducing false inclusion rates and supporting uncertainty-aware abstention under evidence limitations. We performed a retrospective proof-of-concept evaluation using interventional NSCLC trial eligibility text and structured synthetic patient summaries covering variations in stage, biomarker status, prior therapy, ECOG performance status, and exclusion-relevant factors. A total of 120 patient-trial pairs were manually reviewed and split into a development set (n = 30) and held-out test set (n = 90). Five configurations were compared: rule-based baseline, rule-based with conservative Safety Agent gating, single-agent LLM (GPT-4o-mini), and two lightweight multi-agent variants (M1, M2). The primary safety metric was false inclusion rate; secondary metrics were accuracy and uncertain rate. Secondary evaluations included a TCGA-derived stress test under incomplete evidence and a real-text external validation on 72 patient-trial pairs from PMC Open Access NSCLC case reports. Safety Agent gating reduced the false inclusion rate from 11.63% to 1.16% in the combined 120-pair cohort (McNemar's test, p = 0.002) and from 12.12% to 1.52% in the held-out set, accompanied by an increase in uncertain rate. The single-agent LLM (GPT-4o-mini) achieved the highest observed accuracy (0.74, 95% CI 0.66-0.81), with no observed false inclusion among 86 gold-standard ineligible pairs (95% CI 0.0-4.3%) and an uncertain rate of 7.5%. Multi-agent M1 and M2 also achieved zero observed false inclusion on the structured cohort with distinct uncertainty profiles (25.8% and 1.7% uncertain, respectively); in the PMC real-text validation, M1 maintained zero false inclusion (0/44; 95% CI 0.0-8.0%) while M2 produced one false inclusion arising from an anomalous Safety Agent label upgrade on ambiguous unstructured text. A multi-provider evaluation across five LLM configurations spanning four providers found four of five achieving zero observed false inclusion on the full cohort. In the TCGA-derived stress test with missing biomarker and ECOG data, all evaluated systems showed zero observed false inclusion. In this proof-of-concept evaluation under controlled synthetic-input conditions, conservative safety gating substantially reduced false inclusion in rule-based pre-screening, and all LLM-based configurations achieved zero or near-zero false inclusion on the structured synthetic cohort. These findings support explicit false inclusion monitoring, uncertainty-aware abstention, and multi-dimensional evaluation as useful design principles for AI-assisted oncology trial pre-screening. However, results are limited to structured synthetic data, a single LLM family, and small manually reviewed cohorts; prospective multi-site validation on diverse unstructured EHR documentation is needed before safety or utility claims can be generalized.

分类与指标

研究类型
AI/ML
病种
肺癌
JCR 分区
Q1
影响因子
5.0
新锐分区
2区