开发并验证预测肺癌患者对PD-1抑制剂客观反应的机器学习模型
Development and validation of a machine learning model for predicting objective response to PD-1 inhibitors in lung cancer patients.
作者
作者单位
- Department of Pharmacy, The Second Affiliated Hospital, Guangzhou Medical University, Guangzhou, China.
- Department of Pharmacy, Xiaolan Clinical Institute of Shantou University Medical College, Zhongshan, China.
- Department of Clinical Pharmacy, Guangzhou Medical University, Guangzhou, China.
摘要
中文
鉴于肺癌免疫治疗疗效存在显著异质性,识别可靠的预测因素至关重要。本研究旨在开发一种机器学习(ML)模型,以识别接受程序性细胞死亡蛋白1(PD-1)抑制剂治疗的肺癌患者获得客观反应的预测因素。对2021年11月至2024年11月期间在广州医科大学第二附属医院接受PD-1抑制剂治疗的肺癌患者的数据进行了回顾性分析。根据治疗效果将患者分为反应组[完全缓解(CR)+部分缓解(PR)]和无反应组[疾病稳定(SD)+疾病进展(PD)]。经过最小绝对收缩和选择算子(Lasso)分析筛选临床特征后,使用10种ML算法评估前10个反应相关因素以构建模型。使用包括受试者工作特征(ROC)曲线下面积(AUC)、准确度、敏感性、特异性和F1分数在内的指标严格评估模型性能。SHapley Additive exPlanations(SHAP)算法分析了最佳模型中的特征贡献。本研究纳入了212名接受PD-1抑制剂治疗的肺癌患者。在10个模型中,分类提升(CatBoost)模型表现出最佳的整体预测性能,训练集AUC为0.961,验证集AUC为0.871。SHAP结果表明,低肿瘤-淋巴结-转移(TNM)分期、高血红蛋白水平和有手术史的患者更可能对肺癌治疗获得客观反应。此外,Kaplan-Meier分层显示,高反应概率的患者表现出显著更长的无进展生存期。本研究成功开发并验证了一个用于预测接受PD-1抑制剂治疗的肺癌患者客观反应的CatBoost模型。该模型在预测短期客观反应方面显示出高效能,为临床评估提供了有力的工具。配套的基于网络的风险平台作为假设生成和决策支持工具,在支持接受PD-1抑制剂的肺癌患者的个体化筛选和风险分层方面具有初步临床潜力。
English
Given the substantial heterogeneity in the efficacy of immunotherapy for lung cancer, identifying reliable predictive factors is crucial. This study aims to develop a machine learning (ML) model to identify predictors of achieving objective response in lung cancer patients receiving programmed cell death protein 1 (PD-1) inhibitor therapy. A retrospective analysis was conducted on data from lung cancer patients treated with PD-1 inhibitors at The Second Affiliated Hospital of Guangzhou Medical University between November 2021 and November 2024. Patients were categorised into response group [complete response (CR) + partial response (PR)] and non-response group [stable disease (SD) + progressive disease (PD)] based on treatment efficacy. Following least absolute shrinkage and selection operator (Lasso) analysis to screen clinical characteristics, 10 ML algorithms were employed to evaluate the top 10 response-associated factors for model construction. Model performance was rigorously assessed using metrics including area under the receiver operating characteristic (ROC) curve (AUC), accuracy, sensitivity, specificity, and F1 score. The SHapley Additive exPlanations (SHAP) algorithm analysed feature contributions within the optimal model. This study included 212 lung cancer patients who were treated with PD-1 inhibitors. The categorical boosting (CatBoost) model demonstrated the best overall predictive performance among the 10 models, achieving an AUC of 0.961 in the training set and 0.871 in the validation set. SHAP results indicated that patients with low tumour-node-metastasis (TNM) staging, high haemoglobin levels, and a history of surgery were more likely to achieve objective response to lung cancer treatment. Additionally, Kaplan-Meier stratification showed that patients with a high probability of response exhibited significantly longer progression-free survival. This study successfully developed and validated a CatBoost model to predict the objective response of lung cancer patients receiving PD-1 inhibitor therapy. The model demonstrates high efficacy in predicting short-term objective response, providing a robust tool for clinical evaluation. The accompanying web-based risk platform serves as a hypothesis-generating and decision-supportive tool, holding preliminary clinical potential to support individualised screening and risk stratification in lung cancer patients receiving PD-1 inhibitors.
分类与指标
- 研究类型
- AI/ML
- 病种
- 肺癌
- JCR 分区
- Q2
- 影响因子
- 3.4
- 新锐分区
- 3区