基于潜在扩散的半监督域自适应用于病理图像分类
Semi-Supervised Domain Adaptation with Latent Diffusion for Pathology Image Classification.
作者
摘要
中文
计算病理学中的深度学习模型通常由于域偏移而无法跨队列和机构泛化。现有方法要么未能利用来自目标域的无标签数据,要么依赖图像到图像的转换,这可能扭曲组织结构并损害模型准确性。在这项工作中,我们提出了一个半监督域自适应(SSDA)框架,利用在源域和目标域的无标签数据上训练的潜在扩散模型来生成保留形态且目标感知的合成图像。通过将扩散模型条件于基础模型特征、队列身份和组织制备方法,我们在源域中保留组织结构,同时引入目标域的外观特征。目标感知的合成图像与来自源队列的真实标记图像相结合,随后用于训练下游分类器,然后在目标队列上进行测试。所提出的SSDA框架在肺腺癌生长模式分类任务上得到验证。所提出的增强在目标队列的保留测试集上产生了显著更好的性能,而不会降低源队列性能。该方法将目标队列保留测试集上的加权F1分数从0.611提高到0.706,宏F1分数从0.641提高到0.716。我们的结果表明,基于目标感知扩散的合成数据增强为改进计算病理学中的域泛化提供了一种有前景且有效的方法。代码可在 https://github.com/hsu-lab/ldm-pathology-ssda 获取。
English
Deep learning models in computational pathology often fail to generalize across cohorts and institutions due to domain shift. Existing approaches either fail to leverage unlabeled data from the target domain or rely on image-to-image translation, which can distort tissue structures and compromise model accuracy. In this work, we propose a semi-supervised domain adaptation (SSDA) framework that utilizes a latent diffusion model trained on unlabeled data from both the source and target domains to generate morphology-preserving and target-aware synthetic images. By conditioning the diffusion model on foundation model features, cohort identity, and tissue preparation method, we preserve tissue structure in the source domain while introducing target domain appearance characteristics. The target-aware synthetic images, combined with real, labeled images from the source cohort, are subsequently used to train a downstream classifier, which is then tested on the target cohort. The effectiveness of the proposed SSDA framework is demonstrated on the task of lung adenocarcinoma growth pattern classification. The proposed augmentation yielded substantially better performance on the held-out test set from the target cohort, without degrading source cohort performance. The approach improved the weighted F1 score on the target-cohort held-out test set from 0.611 to 0.706 and the macro F1 score from 0.641 to 0.716. Our results demonstrate that target-aware diffusion-based synthetic data augmentation provides a promising and effective approach for improving domain generalization in computational pathology. Code is available at https://github.com/hsu-lab/ldm-pathology-ssda.
分类与指标
- 研究类型
- AI/ML
- 病种
- 肺癌
- JCR 分区
- Q1
- 影响因子
- 7.7
- 新锐分区
- 1区