Medical image model pre-training method and system based on differential hierarchical attention

By employing a differential hierarchical attention-based pre-training method for medical image models, the problems of inaccurate image-text alignment and insufficient suppression of redundant information in 3D medical images are addressed. This method achieves fine-grained alignment between images and reports and precise localization of abnormal regions, improving the model's performance and interpretability. It is applicable to a variety of downstream clinical tasks.

CN122415577APending Publication Date: 2026-07-17UNIV OF SCI & TECH OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF SCI & TECH OF CHINA
Filing Date
2026-05-18
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies lack fine-grained image-text alignment mechanisms at the sentence or region level in 3D medical imaging. They cannot establish a precise correspondence between abnormal text descriptions and anatomical regions in images, making it difficult to suppress redundant information between adjacent slices. The models are easily affected by background regions and have insufficient ability to locate abnormal regions. The accuracy and interpretability of the generated reports are poor, and they rely on a large amount of meticulous manual annotation.

Method used

A pre-training method for medical image models using differential hierarchical attention is employed. By combining a 3D visual encoder with a differential hierarchical attention mechanism and a text encoder in the medical field, a sentence-level contrastive learning framework is constructed. This achieves accurate alignment between local features of the image and the semantics of the report sentence, suppresses redundant background information, and improves the localization accuracy and interpretability of abnormal areas.

Benefits of technology

Without the need for additional fine annotation, the model achieves precise alignment between local image features and the semantics of the report sentence, improving its ability to capture small lesions, enhancing the localization accuracy of abnormal areas and the interpretability of report generation, and lowering the development threshold for clinical AI applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415577A_ABST
    Figure CN122415577A_ABST
Patent Text Reader

Abstract

本发明提出一种基于差分分层注意力的医学影像模型预训练方法及系统,属于人工智能技术领域,包括:步骤S1:获取三维医学影像及其对应的放射学报告;对放射学报告执行预处理,构建样本集;步骤S2:通过融合有差分分层注意力机制的3D视觉编码器,对三维医学影像执行从局部切片到全局扫描的多尺度特征提取,以建模跨切片差异,获得视觉token特征;步骤S3:通过医学领域预训练的文本编码器,对样本进行编码,生成句子级文本嵌入;步骤S4:构建句子级对比学习损失函数,以实现视觉token特征与句子级文本嵌入的双向对齐,得到预训练模型。本发明方法突出三维影像切片之间的变化区域、抑制冗余信息,从而增强异常区域特征表达能力。
Need to check novelty before this filing date? Find Prior Art