Medical image model pre-training method and system based on differential hierarchical attention
By employing a differential hierarchical attention-based pre-training method for medical image models, the problems of inaccurate image-text alignment and insufficient suppression of redundant information in 3D medical images are addressed. This method achieves fine-grained alignment between images and reports and precise localization of abnormal regions, improving the model's performance and interpretability. It is applicable to a variety of downstream clinical tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2026-05-18
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies lack fine-grained image-text alignment mechanisms at the sentence or region level in 3D medical imaging. They cannot establish a precise correspondence between abnormal text descriptions and anatomical regions in images, making it difficult to suppress redundant information between adjacent slices. The models are easily affected by background regions and have insufficient ability to locate abnormal regions. The accuracy and interpretability of the generated reports are poor, and they rely on a large amount of meticulous manual annotation.
A pre-training method for medical image models using differential hierarchical attention is employed. By combining a 3D visual encoder with a differential hierarchical attention mechanism and a text encoder in the medical field, a sentence-level contrastive learning framework is constructed. This achieves accurate alignment between local features of the image and the semantics of the report sentence, suppresses redundant background information, and improves the localization accuracy and interpretability of abnormal areas.
Without the need for additional fine annotation, the model achieves precise alignment between local image features and the semantics of the report sentence, improving its ability to capture small lesions, enhancing the localization accuracy of abnormal areas and the interpretability of report generation, and lowering the development threshold for clinical AI applications.
Smart Images

Figure CN122415577A_ABST