Mutual learning method and system between multi-view medical images and text reports
By building a mutual learning model, the cross-modal relationship problem between multi-view medical images and text reports is solved, richer feature representation and more accurate diagnosis results are achieved, and the model's robustness to missing or damaged data is enhanced.
Patent Information
- Application Number
- CN202411857618.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-12-17
AI Technical Summary
Existing visual language-based models cannot be effectively applied to the cross-modal relationship between multi-view medical images and corresponding text reports, ignoring the importance of multi-view images in clinical practice, resulting in information redundancy and difficulty in capturing lesion semantics or personalized features.
A mutual learning model is constructed, and pre-training and joint training are performed through image reconstruction tasks, report reconstruction tasks and multi-view alignment tasks. Image mask and report mask technology are used, combined with the exchange multimodal fusion method to enhance the cross-modal fusion of visual and text features.
It improves the model's robustness to missing or damaged data, enhances the model's ability to explore lesion information, and optimizes the accuracy and comprehensiveness of diagnostic results.