Mutual learning method and system between multi-view medical images and text reports

By building a mutual learning model, the cross-modal relationship problem between multi-view medical images and text reports is solved, richer feature representation and more accurate diagnosis results are achieved, and the model's robustness to missing or damaged data is enhanced.

CN119694479BActive Publication Date: 2025-10-03CHONGQING UNIV
2 Cites 0 Cited by

Patent Information

Application Number
CN202411857618.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-10-03
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

Existing visual language-based models cannot be effectively applied to the cross-modal relationship between multi-view medical images and corresponding text reports, ignoring the importance of multi-view images in clinical practice, resulting in information redundancy and difficulty in capturing lesion semantics or personalized features.

Method used

A mutual learning model is constructed, and pre-training and joint training are performed through image reconstruction tasks, report reconstruction tasks and multi-view alignment tasks. Image mask and report mask technology are used, combined with the exchange multimodal fusion method to enhance the cross-modal fusion of visual and text features.

Benefits of technology

It improves the model's robustness to missing or damaged data, enhances the model's ability to explore lesion information, and optimizes the accuracy and comprehensiveness of diagnostic results.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and specifically discloses a mutual learning method and system between multi-view medical images and text reports. By constructing a mutual learning model and dividing the learning process of the mutual learning model into an image reconstruction task, a report reconstruction task, and a multi-view alignment task, the image reconstruction task obtains the feature representation of the multi-view medical image, and the report reconstruction task obtains the feature representation of the text report and cross-modally fuses it with the image feature representation to obtain a reconstructed text report. The integration of multimodal reconstruction tasks enables the model to learn richer and more detailed feature representations, thereby improving its robustness to lost or damaged data. In addition, the exchange-based multimodal fusion method in cross-modal text reconstruction aims to fully integrate visual features and enrich the semantic representation of domain-specific knowledge. Through pre-training and joint training, the model performance is optimized so that more accurate and comprehensive diagnostic results can be obtained after applying the mutual learning model.
Need to check novelty before this filing date? Find Prior Art