Explainable medical multimodal search enhanced generation method, system, device, medium

By combining regional anomaly detection and multimodal similarity graph retrieval with a visual language model trained with enhanced reasoning, the inaccuracy of regional retrieval and the lack of transparency in decision-making in existing medical image analysis technologies are solved, achieving high-precision and interpretable medical image analysis results.

CN121834023BActive Publication Date: 2026-06-09ANHUI PROVINCIAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI PROVINCIAL HOSPITAL
Filing Date
2026-03-13
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing medical image analysis technologies lack high-precision regional-level retrieval, transparent information utilization, and enhanced reasoning capabilities, resulting in inaccurate outputs and opaque decision-making processes in multimodal medical tasks, thus hindering the widespread application of medical artificial intelligence systems in clinical settings.

Method used

We employ a method that combines regional anomaly detection, multimodal knowledge base similarity retrieval, and enhanced reasoning training. We locate anomaly regions using a dedicated detection model, construct a multimodal similarity map for regional retrieval, and generate interpretable output through a two-stage trained visual language model.

Benefits of technology

It achieves high-precision regional-level retrieval, improves the relevance and traceability of retrieval results, enhances the logical credibility and clinical interpretability of generated content, and ensures the accuracy and transparency of generated results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834023B_ABST
    Figure CN121834023B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical image analysis, and discloses an interpretable medical multimodal retrieval enhanced generation method, system, device and medium, which comprises the following steps: performing abnormal area detection on input medical images to obtain all candidate abnormal areas; performing region-level similarity retrieval in a pre-constructed multimodal knowledge base based on the abnormal areas to obtain relevant reference text information; the multimodal knowledge base comprises region-level medical image features and text features, and a multimodal similarity graph is constructed based on the image features and the text features; the region-level similarity retrieval is performed based on the multimodal similarity graph; the medical images, the position information of the abnormal areas and the retrieved reference text information are combined into model input according to a preset format; and the model input is input into a visual language model trained through enhanced reasoning, which generates interpretable output containing reasoning logic and a final conclusion according to the input. The logical reliability and clinical interpretability of the generated content are enhanced.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Image retrieval method and device, electronic equipment and storage medium

    CN113553460A

  • Medical image report generation method and system and computer storage medium

    CN117954041A