Medical visual question answering method and device, computer device and storage medium

By combining medical image question pairs and the response text output by multiple specialist physician models, and utilizing image enhancement and multimodal feature extraction modules, more accurate medical visual question-answering answers are generated. This solves the bias and illusion problems of single models in medical visual question answering, and improves the accuracy and reliability of the answers.

CN122388079APending Publication Date: 2026-07-14SOUTH CHINA NORMAL UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTH CHINA NORMAL UNIV
Filing Date
2026-03-18
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

In existing technologies, a single model architecture is difficult to integrate diverse diagnostic reasoning perspectives, and cannot effectively alleviate the bias and illusion problems in medical visual question answering, thus limiting the accuracy and reliability of medical visual question answering.

Method used

By combining medical image question pairs and the response text output by multiple specialist physician models, more accurate response text is generated through image enhancement, multimodal feature extraction, and response generation modules, overcoming the bias and illusion defects caused by single-model reasoning.

Benefits of technology

It improves the accuracy and reliability of medical visual question answering answers, and reduces the error and uncertainty of a single model through the collaborative work of multiple doctor models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122388079A_ABST
    Figure CN122388079A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of visual question answering, and particularly relates to a medical visual question answering method, obtaining a medical image question pair to be answered, a plurality of doctor answer texts of the medical image question pair and a preset question and answer model, performing image enhancement on a medical image according to the medical image question pair to be answered, and obtaining an enhanced medical image; combining the enhanced medical image, the medical question and the plurality of doctor answer texts to construct a multi-modal input set; inputting the multi-modal input set into an encoding module for encoding to obtain a multi-modal encoding representation; inputting the multi-modal encoding representation into a multi-modal feature extraction module for feature extraction to obtain a multi-modal feature representation; inputting the multi-modal feature representation into an answer generation module for target answer text generation to obtain a target answer text of the medical image question pair to be answered, overcoming the deviation and illusion defects brought by single model reasoning, and improving the accuracy and reliability of the medical visual question answering answer.
Need to check novelty before this filing date? Find Prior Art