Medical visual question answering method and device, computer device and storage medium
By combining medical image question pairs and the response text output by multiple specialist physician models, and utilizing image enhancement and multimodal feature extraction modules, more accurate medical visual question-answering answers are generated. This solves the bias and illusion problems of single models in medical visual question answering, and improves the accuracy and reliability of the answers.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTH CHINA NORMAL UNIV
- Filing Date
- 2026-03-18
- Publication Date
- 2026-07-14
AI Technical Summary
In existing technologies, a single model architecture is difficult to integrate diverse diagnostic reasoning perspectives, and cannot effectively alleviate the bias and illusion problems in medical visual question answering, thus limiting the accuracy and reliability of medical visual question answering.
By combining medical image question pairs and the response text output by multiple specialist physician models, more accurate response text is generated through image enhancement, multimodal feature extraction, and response generation modules, overcoming the bias and illusion defects caused by single-model reasoning.
It improves the accuracy and reliability of medical visual question answering answers, and reduces the error and uncertainty of a single model through the collaborative work of multiple doctor models.
Smart Images

Figure CN122388079A_ABST