A method, system, and storage medium for clinical diagnostic medical visual question answering based on multi-margin collaborative debiasing.
By conducting multi-level analysis of the medical visual question answering model and adding MAM and DCL modules, the multimodal representation was optimized, which solved the bias problem of the model in out-of-distribution clinical diagnosis and improved the accuracy and robustness of medical visual question answering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
- Filing Date
- 2026-04-16
- Publication Date
- 2026-06-02
AI Technical Summary
Existing medical visual question answering models tend to over-rely on text questions in out-of-distribution clinical diagnosis, ignoring key lesion features in medical images, leading to medical language bias and affecting the reliability and accuracy of diagnosis, especially in the diagnosis of rare diseases or complex cases.
This paper theoretically analyzes the formation mechanism of medical language bias from three aspects: modal gradient imbalance, feature fusion bias, and classifier weight direction bias. The MAM mechanism and DCL module are added to the UpDn model to optimize the multimodal representation of the model and enhance its robustness and accuracy.
It effectively alleviates medical language bias, improves the model's prediction accuracy and robustness on out-of-distribution clinical diagnostic data, and enhances diagnostic performance for rare or complex cases.
Smart Images

Figure CN122135946A_ABST