基于多模态大语言模型的交互方法、电子设备及存储介质
By training a multimodal large language model with training sample data aligned by visual dependence and evidence consistency, the problem of insufficient visual information understanding is solved, the accuracy and credibility of visual interaction tasks are improved, and the phenomenon of language prior answering is suppressed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG GEELY HLDG GRP CO LTD
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-17
AI Technical Summary
Multimodal large language models have significant technical deficiencies in the deep understanding of visual information, making it difficult to interact based on real visual information. This leads to language priors and "stealing answers," failing to truly reflect the objective information in the visual scene, especially in open-domain question answering tasks.
By using training sample data driven by visual dependence and evidence consistency, we can perform aligned training on multimodal large language models to improve their accuracy and credibility in visual interaction tasks, construct strong visual quantitative scales and evidence consistency evaluation, and suppress language prior eavesdropping.
It enhances the perception and reasoning capabilities of multimodal large language models in tasks with high visual dependence, ensures that interaction results are based on real visual information, and improves the accuracy and credibility of visual interaction tasks.
Smart Images

Figure CN122414418A_ABST