基于多模态大语言模型的交互方法、电子设备及存储介质

By training a multimodal large language model with training sample data aligned by visual dependence and evidence consistency, the problem of insufficient visual information understanding is solved, the accuracy and credibility of visual interaction tasks are improved, and the phenomenon of language prior answering is suppressed.

CN122414418APending Publication Date: 2026-07-17ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG GEELY HLDG GRP CO LTD
Filing Date
2026-06-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Multimodal large language models have significant technical deficiencies in the deep understanding of visual information, making it difficult to interact based on real visual information. This leads to language priors and "stealing answers," failing to truly reflect the objective information in the visual scene, especially in open-domain question answering tasks.

Method used

By using training sample data driven by visual dependence and evidence consistency, we can perform aligned training on multimodal large language models to improve their accuracy and credibility in visual interaction tasks, construct strong visual quantitative scales and evidence consistency evaluation, and suppress language prior eavesdropping.

Benefits of technology

It enhances the perception and reasoning capabilities of multimodal large language models in tasks with high visual dependence, ensures that interaction results are based on real visual information, and improves the accuracy and credibility of visual interaction tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122414418A_ABST
    Figure CN122414418A_ABST
Patent Text Reader

Abstract

本申请提供一种基于多模态大语言模型的交互方法、电子设备及存储介质,涉及人工智能技术领域,通过获取图像数据和与所述图像数据相关联的提问数据;基于预设的多模态大语言模型对所述图像数据和所述提问数据进行视觉交互任务处理,得到所述提问数据的多模态交互结果;所述多模态大语言模型基于视觉依赖性和证据一致性驱动的训练样本数据进行对齐训练得到,所述视觉依赖性用于表征视觉交互任务对视觉证据的依赖程度,所述证据一致性用于表征多模态交互结果与视觉证据之间的对齐程度。采用本申请能够提升多模态大语言模型执行视觉交互任务的准确性与可信度。
Need to check novelty before this filing date? Find Prior Art