Multi-modal large language model generation method and system based on relative attention cutting and judgment enhancement

By employing relative attention pruning and external judgment enhancement mechanisms, the problems of fine-grained visual information capture and hallucination prediction in multimodal large language models are solved, thereby improving the model's visual understanding ability and stability.

CN120873111APending Publication Date: 2025-10-31HUBEI UNIV OF EDUCATION
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510734299.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing multimodal large language models struggle to capture fine-grained visual information in images, are prone to producing hallucination predictions, and lack dynamic error correction mechanisms, resulting in low model reliability.

Method used

By employing relative attention pruning and external judgment enhancement mechanisms, the initial answer is corrected through relative attention pruning and external guidance models, thereby improving fine-grained visual perception capabilities and dynamically correcting erroneous predictions.

Benefits of technology

It significantly improves the model's fine-grained visual understanding ability, suppresses hallucination prediction, and enhances stability and robustness in complex visual scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873111A_ABST
    Figure CN120873111A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal large language model generation method and system based on relative attention cutting and judgment enhancement, and belongs to the technical field of artificial intelligence, and the method comprises the steps: S1, inputting text features and text questions obtained after input image conversion into a large language model together, generating an initial answer, and obtaining an initial answer; comparing the initial answer with a standard answer by adopting an external guidance model to obtain an initial judgment signal; s2, calculating and cutting an input image, a text problem and a general problem to obtain a relative attention graph; and S3, based on the input image, the text question, the standard answer, the initial judgment signal and the relative attention map, correcting the initial answer until the large language model converges. According to the method, the fine-grained visual understanding capability is remarkably improved, the illusion prediction problem is inhibited, and the stability and robustness in a complex visual scene are improved.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Cited By

  • Multi-modal large model illusion detection and suppression method based on attention time sequence difference

    CN121438067A