Multi-modal large language model generation method and system based on relative attention cutting and judgment enhancement
By employing relative attention pruning and external judgment enhancement mechanisms, the problems of fine-grained visual information capture and hallucination prediction in multimodal large language models are solved, thereby improving the model's visual understanding ability and stability.
CN120873111APending Publication Date: 2025-10-31HUBEI UNIV OF EDUCATION
Patent Information
- Application Number
- CN202510734299.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-10-31
AI Technical Summary
Technical Problem
Existing multimodal large language models struggle to capture fine-grained visual information in images, are prone to producing hallucination predictions, and lack dynamic error correction mechanisms, resulting in low model reliability.
Method used
By employing relative attention pruning and external judgment enhancement mechanisms, the initial answer is corrected through relative attention pruning and external guidance models, thereby improving fine-grained visual perception capabilities and dynamically correcting erroneous predictions.
Benefits of technology
It significantly improves the model's fine-grained visual understanding ability, suppresses hallucination prediction, and enhances stability and robustness in complex visual scenes.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN120873111A_ABST
Abstract
The invention discloses a multi-modal large language model generation method and system based on relative attention cutting and judgment enhancement, and belongs to the technical field of artificial intelligence, and the method comprises the steps: S1, inputting text features and text questions obtained after input image conversion into a large language model together, generating an initial answer, and obtaining an initial answer; comparing the initial answer with a standard answer by adopting an external guidance model to obtain an initial judgment signal; s2, calculating and cutting an input image, a text problem and a general problem to obtain a relative attention graph; and S3, based on the input image, the text question, the standard answer, the initial judgment signal and the relative attention map, correcting the initial answer until the large language model converges. According to the method, the fine-grained visual understanding capability is remarkably improved, the illusion prediction problem is inhibited, and the stability and robustness in a complex visual scene are improved.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Cited By
Multi-modal large model illusion detection and suppression method based on attention time sequence difference
CN121438067A