A visual language large model attribution method based on semantic perception optimization
By co-designing multi-scale interpretation aggregation and activation ranking relevance modules, the problem of insufficient spatial context capture and pre-token interference in visual question answering and image description tasks of multimodal large language models is solved. High-quality, noise-resistant attribution graphs are generated, improving the fidelity and robustness of visual attribution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUN YAT SEN UNIV
- Filing Date
- 2026-05-23
- Publication Date
- 2026-07-17
AI Technical Summary
Existing multimodal large language models lack spatial context capture and pre-token interference suppression in visual question answering and image description tasks, resulting in poor spatial coherence and severe noise pollution in attribution graphs, making it difficult to accurately reflect the causal relationship between visual evidence and text output.
A collaborative design of Multiscale Explanation Aggregation (MSEA) and Activation Ranking Relevance (ARC) modules is adopted. Through multiscale attribution graph fusion and semantic relevance quantification, aggregated attribution graphs are generated and interference is suppressed to ensure that the explanation results are strongly correlated with the target semantics.
It significantly enhances the spatial coherence and semantic focusing capabilities of attribution graphs, substantially improves the fidelity and robustness of visual attribution for multimodal large models, is applicable to different architectures and tasks, and has controllable computational costs.
Smart Images

Figure CN122416147A_ABST