A visual language large model attribution method based on semantic perception optimization

By co-designing multi-scale interpretation aggregation and activation ranking relevance modules, the problem of insufficient spatial context capture and pre-token interference in visual question answering and image description tasks of multimodal large language models is solved. High-quality, noise-resistant attribution graphs are generated, improving the fidelity and robustness of visual attribution.

CN122416147APending Publication Date: 2026-07-17SUN YAT SEN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUN YAT SEN UNIV
Filing Date
2026-05-23
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing multimodal large language models lack spatial context capture and pre-token interference suppression in visual question answering and image description tasks, resulting in poor spatial coherence and severe noise pollution in attribution graphs, making it difficult to accurately reflect the causal relationship between visual evidence and text output.

Method used

A collaborative design of Multiscale Explanation Aggregation (MSEA) and Activation Ranking Relevance (ARC) modules is adopted. Through multiscale attribution graph fusion and semantic relevance quantification, aggregated attribution graphs are generated and interference is suppressed to ensure that the explanation results are strongly correlated with the target semantics.

Benefits of technology

It significantly enhances the spatial coherence and semantic focusing capabilities of attribution graphs, substantially improves the fidelity and robustness of visual attribution for multimodal large models, is applicable to different architectures and tasks, and has controllable computational costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416147A_ABST
    Figure CN122416147A_ABST
Patent Text Reader

Abstract

The application provides a visual language large model attribution method based on semantic perception optimization, comprising the following steps: obtaining an input image, and obtaining a visual token set after model tokenization processing; calculating an attribution score of each visual token to a target text token, arranging the spatial position in the original image to form a single-scale attribution map; generating different resolution versions by performing multi-scale scaling on the input image, generating an aggregated attribution map after generating a single-scale attribution map for each version; calculating the output probability distribution and semantic correlation score of the pre-token and the target token, and generating an interference aggregated map by weighted aggregation of the aggregated attribution map of the pre-token; determining an adaptive suppression intensity by least squares optimization, subtracting the corresponding multiple of the interference aggregated map from the aggregated attribution map to obtain a final attribution map. The application solves the problems of insufficient spatial context capture and ineffective suppression of pre-token interference in the visual attribution method by generating an attribution map through decoding the visual token hidden state.
Need to check novelty before this filing date? Find Prior Art