Image-text processing method and device for marking compression framework

By introducing a marker compression framework in multimodal large language models, using attention mechanisms to screen key visual markers and guide clustering and merging through text information, the problems of computational complexity and inference cost in high-resolution images and video processing are solved, and more efficient inference and performance improvements are achieved.

CN120235250AActive Publication Date: 2025-07-01ZHEJIANG YOULU ROBOT TECH CO LTD

Patent Information

Application Number
CN202510715518.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

When processing high-resolution images and videos, the existing multimodal large language models have significantly increased computational complexity and inference costs due to the excessive number of visual markers. The existing accelerated inference methods lack comprehensive consideration of the entire model process, resulting in performance degradation.

Method used

A graphic and text processing method of marking compression framework is provided. Visual mark sequences are generated through visual encoder, and the importance of visual marks is analyzed using CLS attention and local attention mechanisms, the first K key marks are selected, and the clustering and merging of visual marks are guided through text information to reduce the number of visual marks.

Benefits of technology

The inference efficiency of MLLMs is significantly improved, the number of visual markers is reduced, while key visual information and visual-text alignment are retained. Experiments show that while reducing computational costs, model performance is maintained or improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235250A_ABST
    Figure CN120235250A_ABST
Patent Text Reader

Abstract

The invention discloses an image-text processing method and device for marking a compression framework. The method comprises the following steps of: extracting visual features; a visual mark screening processing step; a text feature extraction step; and a multi-modal fusion and model processing step. The method has the beneficial effects that the inference efficiency of the MLLMs is remarkably improved under the condition that the visual mark compression frame does not need additional training; through global and local information fusion of the DVTS module and text guide supplement of the TGVC module, the number of visual marks is greatly reduced, meanwhile, key visual information is reserved, and visual-text alignment is enhanced; experiments show that in various image and video benchmark tests, compared with an existing method, the framework has the advantages that the calculation cost is greatly reduced, the model performance is maintained and even improved, and the framework has remarkable technical advantages and application potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and specifically provides a method and device for processing images and texts based on a marker compression framework. Background Art

[0002] 1. Multimodal large language models (MLLMs) integrate visual signals with language models to achieve the understanding and interaction of content such as images and videos. Its main structure includes a visual encoder, a visual projector, and a base language model. By converting visual information into visual tokens and combining them with text tokens, the input is processed by the language model.

[0003] 2. When existing MLLMs process high-resolution images and videos, due to the excessive number of visual tokens, the computational complexity and inference cost increase significantly. For example, when processing high-resolution images, the number of visual tokens may reach thousands, which greatly increases the computational amount of the model and limits its deployment efficiency in practical applications. In addition, existing acceleration inference methods often only focus on a certain part of the model when reducing the number of visual tokens, such as only optimizing in the visual encoding stage or the language model decoding stage, lacking comprehensive consideration of the entire model process, resulting in performance degradation. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for processing images and texts based on a marker compression framework to solve the problems proposed in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions: In the first aspect, an embodiment of the present invention provides a method for processing images and texts based on a marker compression framework, including: Visual feature extraction step: encoding the input image through a visual encoder to generate a visual token sequence; Visual token screening and processing step: Analyzing the importance of the visual token sequence based on the CLS attention and local attention mechanisms, sorting the visual tokens, and selecting the top K key tokens; Using the remaining visual tokens to guide the clustering and merging of visual tokens with text information, and extracting the feature information of the remaining tokens by repeating the processing N times; Text feature extraction step: encoding the input text question through a text encoder to generate text features; Multimodal fusion and model processing step: fusing the top K key tokens, the remaining tokens processed by guiding the clustering and merging of visual tokens with text information, and the text features, inputting them into a large language model, and performing intermediate layer processing on the fused features to complete the calculation of the image and text task.

[0006] Preferably, the importance of the visual token sequence is analyzed based on the CLS attention and local attention mechanism, the visual tokens are sorted, and the top K key tokens are selected, including: Input the special token CLS and the visual token sequence into the attention mechanism to calculate the attention weights of CLS for each visual token; Apply the attention mechanism within the visual token sequence to calculate the local attention weights of the tokens with surrounding neighboring tokens; Fuse the CLS attention weights and the local attention weights to generate the comprehensive importance scores for each visual token; Sort the visual tokens according to the comprehensive scores, and select the top K tokens with the highest scores as the key tokens, retaining the most representative global semantic information and key local details in the image.

[0007] Preferably, the remaining visual tokens are used to guide the clustering and merging of visual tokens by text information, and the feature information of the remaining tokens is extracted by repeating the processing N times, including: Perform bidirectional similarity calculation on the remaining visual tokens and text tokens; Based on the bidirectional similarity results, sort the remaining visual tokens and select the top R visual tokens as the clustering centers; Assign the remaining visual tokens to the corresponding clustering centers according to the similarity and perform the clustering operation towards the centers; For the visual tokens after clustering and merging, repeatedly execute the process of "similarity calculation → determine new centers → re-cluster and merge" in a loop for N times.

[0008] Preferably, the bidirectional similarity calculation of the remaining visual tokens and text tokens includes: Calculate the similarity from visual tokens to text tokens to measure the matching degree between visual content and text semantics; Calculate the similarity from text tokens to visual tokens to clarify the guiding emphasis direction of the text on visual tokens.

[0009] Preferably, for the visual tokens after clustering and merging, repeatedly execute the process of "similarity calculation → determine new centers → re-cluster and merge" in a loop for N times, including: Each iteration further purifies the visual tokens, gradually strengthens the core features under the guidance of the text, and eliminates irrelevant details; Through multiple rounds of optimization, the visual tokens are made to better fit the text semantic requirements, and finally, visually guided and feature-condensed visual tokens are output.

[0010] Preferably, the intermediate layer processing of the fused features includes: Filter key visual tokens through global semantic importance and local spatial continuity; Then use text information to guide the clustering and merging of visual tokens.

[0011] In a second aspect, an embodiment of the present invention provides a graphic and text processing device for a token compression framework, including: Visual feature extraction module: Encode the input image through a visual encoder to generate a visual token sequence; Visual token screening and processing module: Analyze the importance of the visual token sequence based on CLS attention and local attention mechanisms, sort the visual tokens, and select the top K key tokens; Use text information to guide the clustering and merging of the remaining visual tokens, and extract the feature information of the remaining tokens by repeating the processing N times; Text feature extraction module: Encode the input text question through a text encoder to generate text features; Multi-modal fusion and model processing module: Fuse the top K key tokens, the remaining tokens processed by using text information to guide the clustering and merging of visual tokens, and text features, input them into a large language model, and perform intermediate layer processing on the fused features to complete the graphic and text task calculation.

[0012] In a third aspect, an embodiment of the present invention provides an electronic device, including at least one processor, and the processor is communicatively connected to at least one memory, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of the above embodiments.

[0013] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the method according to any one of the above embodiments when executed.

[0014] Compared with the prior art, the beneficial effects of the present invention are: Without additional training, the visual token compression framework of the present invention significantly improves the inference efficiency of MLLMs; Through the global and local information fusion of the DVTS module and the text-guided supplement of the TGVC module, the number of visual tokens is greatly reduced, while key visual information is retained and visual-text alignment is enhanced; Experiments show that, compared with existing methods, this framework significantly reduces the computational cost while maintaining or even improving the model performance in a variety of image and video benchmark tests, and has significant technical advantages and application potential. Description of the Drawings

[0015] Figure 1It is the flowchart of the method of the embodiment of the present invention; Figure 2 It is the framework diagram of the method of the embodiment of the present invention; Figure 3 It is the flowchart of the visual dominant marker selection of the present invention; Figure 4 It is the device module diagram of the present invention. Specific embodiments

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0017] Please refer to Figure 1 , the present invention provides a technical solution: a graphic and text processing method for a marker compression framework, including: S100: Visual feature extraction step: Encode the input image through a visual encoder to generate a visual marker sequence.

[0018] S200: Visual marker screening and processing step: Analyze the importance of the visual marker sequence based on the CLS attention and local attention mechanisms, sort the visual markers, and select the top K key markers; Use the remaining visual markers to guide the clustering and merging of visual markers with text information, and extract the feature information of the remaining markers by repeating the processing N times.

[0019] S300: Text feature extraction step: Encode the input text question through a text encoder to generate text features.

[0020] S400: Multimodal fusion and model processing step: Fuse the top K key markers, the remaining markers processed by guiding the clustering and merging of visual markers with text information, and the text features, input them into a large language model, and perform intermediate layer processing on the fused features to complete the graphic and text task calculation.

[0021] Please refer to Figure 2 , Figure 2 shows the system framework corresponding to the graphic and text processing method of the marker compression framework in the embodiments of the present invention, and details the process of processing graphic and text features by the above method.

[0022] First is visual information processing. The visual encoder obtains the input image (such as a picture of an airplane) and generates a sequence of visual tokens. The importance of the tokens is analyzed through CLS attention (blue connection) and local attention (red connection) to select visually dominant tokens. The tokens are sorted, and the top K key tokens are selected. The remaining tokens are subjected to feature extraction through the TGVC module (processed N times), and finally, the visual features are integrated through the token mapper.

[0023] Secondly, text information processing. Text encoder: processes the text question (such as "What color is the airplane in the picture?"), generating text features.

[0024] Thirdly, multi-modal fusion and large language model processing. The visual tokens and text features are fused and then input into the large language model. Inside the model, the intermediate layer features are further processed through the DVTS and T6VC modules to integrate visual and language information, and finally, the image-text question answering task is completed.

[0025] The overall architecture embodies the image-text interaction logic of "visual token screening → multi-modal feature fusion → large language model reasoning", aiming to solve the correlation analysis problem of image content understanding and text questions.

[0026] In an embodiment of the present invention, the framework mainly includes two plug-in modules: the Dominant Vision Token Selection (DVTS) module and the Text-Guided Vision Complement (TGVC) module. The DVTS module screens key visual tokens through global semantic importance and local spatial continuity to ensure the integrity of visual information and key details are retained. The TGVC module uses text information to guide the clustering and merging of visual tokens to supplement visual information related to text instructions and enhance the alignment between visual and text representations. These two modules can be seamlessly inserted between any two layers of the visual encoder and the LLM and can play a role in both the visual encoding and LLM decoding stages.

[0027] Specifically, in the DVTS module, the local semantic importance is to extract the attention weights from the penultimate layer of the CLIP-based visual encoder and use the attention distribution of the CLS token to measure the global importance score of each visual token:

[0028] where H is the number of attention heads, is the attention score of the CLS token in the h-th attention head for the i-th visual token.

[0029] Local Spatial Continuity: Local spatial continuity is captured by a local label affinity measurement algorithm. This algorithm uses a dual kernel function to simultaneously consider feature similarity and positional proximity, and calculates the local importance score:

[0030] where F xy and P xy are the feature vector and spatial coordinates of the label at position (x, y) respectively, σf and σp are the standard deviations of feature and position differences, w1, w2 and w3 are balance parameters, and (x, y) and (u, v) are two coordinates in the image feature space. and are the calculated intermediate values of feature similarity and spatial similarity, is the final local importance score obtained by normalizing the intermediate value.

[0031] Adaptive Variance Weighting: Combine the global and local importance scores and use an adaptive variance weighting mechanism to determine the final importance score S i :

[0032] where and are the variances of the global and local importance scores respectively, represents the global importance score, represents the local importance score, and α is the weighting coefficient.

[0033] Specifically, in the TGVC module: Determine the clustering centers: Calculate the similarity between the remaining visual labels and the text features, and select the top R tokens with the highest similarity as the clustering centers:

[0034] where V r is the remaining visual token, T is the text feature, d is the feature dimension.

[0035] Label Assignment: For each remaining token, calculate its assignment score to each clustering center:

[0036] where is the similarity between the visual token and the text, is the clustering center and the similarity with the text.

[0037] Cluster aggregation: For each cluster center, aggregate the assigned tokens through weighted averaging to obtain the final visual supplementary token:

[0038] Repeat the iteration T times to refine the clusters, where represents the j th final visual supplementary token obtained by weighted aggregation of the cluster center, representing the visual feature representation after fusing the relevant information within the cluster, cluster( j ) represents the and the visual token set to form the j th cluster.

[0039] Please refer to Figure 3 , the processing process of TGVC is as follows: 1. Similarity calculation: Build the basis for visual and text association Bidirectional similarity calculation between visual tokens and text tokens: Calculate the similarity between the remaining visual tokens and text tokens, and select the top R visual tokens as the centers based on the similarity results. These center tokens will become the core anchor points for subsequent clustering.

[0040] At the same time, calculate the similarity from text tokens to visual tokens to further clarify the guiding relationship of text information to visual tokens, providing a two-way association basis for clustering.

[0041] 2. Token clustering: Generate supplementary visual tokens Cluster operation towards the center: Using the top R visual tokens selected as the centers, distribute the remaining visual tokens around each center through a clustering algorithm (such as similarity-based aggregation). Each center and its clustered visual tokens form a local set, and by the method of "clustering towards the center", the scattered remaining visual tokens are integrated into a structured supplementary token group.

[0042] Generation of supplementary visual tokens: After clustering, generate supplementary visual tokens with text guiding significance. These tokens not only retain the original visual information but also incorporate the guiding direction of text semantics, providing more accurate visual feature supplementation for subsequent multimodal fusion and large language model calculations.

[0043] In the embodiments of the present invention, based on the CLS attention and local attention mechanisms, analyze the importance of the visual token sequence, sort the visual tokens, and select the top K key tokens, including: Input the special token CLS and the visual token sequence into the attention mechanism, and calculate the attention weights of CLS for each visual token; Apply an attention mechanism to the visual tag sequence and calculate the local attention weights between the tag and its surrounding neighboring tags; The CLS attention weights are fused with the local attention weights to generate a comprehensive importance score for each visual marker; The visual tags are sorted according to the comprehensive scores, and the top K tags with the highest scores are selected as key tags to retain the most representative global semantic information and key local details in the image.

[0044] In an embodiment of the present invention, the remaining visual tags are clustered and merged using text information to guide the visual tags, and feature information of the remaining tags is extracted by repeating the process N times, including: Perform bidirectional similarity calculation on the remaining visual tags and text tags; Based on the bidirectional similarity results, the remaining visual markers are sorted and the first R visual markers are selected as cluster centers; Assign the remaining visual markers to the corresponding cluster centers according to their similarity, and perform clustering operations toward the centers; For the visual markers after clustering and merging, the process of “similarity calculation → determination of new center → re-clustering and merging” is executed again in a loop, and repeated N times.

[0045] In the embodiment of the present invention, the parameter K is the number of key tags, N is the number of clustering iterations, and R is the number of cluster centers. K, N, and R can be set as: K = total number of image tags × 0.2, N = 3, and R = number of remaining tags × 0.1.

[0046] In an embodiment of the present invention, performing bidirectional similarity calculation on the remaining visual mark and the text mark includes: Calculate the similarity between visual tags and text tags to measure the degree of match between visual content and text semantics; Calculate the similarity between text tags and visual tags to clarify the direction in which the text guides the visual tags.

[0047] In an embodiment of the present invention, for the visual markers after clustering and merging, the process of "similarity calculation→determining new center→re-clustering and merging" is cyclically executed again, and repeated N times, including: Each iteration further refines the visual markers, gradually strengthening the core features guided by the text and removing irrelevant details; Through multiple rounds of optimization, the visual markers are made to better meet the semantic requirements of the text, and finally visual markers guided by the text and with concise features are output.

[0048] In an embodiment of the present invention, the fusion features are processed at an intermediate level, including: Filter key visual landmarks by global semantic importance and local spatial continuity; Then, the text information is used to guide the clustering and merging of visual markers.

[0049] Please refer to Figure 4 , an embodiment of the present invention provides a graphic and text processing device for a marker compression framework, including: Visual feature extraction module 100: Encode the input image through a visual encoder to generate a visual marker sequence; Visual marker screening and processing module 200: Analyze the importance of the visual marker sequence based on the CLS attention and local attention mechanisms, sort the visual markers, and select the top K key markers; Use the remaining visual markers to guide the clustering and merging of visual markers with text information, and extract the feature information of the remaining markers by repeating the processing N times; Text feature extraction module 300: Encode the input text question through a text encoder to generate text features; Multi-modal fusion and model processing module 400: Fuse the top K key markers, the remaining markers processed by guiding the clustering and merging of visual markers with text information, and the text features, input them into a large language model, and perform intermediate layer processing on the fused features to complete the graphic and text task calculation.

[0050] An embodiment of the present invention also provides an electronic device, including at least one processor, the processor is communicatively connected to at least one memory, wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of the above embodiments.

[0051] An embodiment of the present invention also provides a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the processor to execute the method according to any one of the above embodiments when executed.

[0052] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A graphic and text processing method for a mark compression framework, characterized in that, Including: Visual feature extraction step: Encode the input image through a visual encoder to generate a sequence of visual tokens; Visual token screening and processing step: Analyze the importance of the visual token sequence based on CLS attention and local attention mechanisms, sort the visual tokens, and select the top K key tokens; Use text information to guide the clustering and merging of the remaining visual tokens, and extract the feature information of the remaining tokens through N repeated processes; Text feature extraction step: Encode the input text question through a text encoder to generate text features; Multi-modal fusion and model processing step: Fuse the top K key tokens, the remaining tokens after using text information to guide the clustering and merging of visual tokens, and text features, input them into a large language model, and perform intermediate layer processing on the fused features to complete the calculation of the graphic and text task.

2. The graphic and text processing method of a marker compression framework according to claim 1, characterized in that: The step of analyzing the importance of the visual token sequence based on CLS attention and local attention mechanisms, sorting the visual tokens, and selecting the top K key tokens includes: Input the special token CLS and the visual token sequence into the attention mechanism to calculate the attention weights of CLS for each visual token; Apply the attention mechanism inside the visual token sequence to calculate the local attention weights of the tokens and their surrounding adjacent tokens; Fuse the CLS attention weights and local attention weights to generate the comprehensive importance scores of each visual token; Sort the visual tokens according to the comprehensive scores, and select the top K tokens with the highest scores as key tokens, retaining the most representative global semantic information and key local details in the image.

3. The graphic and text processing method of a marker compression framework according to claim 1, characterized in that: The step of using text information to guide the clustering and merging of the remaining visual tokens and extracting the feature information of the remaining tokens through N repeated processes includes: Perform bidirectional similarity calculation on the remaining visual tokens and text tokens; Based on the bidirectional similarity results, sort the remaining visual tokens and select the top R visual tokens as clustering centers; Assign the remaining visual tokens to the corresponding clustering centers according to the similarity and perform the clustering operation towards the center; For the visual tokens after clustering and merging, repeatedly execute the process of "similarity calculation → determine new centers → re-cluster and merge" again, repeating N times.

4. A method for processing graphics and texts of a tag compression framework according to claim 3, characterized in that: The step of performing bidirectional similarity calculation on the remaining visual tokens and text tokens includes: Calculate the similarity from visual tokens to text tokens to measure the matching degree between visual content and text semantics; Calculate the similarity from text tokens to visual tokens to clarify the guiding emphasis direction of the text on visual tokens.

5. A method for processing graphics and texts of a marker compression framework according to claim 1, characterized in that: The step of repeatedly executing the process of "similarity calculation → determine new centers → re-cluster and merge" again for the visual tokens after clustering and merging, repeating N times includes: Each iteration further purifies the visual tokens, gradually strengthens the core features under text guidance, and eliminates irrelevant details; Through multiple rounds of optimization, make the visual tokens more in line with the text semantic requirements, and finally output the text-guided and feature-condensed visual tokens.

6. The graphic and text processing method of a tag compression framework according to claim 1, characterized in that: The step of performing intermediate layer processing on the fused features includes: Screen key visual tokens through global semantic importance and local spatial continuity; Then use text information to guide the clustering and merging of visual tokens.

7. An image and text processing device for a tag compression framework, characterized in that, Including: Visual feature extraction module: Encodes the input image through a visual encoder to generate a sequence of visual tokens; Visual token screening and processing module: Analyzes the importance of the visual token sequence based on CLS attention and local attention mechanisms, sorts the visual tokens, and selects the top K key tokens; Uses text information to guide the clustering and merging of the remaining visual tokens, and extracts the feature information of the remaining tokens by repeating the processing N times; Text feature extraction module: Encodes the input text question using a text encoder to generate text features; Multimodal fusion and model processing module: Fuses the top K key tokens, the remaining tokens after clustering and merging the visual tokens guided by text information, and the text features, inputs them into a large language model, and performs intermediate layer processing on the fused features to complete the calculation of the image-text task.

8. An electronic device, characterized in that, Includes at least one processor, and the processor is communicatively connected to at least one memory. Among them, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to implement the method according to any one of claims 1 to 6 when executed.

Citation Information

Patent Citations

  • High-efficiency high-resolution image visual mark generation method for multi-modal large model

    CN118735932A

  • Adaptive visual mark pruning method and device based on multi-modal large model

    CN119204137A

  • Multi-modal dialogue abstract method based on multi-level visual guidance

    CN119918545A

  • Cross-attention system and method for fast video-text retrieval task with image clip

    WO2022261570A1

  • Visual tokenization with language models

    WO2025072952A1

Cited By

  • Cross-modal image-text analysis method for machine vision

    CN121210958A

  • Machine vision-oriented cross-modal graphic-text analysis method

    CN121210958B

  • Non-uniform image data processing method and device based on visual language model

    CN122176733A

  • Non-uniform image data processing method and device based on visual language model

    CN122176733B