Multi-modal large model reasoning method and device
By constructing a comprehensive importance matrix of cross-modal and self-attention, and optimizing visual token selection, the problem that visual token selection cannot take into account visual information and text semantics in multimodal large models is solved, and the effective pruning of visual tokens is realized, reducing computational burden and improving model performance.
Patent Information
- Application Number
- CN202510856642.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When selecting visual tokens, the existing multimodal large models cannot consider the significance of visual information and the correlation between text semantics at the same time, resulting in the inability to effectively prune, which increases the computational overhead.
By calculating the cross-modal attention and self-attention between the text vector and the visual token, a comprehensive importance matrix is constructed, and the selection of visual tokens is optimized by combining the mask matrix and the objective function to realize the pruning of the cross-modal and self-attention selection token set.
Effective pruning of visual tokens is implemented, reducing computational burden, improving model performance, and maintaining performance close to the original model when retaining a small number of visual tokens.
Smart Images

Figure CN120354955A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a multi-modal large model inference method and device. Background Art
[0002] In recent years, vision-text large models have performed well in various vision-related tasks, such as visual question answering (VQA), image caption generation, etc. These tasks require the model to simultaneously understand and process visual information (images) and language information (text). Vision-text large models typically consist of three main parts: a visual encoder, a text encoder, and a language model (LLM). The visual encoder is responsible for processing images and generating visual tokens, the text encoder processes text prompts, and the LLM is responsible for generating the final output.
[0003] Existing vision-text large models usually need to process a large number of visual tokens. For example, the LLaVA-NEXT model uses 2,880 visual tokens in single-image tasks, which is significantly higher than the number of text prompt words in conventional tasks, resulting in significant computational overhead and limiting the efficiency of the model in actual deployment.
[0004] Currently, methods such as FastV and SparseVLM select visual tokens by extracting self-attention weights from the LLM layer, but these methods ignore text semantics and are vulnerable to the influence of biased self-attention in the LLM. In addition, FasterVLM proposes using category attention in the visual encoder as an importance metric for visual token pruning, but this method fails to recall non-significant but semantically relevant visual information. Moreover, the text-to-image attention (cross-modal attention) extracted by existing methods (such as FastV and SparseVLM) in the LLM layer is biased and usually concentrates on the end part of the input image, resulting in the inability to accurately identify visual tokens related to text prompts. The root cause of this bias problem is that the attention mechanism in the LLM layer does not fully consider the semantic alignment between text and image.
[0005] In summary, during the inference process of multi-modal large models, it is impossible to select visual tokens that can simultaneously consider the saliency of visual information and the relevance of text semantics, and it is impossible to effectively prune visual tokens. Summary of the Invention
[0006] An embodiment of the present application provides a multi-modal large model inference method and device, which are used to solve the problem that in the process of multi-modal large model inference, the selection of visual tokens cannot simultaneously consider the saliency of visual information and the relevance of text semantics, and it is impossible to effectively prune visual tokens.
[0007] The embodiments of the present application adopt the following technical solutions: On the one hand, an embodiment of the present application provides a multi-modal large model inference method, which includes: determining the cross-modal attention of each visual token for the text prompt according to the semantic relevance between the text vector of the text prompt and each visual token corresponding to the input image; respectively sorting the self-attention and cross-modal attention of each visual token in descending order, and respectively performing cumulative summation on the sorted self-attention and cross-modal attention to obtain a self-attention cumulative sum vector and a cross-modal attention cumulative sum vector; constructing a comprehensive importance matrix according to the self-attention cumulative sum vector and the cross-modal attention cumulative sum vector; and constructing a mask matrix according to the token selection quantity threshold of the input image; performing an element-wise product of the comprehensive importance matrix and the mask matrix to obtain the comprehensive importance scores corresponding to the cross-modal selected token set and the self-attention selected token set; wherein, the number of cross-modal selected tokens and the number of self-attention selected tokens are respectively determined by the row index and column index of the comprehensive importance score; optimizing and solving the cross-modal selected token set and the self-attention selected token set that satisfy the constraint conditions when the objective function is at the maximum value according to the maximization of the comprehensive importance score and a preset objective function; the objective function is used to represent the comprehensive representation of maximizing cross-modal semantic relevance and image internal importance; performing a redirected selection on the multiple visual tokens of the input image according to the cross-modal selected token set and the self-attention selected token set to prune the multiple visual tokens; and inputting the remaining visual tokens and the text vector into a large language model for inference.
[0008] In one example, the determining the cross-modal attention of each visual token for the text prompt according to the semantic relevance between the text vector and each visual token specifically includes: calculating the semantic similarity between the text vector and each visual token, and normalizing the semantic similarity to obtain the initial cross-modal attention; aligning the initial cross-modal attention distribution with the self-attention distribution to obtain the cross-modal attention of each visual token.
[0009] In one example, the calculating the semantic similarity between the text vector and each visual token specifically includes: calculating the cosine value between the text vector and each visual token; and determining the cosine value as the semantic similarity between the text vector and each visual token.
[0010] In one example, a comprehensive importance matrix is constructed based on the self-attention cumulative sum vector and the cross-modal attention cumulative sum vector, specifically including: transposing the self-attention cumulative sum vector, and multiplying the transposed self-attention cumulative sum vector by the cross-modal attention to obtain the comprehensive importance matrix.
[0011] In one example, a mask matrix is constructed according to the token selection quantity threshold of the input image, specifically including: summing the row number and column number where the position of the mask matrix is located; when the sum result is less than or equal to the token selection quantity threshold, filling the element value at the corresponding position with 1; when the sum result is greater than the token selection quantity threshold, filling the element value at the corresponding position with 0.
[0012] In one example, according to the maximization of the comprehensive importance score and the preset objective function, the cross-modal selection token set and the self-attention selection token set that satisfy the constraint conditions when the objective function is at its maximum are optimized and solved, specifically including: constructing a geometric mean form objective function based on the cross-modal attention sum of the selected cross-modal selection token set and the self-attention sum of the selected self-attention selection token set; under the condition that the constraint is that the sum between the size of the cross-modal selection token set and the size of the self-attention selection token set is equal to the token selection quantity threshold, optimizing the cross-modal selection token set and the self-attention selection token set by maximizing the comprehensive importance score to solve the optimal solution of the objective function.
[0013] In one example, the method further includes: determining the value of the row index as the number of cross-modal selection tokens, and determining the value of the column index as the number of self-attention selection tokens; sequentially extracting the top-ranked visual tokens from the descending order of cross-modal attention according to the number of cross-modal selection tokens to obtain the cross-modal selection token set; sequentially extracting the top-ranked visual tokens from the descending order of self-attention according to the number of self-attention selection tokens to obtain the self-attention selection token set.
[0014] In one example, according to the maximization of the comprehensive importance score and the preset objective function, the cross-modal selection token set and the self-attention selection token set that satisfy the constraint conditions when the objective function is at its maximum are optimized and solved, specifically including:
[0015]
[0016] Among them, represents the cross-modal attention of the th token in the unsorted list, represents the self-attention of the th token in the unsorted list, represents the size of the cross-modal selection token set, Represents the size of the set of tokens selected by self-attention, Represents the token selection quantity threshold of the input image, Represents the comprehensive importance matrix, M represents the mask matrix, Are respectively at Row When the comprehensive importance scores of the column elements are the largest, the number of tokens in the cross-modal selected token set and the number of tokens in the self-attention selected token set, , V = , = U, = V.
[0017] In one example, the redirecting and selecting of multiple visual tokens of the input image according to the cross-modal selected token set and the self-attention selected token set specifically includes: taking the union of the cross-modal selected token set and the self-attention selected token set to obtain the selected set of retained visual tokens, and pruning other un-retained tokens.
[0018] On the other hand, the embodiments of the present application provide a multi-modal large model inference device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a multi-modal large model inference method described in any one of the above.
[0019] The above at least one technical solution adopted by the embodiments of the present application can achieve the following beneficial effects: Introduce the cross-modal attention of each visual token for the text prompt. The higher the cross-modal attention of a token, the stronger the semantic correlation between the token and the text, which is important for understanding the text-image alignment. At the same time, calculate the self-attention of each visual token. The higher the self-attention score of a token, the more important information the token has within the image, which is important for understanding the image content.
[0020] On this basis, by cumulative summation and descending order sorting, the combination problem (the problem of two attention selected token sets) can be transformed into a computable form, which can avoid combinatorial explosion and dimensionality reduction. In addition, when sorting in descending order, the sum of the first i tokens is actually the maximum possible sum for that i quantity, which can exclude obviously impossible sets.
[0021] By constructing a mask matrix, it can ensure the effectiveness of the token selection range, that is, ensure that the total number of finally selected tokens calculated does not exceed the token selection quantity threshold, that is, does not exceed the total number of visual tokens of the input image. When the mask returns to zero, illegal solutions are automatically excluded.
[0022] Under the constraints of the above mask matrix, the number of cross-modal selection tokens and the number of self-attention selection tokens are determined by comprehensively considering the row index and column index of the importance scores respectively. Thus, within the effective selection range, the matrix coordinates are set as the optimal encoding of the resource allocation scheme. The optimization process is to locate the best coordinate point in the two-dimensional resource allocation space. The row-column index method transforms the discrete problem into a differentiable matrix extreme value search. This design transforms the combinatorial optimization into a solvable problem, which ensures optimality more than the traditional greedy algorithm.
[0023] By maximizing the comprehensive importance score and solving the optimal solution of the objective function, the visual token selection is dynamically optimized. Thus, a set of cross-modal selection tokens that are relatively important to the text and a set of self-attention selection tokens that are relatively important within the image are screened out. Through the union of the two, the visual tokens that are important both to the text semantics and within the image are retained, thus taking into account the information in both aspects.
[0024] In summary, the visual attention is redirected from the original distribution based only on visual information to the distribution that also considers text semantics, enabling the selection of visual tokens to consider both the saliency of visual information and the relevance of text semantics. The effective pruning of visual tokens is achieved, reducing the computational burden while improving the performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] To illustrate the technical solutions of the present application more clearly, some embodiments of the present application will be described in detail below with reference to the accompanying drawings. In the drawings: Figure 1 is a schematic flowchart of a multi-modal large model inference method provided by an embodiment of the present application; Figure 2 is a schematic framework diagram of a multi-modal large model inference method provided by an embodiment of the present application; Figure 3 is a schematic structural diagram of a multi-modal large model inference device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below in conjunction with specific embodiments and the corresponding accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0027] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0028] Figure 1A flowchart of a multi-modal large model inference method provided by an embodiment of this application. This method can be applied to different business fields, such as the Internet finance business field, the e-commerce business field, the instant messaging business field, the game business field, the official business field, etc. Some input parameters or intermediate results in this process allow manual intervention and adjustment to help improve accuracy.
[0029] The implementation of the analysis method involved in an embodiment of this application can be a terminal device or a server, and this application does not make special restrictions on this. For the convenience of understanding and description, the following embodiments will be described in detail taking the server as an example.
[0030] It should be noted that this server can be a single device or a system composed of multiple devices, that is, a distributed server, and this application does not make specific limitations on this.
[0031] Figure 1 The process in it includes the following steps: S101: Determine the cross-modal attention of each visual token for the text prompt according to the semantic correlation between the text vector of the text prompt and each visual token corresponding to the input image.
[0032] In some embodiments of this application, the text prompt is converted into a text embedding through a text encoder to obtain a text vector.
[0033] Among them, the process of determining the cross-modal attention of each visual token for the text prompt is as follows: First, calculate the semantic similarity between the text vector and each visual token, and normalize the semantic similarity to obtain the initial cross-modal attention.
[0034] Then, align the initial cross-modal attention distribution with the self-attention distribution to obtain the cross-modal attention of each visual token.
[0035] It should be noted that by calculating the similarity between the text embedding (text vector) and the visual embedding (visual token), the semantic association between the text prompt and the input image is determined. The similarity calculation can adopt methods such as dot product and cosine similarity. The dot product calculates the dot product between the text embedding vector and the visual embedding vector, and the cosine similarity calculates the cosine value of the included angle between them.
[0036] Normalize similarity: In order to eliminate the differences in the data distributions of different modalities, the similarity is normalized. For example, use the Softmax operation to normalize the similarity between the text prompt and the input image and align it with the self-attention distribution.
[0037] S102: Descendingly sort the self-attention and cross-modal attention of each visual token respectively, and perform cumulative summation on the sorted self-attention and cross-modal attention respectively to obtain a self-attention cumulative sum vector and a cross-modal attention cumulative sum vector.
[0038] It should be noted that for convenience of selection, descending sorting can be performed. In addition, the step of cumulative summation is to gradually add a set of data in sequence, that is, starting from the first element, successively add the current element to the sum of all previous elements to form a new cumulative sequence.
[0039] Among them, through cumulative summation and descending sorting, the combinatorial problem (the problem of selecting token sets for two kinds of attention) can be transformed into a computable form, which can avoid combinatorial explosion and the dimension reduction effect. In addition, when sorting in descending order, the sum of the first i tokens is actually the maximum possible sum for that i quantity, which can exclude obviously impossible sets.
[0040] In some embodiments of the present application, the formula for cumulative summation is as follows:
[0041] Among them, represents the cumulative sum of the first 0 cross-modal attention sorted tokens, represents the cumulative sum of the first 0 self-attention sorted tokens, represents the cumulative sum of the first t cross-modal attention sorted tokens, represents the th cross-modal attention of the visual token sorted in descending order, represents the cumulative sum of the first t self-attention sorted tokens, represents the th self-attention of the visual token sorted in descending order.
[0042] In some embodiments of the present application, the input image is encoded by a visual encoder, and key visual information is extracted through the self-attention layer therein. That is, during the encoding process, the self-attention layer calculates the mutual relationship between each visual token to generate a feature representation of the image. These self-attention weights reflect the interaction and importance between different parts inside the image. Information between different visual tokens is transmitted based on the self-attention between them, and the CLS token integrates the information of all visual tokens.
[0043] It should be noted that after the encoding is completed, the visual attention of the input image can be extracted, and the visual attention is extracted from the self-attention layer of the visual encoder. The process is as follows: Generally, visual attention is defined as the magnitude of their impact on the final output. Among them, the extraction of visual attention in this application can be achieved by calculating the self-attention weights between the CLS token and the image patches (visual tokens). The CLS token is usually used to represent global image information in the visual Transformer model. Therefore, the self-attention weights between it and each image patch can reflect the contribution of each image patch to the global information, that is, visual attention. These weights constitute a measure of the importance of visual tokens. For example, the importance of each visual token to the CLS token can be extracted from the penultimate self-attention layer of the visual encoder and used as their initial visual attention.
[0044] S103: Construct a comprehensive importance matrix based on the self-attention cumulative sum vector and the cross-modal attention cumulative sum vector; and construct a mask matrix according to the token selection quantity threshold of the input image.
[0045] In some embodiments of this application, the process of constructing the comprehensive importance matrix is as follows: Transpose the self-attention cumulative sum vector, and multiply the transposed self-attention cumulative sum vector by the cross-modal attention to obtain the comprehensive importance matrix.
[0046] It should be noted that the expression is as follows:
[0047] Among them, represents the comprehensive importance matrix, represents the cross-modal attention cumulative sum vector, represents the self-attention cumulative sum vector.
[0048] In some embodiments of this application, the process of constructing the mask matrix is as follows: Sum the row number and column number of the position of the mask matrix. Then, when the sum result is less than or equal to the token selection quantity threshold, fill the element value at the corresponding position with 1.
[0049] When the sum result is greater than the token selection quantity threshold, fill the element value at the corresponding position with 0.
[0050] It should be noted that the token selection quantity threshold can be the number of visual tokens of the input image.
[0051] Among them, the expression for constructing the mask matrix is as follows:
[0052] Among them, M represents the mask matrix. For the element M at row and column, at When it is determined that it can be selected, the value is 1; otherwise, it is 0. K is the token selection quantity threshold.
[0053] It should be noted that the mask matrix is used to limit the selection range to ensure that the selection range is valid, that is, to ensure that the total number of selected tokens does not exceed K. In addition, and are integer indices. Since the size of the mask matrix is determined by the number of visual tokens. For example, when the number of visual tokens is 5, the size of the mask matrix is 5×5.
[0054] Therefore, for the row and the column, actually it can be considered corresponding to the th visual token in the descending order of cross-modal attention, it can be considered corresponding to the th visual token in the descending order of self-attention.
[0055] S104: Perform an element-wise product of the comprehensive importance matrix and the mask matrix to obtain the comprehensive importance scores corresponding to the cross-modal selection token set and the self-attention selection token set, where the number of cross-modal selection tokens and the number of self-attention selection tokens are determined by the row index and column index of the comprehensive importance scores respectively.
[0056] For example, perform an element-wise product for the row and the column. The product result is equivalent to the comprehensive importance score corresponding to the th visual token before the descending order of cross-modal attention and the th visual token before the descending order of self-attention, that is, the number of cross-modal selection tokens is and the number of self-attention selection tokens is .
[0057] In some embodiments of the present application, the value of the row index is determined as the number of cross-modal selection tokens, and the value of the column index is determined as the number of self-attention selection tokens.
[0058] According to the number of cross-modal selection tokens, sequentially extract the top-ranked visual tokens from the descending order of cross-modal attention to obtain the cross-modal selection token set.
[0059] According to the number of self-attention selection tokens, sequentially extract the top-ranked visual tokens from the descending order of self-attention to obtain the self-attention selection token set.
[0060] S105: According to the comprehensive importance score and the preset objective function, optimize and solve the cross-modal selection token set and the self-attention selection token set that satisfy the constraint conditions when the objective function is at its maximum; the objective function is used to represent the comprehensive representation of maximizing cross-modal semantic relevance and image internal importance.
[0061] In some implementations of the present application, the purpose is to find the retained visual tokens with the maximum comprehensive importance through the comprehensive importance score. The specific process is as follows: Based on the sum of cross-modal attention of all visual tokens in the selected cross-modal selection token set and the text vector, and the sum of self-attention of all visual tokens in the selected self-attention selection token set and the CLS token, construct a geometric mean form objective function.
[0062] Under the constraint condition that the sum of the size of the cross-modal selection token set and the size of the self-attention selection token set is equal to the token selection quantity threshold, optimize the cross-modal selection token set and the self-attention selection token set by maximizing the comprehensive importance score to solve the optimal solution of the objective function.
[0063] Among them, the expression is as follows:
[0064]
[0065] Among them, represents the cross-modal attention of the th unsorted token, represents the self-attention of the th unsorted token, represents the size of the cross-modal selection token set, represents the size of the self-attention selection token set, represents the token selection quantity threshold of the input image, represents the comprehensive importance matrix, M represents the mask matrix, are respectively the number of tokens in the cross-modal selection token set and the number of tokens in the self-attention selection token set when the comprehensive importance score of the element in the th row and th column is the largest, , V = , = U, = V.
[0066] It should be noted that the set size refers to the number of selected tokens. The cross-modal selection token set is actually a set of visual tokens selected based on cross-modal attention, and the self-attention selection token set is a set of visual tokens selected based on self-attention.
[0067] It should be noted that since the square root function is monotonically increasing, when the objective function is maximized, it is equivalent to maximizing the product of the sum of cross-modal attention and the sum of self-attention, that is, maximizing the product of the cumulative sum of the top U tokens of cross-modal attention and the cumulative sum of the top V tokens of self-attention. And the comprehensive importance matrix is obtained by multiplying the cumulative sums between self-attention and cross-modal attention. Therefore, finding the maximum value of the comprehensive importance matrix is essentially finding the maximization of the above product.
[0068] Therefore, when obtaining the maximum value of the comprehensive importance matrix, it is actually the maximum value of the objective function. At this time, it is only necessary to further satisfy the constraint conditions.
[0069] In summary, by combining the coordinates of the maximum value of the comprehensive importance matrix to determine the optimal number of selections, by setting , V = , it also realizes maximizing the importance under .
[0070] In addition, in order to comprehensively consider the visual attention of images and the cross-modal attention from text to images, the objective function uses the geometric mean method to measure the comprehensive importance of each visual token. This method is based on the mathematical properties of the geometric mean and can more fairly measure the importance of visual tokens. The visual tokens to be retained are selected according to the comprehensive importance scores of the selected set. This method enables the selection of visual tokens to consider both the saliency of visual information and the relevance of text semantics, which is equivalent to selecting the top V tokens that are important for the interior of the image and the top V tokens that are important for text semantics.
[0071] S106: According to the cross-modal selected token set and the self-attention selected token set, perform redirection selection on multiple visual tokens of the input image to prune the multiple visual tokens.
[0072] In some embodiments of the present application, the process of performing redirection selection on multiple visual tokens of the input image is as follows: Find the union of the cross-modal selected token set and the self-attention selected token set to obtain the selected set of retained visual tokens, and prune the other un-retained tokens.
[0073] It should be noted that the number of finally selected visual tokens may be less than +V because there is a possibility of overlap.
[0074] Among them, the pruning expression is as follows:
[0075] Among them, represents the The order of cross-modal attention for a token under descending order, represents the order of self-attention for the th token under descending order, represents the top U visual tokens, represents the top V visual tokens, represents the selected set of retained tokens, represents the kth visual token, represents the set of retained tokens in descending order.
[0076] In summary, consider the redirection of visual tokens: According to the comprehensive importance score, redirect the visual attention from the original distribution based only on visual information to the distribution that also considers text semantics. This process is achieved by selecting visual tokens that are more relevant to the text prompt, enabling the large language model to better focus on task-related visual information.
[0077] Consider the selection and sorting of visual tokens: The retained visual tokens will be passed to the LLM together with the text tokens for subsequent processing. In this way, the adaptive cross-modal attention redirection module achieves effective pruning of visual tokens, reducing the computational burden while improving the performance of the model.
[0078] S107: Input the retained visual tokens and the text vector into the large language model for inference.
[0079] In some embodiments of the present application, after completing the pruning of visual tokens, the visual tokens and text tokens enter the large language model (LLM) for inference simultaneously. This process only retains the visual tokens remaining after pruning. During the subsequent inference process, only one new text token is generated each time until the LLM outputs the [EOS] token or reaches the upper limit of the text and visual token capacity of the model, and the generation process stops. Since visual tokens are usually much more numerous than text tokens, and this technology greatly reduces the number of input visual tokens, the time expenditure of the overall inference process is significantly reduced.
[0080] It should be noted that although the embodiments of the present application are introduced and described in sequence for steps S101 to S107 with reference to Figure 1 this does not mean that steps S101 to S107 must be executed in a strict order. The reason why the embodiments of the present application introduce and describe steps S101 to S107 in the order shown in Figure 1 is to facilitate those skilled in the art to understand the technical solutions of the embodiments of the present application. In other words, in the embodiments of the present application, the order between steps S101 to S107 can be appropriately adjusted according to actual needs.
[0081] By Figure 1 the method of, a training-free visual token pruning method is proposed. By mimicking the way the human brain processes multimodal information, redundant visual tokens are cropped before inputting into the large model, thus significantly improving the information processing efficiency of the multimodal large model.
[0082] Among them, AdaV (Adaptive Visual Token Pruning) of this application is an acceleration method without additional training, which can be directly applied to existing vision-language models (vision-text large models) without retraining or fine-tuning the model. This greatly reduces the computational cost and time overhead, and also avoids the potential overfitting risk introduced by training. Specifically, by introducing a visual attention redirection mechanism in the pre-LLM stage, the AdaV selector can efficiently prune visual tokens without any training or fine-tuning of the model.
[0083] In addition, the AdaV method significantly reduces the number of visual tokens, thus greatly improving the inference speed of the large language model. When retaining a small number of visual tokens (such as 10% or less), AdaV can still maintain the performance close to the original model. Specifically, on the LLaVA-NEXT-7B model, the performance of AdaV only drops by about 1.51% when retaining 25% visual tokens, and drops by about 3.96% when retaining 10% visual tokens. When retaining 5% visual tokens, the performance of AdaV is still better than other training-independent methods and even better than the fine-tuning method VisionZip.
[0084] The technical solution is specifically as follows: Introduce cross-modal attention of each visual token for text prompts. The higher the cross-modal attention of a token, the stronger the semantic correlation between the token and the text, which is important for understanding text-image alignment. At the same time, calculate the self-attention of each visual token. The higher the self-attention score of a token, the more important information the token has within the image, which is important for understanding the image content.
[0085] On this basis, by cumulative summation and descending order sorting, the combination problem (the problem of selecting token sets by two kinds of attention) can be transformed into a computable form, which can avoid combinatorial explosion and dimensionality reduction. In addition, when sorting in descending order, the sum of the first i tokens is actually the maximum possible sum under the quantity of i, which can exclude obviously impossible sets.
[0086] By constructing a mask matrix, the token selection range can be ensured to be valid, that is, to ensure that the total number of finally selected tokens calculated does not exceed the token selection quantity threshold, that is, does not exceed the total number of visual tokens of the input image. When the mask returns to zero, illegal solutions are automatically excluded.
[0087] Under the constraint of the above mask matrix, the number of cross-modal selected tokens and the number of self-attention selected tokens are respectively determined by integrating the row index and column index of the importance score. Thus, within the effective selection range, the matrix coordinates are set as the optimal encoding of the resource allocation scheme. The optimization process is to locate the best coordinate point in the two-dimensional resource allocation space. The row-column index method transforms the discrete problem into a differentiable matrix extreme value search. This design transforms the combinatorial optimization into a solvable problem, ensuring optimality more than the traditional greedy algorithm.
[0088] By maximizing the comprehensive importance score and solving the optimal solution of the objective function, the visual token selection is dynamically optimized, so as to screen out the set of cross-modal selected tokens that are relatively important to the text and the set of self-attention selected tokens that are relatively important within the image. And through the union of the two, the visual tokens that are important both to the text semantics and within the image are retained, thus taking into account the information of both aspects.
[0089] In summary, the visual attention is redirected from the original distribution based only on visual information to the distribution that also considers text semantics, enabling the selection of visual tokens to consider both the saliency of visual information and the relevance of text semantics at the same time. The effective pruning of visual tokens is achieved, reducing the computational burden while improving the performance of the model.
[0090] More intuitively, Figure 2 is a schematic diagram of a multi-modal large model inference framework provided by an embodiment of the present application.
[0091] In Figure 2 it shows the process that the adaptive cross-modal attention redirection module redirects the visual attention from the original distribution based only on visual information to the distribution that also considers text semantics, dynamically optimizes the visual token selection, and effectively prunes all visual tokens of the input image based on the comprehensive importance of the selected token set.
[0092] Based on the same idea, some embodiments of the present application also provide devices and non-volatile computer storage media corresponding to the above method.
[0093] Figure 3 is a schematic diagram of the structure of a multi-modal large model inference device provided by an embodiment of the present application, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a multi-modal large model inference method described in any one of the above.
[0094] A non-volatile computer storage medium for multi-modal large model inference provided by some embodiments of the present application stores computer-executable instructions that can execute a multi-modal large model inference method described in any one of the above.
[0095] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the partial description of the method embodiments for the relevant parts.
[0096] The devices and media provided by the embodiments of the present application correspond one-to-one with the methods. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.
[0097] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0098] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0099] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 one or more of the blocks.
[0100] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 one or more of the blocks.
[0101] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0102] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0103] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0104] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.
[0105] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the technical principle of the present application shall fall within the protection scope of the present application.
Claims
1. A multi-modal large model inference method, characterized in that, The method includes: Determining the cross-modal attention of each visual token with respect to the text prompt according to the semantic relevance between the text vector of the text prompt and each visual token corresponding to the input image; Descendingly sorting the self-attention and cross-modal attention of each visual token respectively, and cumulatively summing the sorted self-attention and cross-modal attention respectively to obtain a self-attention cumulative sum vector and a cross-modal attention cumulative sum vector; Constructing a comprehensive importance matrix according to the self-attention cumulative sum vector and the cross-modal attention cumulative sum vector; and constructing a mask matrix according to the token selection quantity threshold of the input image; Performing an element-wise multiplication of the comprehensive importance matrix and the mask matrix to obtain the comprehensive importance scores corresponding to the cross-modal selected token set and the self-attention selected token set; wherein, the number of cross-modal selected tokens and the number of self-attention selected tokens are determined by the row index and column index of the comprehensive importance scores respectively; Optimally solving for the cross-modal selected token set and the self-attention selected token set that satisfy the constraint conditions when the objective function is at its maximum according to the maximization of the comprehensive importance scores and a preset objective function; the objective function is used to represent the comprehensive representation of maximizing the cross-modal semantic relevance and the internal importance of the image; Performing a redirected selection on the multiple visual tokens of the input image according to the cross-modal selected token set and the self-attention selected token set to prune the multiple visual tokens; Inputting the retained visual tokens and the text vector into a large language model for inference.
2. The method according to claim 1, characterized in that The determining the cross-modal attention of each visual token with respect to the text prompt according to the semantic relevance between the text vector and each visual token specifically includes: Calculating the semantic similarity between the text vector and each visual token, and normalizing the semantic similarity to obtain the initial cross-modal attention; Aligning the initial cross-modal attention distribution with the self-attention distribution to obtain the cross-modal attention of each visual token.
3. The method according to claim 2, wherein The calculating the semantic similarity between the text vector and each visual token specifically includes: Calculating the cosine value between the text vector and each visual token; Determining the cosine value as the semantic similarity between the text vector and each visual token.
4. The method according to claim 1, wherein The constructing a comprehensive importance matrix according to the self-attention cumulative sum vector and the cross-modal attention cumulative sum vector specifically includes: Transposing the self-attention cumulative sum vector, and multiplying the transposed self-attention cumulative sum vector by the cross-modal attention to obtain the comprehensive importance matrix.
5. The method according to claim 1, wherein The constructing a mask matrix according to the token selection quantity threshold of the input image specifically includes: Summing the row number and column number of the position of the mask matrix; When the summation result is less than or equal to the token selection quantity threshold, filling the element value at the corresponding position with 1; When the summation result is greater than the token selection quantity threshold, filling the element value at the corresponding position with 0.
6. The method according to claim 1, characterized in that, The optimally solving for the cross-modal selected token set and the self-attention selected token set that satisfy the constraint conditions when the objective function is at its maximum according to the maximization of the comprehensive importance scores and a preset objective function specifically includes: Construct a geometric mean form objective function based on the cross-modal attention sum of the selected cross-modal selection token set and the self-attention sum of the selected self-attention selection token set; Under the constraint that the sum between the size of the cross-modal selection token set and the size of the self-attention selection token set is equal to the token selection number threshold, optimize the cross-modal selection token set and the self-attention selection token set by maximizing the comprehensive importance score to solve the optimal solution of the objective function.
7. The method according to claim 1, wherein The method further includes: Determine the value of the row index as the number of cross-modal selection tokens, and determine the value of the column index as the number of self-attention selection tokens; According to the number of cross-modal selection tokens, sequentially extract the top-ranked visual tokens from the cross-modal attention descending order to obtain the cross-modal selection token set; According to the number of self-attention selection tokens, sequentially extract the top-ranked visual tokens from the self-attention descending order to obtain the self-attention selection token set.
8. The method according to claim 1, wherein Optimize and solve the cross-modal selection token set and the self-attention selection token set that satisfy the constraint conditions when the objective function is at its maximum according to the maximization of the comprehensive importance score and the preset objective function, specifically including: Among them, represents the cross-modal attention of the th unsorted token, represents the self-attention of the th unsorted token, represents the size of the cross-modal selected token set, represents the size of the self-attention selected token set, represents the token selection quantity threshold of the input image, represents the comprehensive importance matrix, M represents the mask matrix, are respectively the number of tokens in the cross-modal selected token set and the number of tokens in the self-attention selected token set when the comprehensive importance score of the element in the th row and th column is the largest, , V = = U, = V.
9. The method according to claim 1, wherein The re-directive selection of multiple visual tokens of the input image according to the cross-modal selection token set and the self-attention selection token set specifically includes: Take the union of the cross-modal selection token set and the self-attention selection token set to obtain the selected retained visual token set, and prune other un-retained tokens.
10. A multimodal large model inference device, characterized in that, Includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a multi-modal large model inference method according to any one of claims 1-9 above.
Citation Information
Patent Citations
Visual token pruning method based on graph information propagation
CN119919768A
Enhanced model explanations using dynamic tokenization for entity matching models
US20240177053A1
Cited By
Multi-modal pruning method based on multi-granularity multi-encoder cooperation mechanism
CN121279386A