A visual token pruning method based on diversity and alignment perception
Patent Information
- Application Number
- CN202611053260.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-09-25
AI Technical Summary
其缺陷在于:评分过程通常与用户文本指令无关,容易保留视觉上显著但与问题无关的背景区域,也可能重复选择相似令牌,导致有限预算未被有效利用
第一,本发明同时解决视觉多样性不足与指令对齐不足的问题。现有自加权方法容易保留与指令无关的显著背景,单纯多样性方法容易选择离群背景,单纯指令对齐方法又容易重复保留相邻局部区域。本发明将基于LogDet函数的多样性效用项和基于令牌级语义相似度的文本指令对齐效用项统一在同一子模目标中,使保留的视觉令牌既能覆盖非冗余视觉区域,又能围绕用户指令提供关键证据,从而缓解多样性与对齐性的固有冲突。
Smart Images

Figure CN122819487A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and multimodal large language model reasoning acceleration technology, and in particular to a visual token pruning method for multimodal large language models, which can be applied to scenarios such as visual language models, image question answering, visual reasoning, multimodal dialogue, and efficient multimodal reasoning on resource-constrained devices. Background Technology
[0002] Existing Multimodal Large Language Models (MLLMs) typically begin by converting image or video frames into a large number of visual tokens using a visual encoder. These visual tokens, along with text instructions, are then input into the large language model decoder. While these models have achieved good results in tasks such as image-text question answering, visual reasoning, OCR understanding, and video question answering, their main computational bottleneck stems from the visual modality. To capture fine-grained information in high-resolution images, the visual encoder often needs to generate hundreds to thousands of visual tokens. For example, common models such as LLaVA-1.5, LLaVA-NeXT, and Qwen2.5-VL can encode an image into 576, 2880, or approximately 2691 visual tokens, respectively. Since the computational complexity of self-attention in Transformer-based large language model decoders increases quadratically with sequence length, a large number of visual tokens significantly increases inference latency, memory usage, and deployment costs, limiting the application of multimodal large language models in edge devices, online services, and resource-constrained scenarios.
[0003] To alleviate the aforementioned problems, existing technologies typically employ visual token pruning or compression methods, removing a portion of visual tokens before or during model inference, retaining only the image regions deemed more important. Compared to compression methods that require retraining the visual projector or modifying the model structure, training-independent visual token pruning methods offer plug-and-play advantages, making them more suitable for deployment in existing multimodal large language models.
[0004] Existing token pruning schemes that do not require training can be broadly classified into four categories: The first category is the self-weighted pruning method. This type of method typically selects visual tokens based on the attention weights within the visual encoder or large language model, CLS token attention, feature norm, or other heuristic importance scores. Its drawbacks are: the scoring process is usually independent of the user's text instructions, it easily preserves visually significant but irrelevant background areas, and it may repeatedly select similar tokens, resulting in inefficient use of the limited budget.
[0005] The second category is diversity-based pruning methods. These methods treat visual tokens as a whole set, tending to retain tokens that differ significantly from each other in the visual feature space to cover a wider spatial or semantic region. Their drawback is that simply pursuing diversity may lead to the selection of outliers or complex backgrounds in the feature space, thereby removing crucial local details needed to answer the user's question, resulting in misaligned instructions.
[0006] The third category is instruction-alignment-based pruning methods. These methods utilize the attention or similarity relationship between text instructions and visual tokens to prioritize the preservation of visual regions that are semantically relevant to the text. Their drawback is that they often employ additive or modular scoring, ignoring information redundancy between similar tokens. If multiple adjacent image patches correspond to the same target object, they may receive similar alignment scores, leading to repeated preservation, high token redundancy, and reduced coverage of the global context.
[0007] The fourth category is hybrid pruning methods. Some existing methods attempt to consider visual saliency, diversity, and text relevance simultaneously, but most rely on empirical fusion rules or linear weight aggregation, lacking explicit modeling of diminishing marginal returns and failing to provide provable guarantees of approximate optimality. Especially at high compression rates, how to balance visual diversity and instruction alignment within a limited token budget still lacks a unified, stable, and training-independent technical solution.
[0008] Therefore, existing visual token pruning techniques for multimodal large language models still suffer from the following problems: First, self-weighted methods lack query-awareness; second, diversity methods easily lose key evidence related to instructions; third, instruction alignment methods tend to repeatedly select redundant local regions; and fourth, hybrid methods are mostly heuristic fusions, making it difficult to achieve a stable balance between diversity and alignment. There is an urgent need for a training-irrelevant visual token pruning method that can unify visual diversity, text instruction alignment, and diminishing marginal returns mechanisms under a single optimization objective. Summary of the Invention
[0009] Therefore, it is necessary to provide a visual token pruning method based on diversity and alignment awareness to address the aforementioned technical problems.
[0010] Firstly, this application provides a visual token pruning method based on diversity and alignment awareness. Applied to a multimodal large language model, the multimodal large language model includes a visual encoder, projector, tokenizer, and Transformer decoder. The method includes: Obtain the visual token feature sequence output by the visual encoder after encoding the input image, and obtain the text token representation output by the word segmenter after processing the text instruction; Construct a similarity matrix between any two visual tokens in the visual token feature sequence. The similarity matrix is used to characterize the degree of similarity between the visual tokens in the semantic space after the projector mapping. Calculate the semantic alignment score between each visual token and the global representation of the text instruction; A diversity utility term is constructed based on the similarity matrix, and an alignment utility term is constructed based on the semantic alignment score. The diversity utility term and the alignment utility term are then weighted and fused into a unified objective function. Given a visual token budget, with the goal of maximizing the unified objective function, visual tokens are iteratively selected from the visual token feature sequence. In each round, candidate visual tokens that maximize the marginal benefit of the unified objective function are added to the selected set until the number of visual tokens in the selected set reaches the visual token budget. The visual tokens and text tokens from the selected set are input together into the Transformer decoder for inference, generating an inference result.
[0011] Optionally, in one embodiment of this application, the diversity utility term is a determinant logarithmic function based on the principal submatrix composed of the selected visual token indices in the similarity matrix, and the value of the determinant logarithmic function increases as the average similarity between visual tokens in the selected set decreases.
[0012] Optionally, in one embodiment of this application, the similarity matrix is constructed as follows: each visual token is mapped to the semantic space of the multimodal large language model via the projector and normalized; the dot product between the normalized visual token representations is calculated; and the dot product value is used as the element at the corresponding position in the similarity matrix.
[0013] Optionally, in one embodiment of this application, the semantic alignment score is the dot product between the normalized representation of each visual token and the global normalized representation of the text instruction, wherein the global representation of the text instruction is the pooling result of the text token representation in the sequence dimension or the global representation vector output by the tokenizer.
[0014] Optionally, in one embodiment of this application, the marginal benefit is obtained by weighted summation of the new diversity contribution and the new alignment contribution, wherein the new diversity contribution is the increase in the diversity utility item after the candidate visual token is added to the selected set, and the new alignment contribution is the semantic alignment score of the candidate visual token.
[0015] Optionally, in one embodiment of this application, the new diversity contribution is measured by the conditional variance of the candidate visual tokens under the selected set conditions. The conditional variance is obtained by recursively updating the residual covariance matrix through Schur complement. After each round of selection, the residual covariance matrix is updated according to the selected visual tokens for the calculation of conditional variance in the next round.
[0016] Optionally, in one embodiment of this application, in the weighted fusion of the unified objective function, the weight coefficients of the diversity utility term and the alignment utility term are preset constants, and the values of the weight coefficients are determined according to the compression requirements of the inference task type or deployment scenario.
[0017] Secondly, this application also provides a visual token pruning system based on diversity and alignment awareness, applied to a multimodal large language model. The multimodal large language model includes a visual encoder, a projector, a tokenizer, and a Transformer decoder. The system includes: The input acquisition module is used to acquire the visual token feature sequence output by the visual encoder after encoding the input image, and to acquire the text token representation output by the word segmenter after processing the text instruction; A similarity construction module is used to construct a similarity matrix between any two visual tokens in the visual token feature sequence. The similarity matrix is used to characterize the degree of similarity between the visual tokens in the semantic space after the projector mapping. An alignment calculation module is used to calculate the semantic alignment score between each visual token and the global representation of the text instruction; The objective construction module is used to construct a diversity utility term based on the similarity matrix, construct an alignment utility term based on the semantic alignment score, and weight and fuse the diversity utility term and the alignment utility term into a unified objective function. The token selection module is used to iteratively select visual tokens from the visual token feature sequence under a given visual token budget, with the goal of maximizing the unified objective function. In each round, the candidate visual token that maximizes the marginal benefit of the unified objective function is added to the selected set until the number of visual tokens in the selected set reaches the visual token budget. The inference output module is used to input the visual tokens and text tokens from the selected set into the Transformer decoder for inference and generate inference results.
[0018] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the steps of the methods described in the various embodiments above.
[0019] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the methods described in the various embodiments above.
[0020] The aforementioned visual token pruning method based on diversity and alignment awareness has the following advantages compared to existing technologies: First, this invention simultaneously addresses the problems of insufficient visual diversity and insufficient instruction alignment. Existing self-weighted methods tend to retain salient backgrounds unrelated to the instruction, simple diversity methods tend to select outlier backgrounds, and simple instruction alignment methods tend to repeatedly retain adjacent local regions. This invention unifies the diversity utility term based on the LogDet function and the text instruction alignment utility term based on token-level semantic similarity in the same sub-module target, enabling the retained visual tokens to both cover non-redundant visual regions and provide key evidence around user instructions, thereby alleviating the inherent conflict between diversity and alignment.
[0021] Second, this invention can significantly reduce the inference overhead of multimodal large language models. By reducing the number of visual tokens before the large language model decoder, this invention can directly reduce the length of subsequent self-attention sequences, thereby reducing inference latency and memory usage. Since this method does not introduce additional trainable branches and does not require retraining the visual encoder or large language model, it is suitable for rapid deployment on existing models.
[0022] Third, this invention is training-independent, architecture-independent, and plug-and-play. It utilizes only the existing visual tokens, text representations, and projectors in the original model to calculate the similarity matrix and relevance scores without updating model parameters, thus adapting to different architectures such as LLaVA-1.5, LLaVA-NeXT, Qwen2.5-VL, and Video-LLaVA. For engineering deployment, this method can be integrated as an independent inference acceleration module into existing multimodal large language model services.
[0023] Fourth, this invention maintains a good accuracy-efficiency trade-off across various multimodal models and tasks. The method was validated on six architectures: LLaVA-1.5-7B / 13B, LLaVA-NeXT-7B / 13B, Qwen2.5-VL-7B, and Video-LLaVA-7B. It outperforms several existing visual token pruning methods on various benchmarks, including visual question answering, graph-text reasoning, target illusion detection, and video question answering. Particularly on LLaVA-NeXT-7B, even with only about 5.60% of the visual tokens retained, the method still maintains approximately 95.43% of the original model's accuracy, demonstrating high compression efficiency and inference performance preservation.
[0024] Fifth, this invention is better able to preserve key visual evidence in high compression scenarios. With extremely low visual token budgets, traditional methods often lose key regions due to excessive pursuit of local alignment or blind dispersion. The sub-model objective of this invention can suppress redundant tokens through a diminishing marginal returns mechanism and strengthen key regions through instruction alignment scores, thus preserving a relatively complete reasoning basis for the model even at high compression rates. Attached Figure Description
[0025] Figure 1 This is a flowchart illustrating a visual token pruning method based on diversity and alignment awareness in one embodiment. Figure 2 This is a structural block diagram of a visual token pruning system based on diversity and alignment awareness in one embodiment; Figure 3 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0027] In the embodiments of this application, such as Figure 1 As shown, a visual token pruning method based on diversity and alignment awareness is provided. It can be deployed as a training-independent plug-and-play module in existing multimodal large language models. It is suitable for model architectures such as LLaVA series, Qwen2.5-VL and Video-LLaVA, where the decoder is driven by both visual tokens and text tokens.
[0028] In this embodiment, the multimodal large language model includes a visual encoder, a projector, a tokenizer, and a Transformer decoder. Given an input image and text instructions, the input image is first encoded using the visual encoder in the multimodal large language model to obtain the original visual token feature sequence. Let the total number of visual tokens be... , No. Each visual token feature is denoted as The set of indices corresponding to all visual tokens is denoted as . .
[0029] Simultaneously, the text instructions are processed using a tokenizer to obtain a sequence of text token representations, denoted as... Under the condition of a complete input without visual token pruning, the visual token sequence and text token are input together into a multimodal large language model, and the model output can be represented as follows: ,in The parameter is A multimodal large language model.
[0030] In this embodiment of the application, the goal of the method is to obtain a complete set of visual tokens without updating the model parameters. Select a smaller subset of visual tokens This aims to make the pruned model output as close as possible to the model output under the full visual token input. Let the number of tokens retained be... ( Then, the visual token pruning problem can be represented as:
[0031] in, This represents a measure of the output difference between the complete input and the pruned input. Since directly solving this combinatorial optimization problem has high computational complexity, this invention further transforms it into a visual token subset utility maximization problem.
[0032] In one embodiment of this application, to avoid the pruning process resulting in visual tokens being concentrated in the same local region, the present invention models the diversity of the candidate visual token set. Let... This represents a projector in a multimodal large language model, used to map visual token features to the semantic space of the large language model. The representation of a visual token after projection and normalization is as follows: Based on the projected visual token representations, a similarity matrix is constructed between the visual tokens. , of which Line 1 The elements of the column are:
[0033] For any subset of candidate visual tokens Take the similarity matrix from Principal submatrix obtained by indexing And define the visual diversity goal as:
[0034] in, It is the numerical stability constant. This is the identity matrix. This term measures the degree of information complementarity among the tokens in the selected visual token set. When a candidate token is highly similar to the selected tokens, its contribution to the addition is small; when a candidate token can supplement new visual regions or semantic information, its contribution to the addition is large. The value of the logarithm of the determinant increases as the average similarity among the visual tokens in the selected set decreases, that is, the less similar the selected tokens are to each other, the higher the value of the diversity utility term.
[0035] In one embodiment of this application, the similarity matrix is constructed as follows: each visual token is mapped to the semantic space of the multimodal large language model via the projector and normalized; the dot product between the normalized visual token representations is calculated; and the dot product value is used as the element at the corresponding position in the similarity matrix. Specifically, for any... The and the first Each visual token, and its similarity matrix elements This is the dot product as represented by the normalized form.
[0036] In one embodiment of this application, to avoid retaining background tokens irrelevant to the current task solely based on visual diversity, the present invention further calculates the semantic alignment degree between each visual token and the text instruction. In one embodiment of this application, the semantic alignment score is the dot product between the normalized representation of each visual token and the globally normalized representation of the text instruction. Specifically, let the global representation of the text instruction be... The normalized visual token representation and global text representation are obtained from the pooling result of the text token representation along the sequence dimension or the global representation vector output by the tokenizer. and Then the first The correlation score between a visual token and a text instruction is defined as follows:
[0037] To stabilize the numerical scale among different input samples, the correlation scores are subjected to min-max normalization:
[0038] in, and Let represent the minimum and maximum relevance scores of all visual tokens in the current image, respectively. For any subset of candidate tokens... The text instruction alignment target is defined as:
[0039] This item is used to guide the pruning process to prioritize retaining visual tokens associated with text instructions.
[0040] In one embodiment of this application, the present invention performs a weighted fusion of the visual diversity objective and the text instruction alignment objective to construct a unified visual token pruning objective function:
[0041] Expanding the above two items, we get:
[0042] in, The first term is a balancing coefficient used to adjust the weight between visual diversity and text instruction alignment. The first term encourages the selected token to cover more non-redundant visual information, while the second term encourages the selected token to remain semantically relevant to the current text instruction. In one embodiment of this application, the weight coefficients of the diversity utility term and the alignment utility term are preset constants, and the values of these weight coefficients are determined based on the compression requirements of the inference task type or deployment scenario.
[0043] In one embodiment of this application, given a token budget Under these conditions, this invention employs a greedy selection strategy to solve for the visual token subset. Let the first... The set of tokens already selected in the round is Candidate token is Then The marginal benefit of joining the current set is:
[0044] make To normalize the similarity matrix, based on the determinant chain decomposition, the diversity score can be expressed as:
[0045] in, This indicates a selection order for the chosen token indexes. Indicates the first The set of tokens already retained before the round of selection. Indicates candidate token The conditional variance, given the current set of selected data, is used to measure the additional visual information it introduces. Therefore, the... In the greedy choice round, Add to the currently selected set The marginal benefit can be written as:
[0046] The first term represents the contribution of candidate tokens to the diversity of the selected token set, and the second term represents the alignment contribution between candidate tokens and text instructions. This serves as a balancing coefficient. Therefore, in each round of selection, this invention prioritizes retaining visual tokens that simultaneously possess high levels of newly added visual information and strong textual relevance. In one embodiment of this application, the marginal benefit is obtained by a weighted sum of the newly added diversity contribution and the newly added alignment contribution. The newly added diversity contribution is the increase in the diversity utility item after a candidate visual token is added to the selected set, and the newly added alignment contribution is the semantic alignment score of the candidate visual token.
[0047] In one embodiment of this application, to avoid recalculating the conditional variance from scratch or explicitly performing matrix inversion in each round, the present invention employs Schur complement recursively to update the residual covariance matrix. After the first update, the [number]th The greedy selection of the visual token that maximizes marginal returns in the wheel can be simplified to:
[0048] in Indicates the preceding After the first round of updates, the regularized similarity matrix is updated. The diagonal elements represent the remaining variance of the candidate tokens under the current conditions.
[0049] Then update the selected token set. And based on the selected token Update the residual covariance matrix.
[0050] Repeat the above process until the condition is met. Ultimately, a set of visual token indices is obtained. In one embodiment of this application, the contribution of new diversity is measured by the conditional variance of candidate visual tokens under the selected set conditions. The conditional variance is obtained by recursively updating the residual covariance matrix through Schur complement. After each round of selection, the residual covariance matrix is updated according to the selected visual token for the calculation of conditional variance in the next round.
[0051] In this embodiment, after completing the visual token selection, only the retained visual token and the text token are input together into the multimodal large language model decoder to obtain the pruned model output. Because the number of visual tokens in the input language model decoder is large, the number of tokens is determined by... Reduce to ,and Therefore, it can significantly reduce the sequence length, computational overhead, and memory usage during the inference phase.
[0052] In this embodiment, the method can be integrated into various multimodal large language model architectures. For the LLaVA-1.5-7B model, the image generates 576 visual tokens via a visual encoder. For the LLaVA-NeXT-7B model, the high-resolution image generates 2880 visual tokens via a visual encoder. For the Qwen2.5-VL-7B model, the visual encoder generates approximately 2691 visual tokens. Across all the above architectures, this invention can achieve joint evaluation of visual token diversity and alignment with a greedy selection process, with an overhead not exceeding a certain proportion of the original model's inference time, while maintaining a high degree of consistency in the original model's output distribution after compression.
[0053] Taking the application of the LLaVA-NeXT-7B model in visual question answering and image-text reasoning tasks as an example, given an input image and user text instructions, the system first obtains the original visual token sequence through a visual encoder and the text token sequence through a token segmenter. Then, following the method described in this invention, a visual token similarity matrix is constructed, and the alignment score between each visual token and the global representation of the text instruction is calculated. A greedy selection is performed within a given token budget (e.g., 5.60% of the original number of tokens). After the selection is completed, only the retained visual tokens and text tokens are input together into the large language model decoder. After self-attention calculation and autoregression generation, the answer is output.
[0054] In experimental verification, this invention was evaluated on six architectures: LLaVA-1.5-7B / 13B, LLaVA-NeXT-7B / 13B, Qwen2.5-VL-7B, and Video-LLaVA-7B, covering multiple benchmark tasks including visual question answering (VQAv2, GQA), image-text reasoning (MMBench, MM-Vet), target illusion detection (POPE), and video question answering (MSVD-QA, MSRVTT-QA). Compared with existing visual token pruning methods that do not require training, this invention achieves better or comparable accuracy across all tasks. Especially under high compression conditions, i.e., when the number of retained tokens is less than 10% of the original number, this invention can still maintain more than 95% of the accuracy of the original model, significantly outperforming methods based solely on attention ranking or diversity.
[0055] In summary, this invention provides a training-independent visual token pruning method that unifies visual diversity and text instruction alignment within a submodal optimization framework. By suppressing redundant selection using a diversity utility term based on the LogDet function, enhancing instruction relevance through token-level semantic alignment scores, and achieving efficient greedy selection through Schur complement recursive updates, this invention significantly reduces inference latency and memory consumption while maintaining the inference accuracy of multimodal large language models, making it suitable for various resource-constrained deployment scenarios.
[0056] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0057] Based on the same inventive concept, this application also provides a visual token pruning system based on diversity and alignment awareness for implementing the aforementioned visual token pruning method based on diversity and alignment awareness. The solution provided by this system is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of a visual token pruning system based on diversity and alignment awareness provided below can be found in the limitations of the visual token pruning method based on diversity and alignment awareness described above, and will not be repeated here.
[0058] In one embodiment, such as Figure 2 As shown, a visual token pruning system 200 based on diversity and alignment awareness is provided, including: an input acquisition module 201, a similarity construction module 202, an alignment calculation module 203, a target construction module 204, a token selection module 205, and an inference output module 206, wherein: The input acquisition module is used to acquire the visual token feature sequence output by the visual encoder after encoding the input image, and to acquire the text token representation output by the word segmenter after processing the text instruction; A similarity construction module is used to construct a similarity matrix between any two visual tokens in the visual token feature sequence. The similarity matrix is used to characterize the degree of similarity between the visual tokens in the semantic space after the projector mapping. An alignment calculation module is used to calculate the semantic alignment score between each visual token and the global representation of the text instruction; The objective construction module is used to construct a diversity utility term based on the similarity matrix, construct an alignment utility term based on the semantic alignment score, and weight and fuse the diversity utility term and the alignment utility term into a unified objective function. The token selection module is used to iteratively select visual tokens from the visual token feature sequence under a given visual token budget, with the goal of maximizing the unified objective function. In each round, the candidate visual token that maximizes the marginal benefit of the unified objective function is added to the selected set until the number of visual tokens in the selected set reaches the visual token budget. The inference output module is used to input the visual tokens and text tokens from the selected set into the Transformer decoder for inference and generate inference results.
[0059] In one embodiment of this application, the diversity utility term is a logarithmic determinant function based on the principal submatrix composed of the selected visual token indices in the similarity matrix, and the value of the logarithmic determinant function increases as the average similarity between visual tokens in the selected set decreases.
[0060] In one embodiment of this application, the similarity matrix is constructed as follows: each visual token is mapped to the semantic space of the multimodal large language model via the projector and normalized; the dot product between the normalized visual token representations is calculated; and the dot product value is used as the element at the corresponding position in the similarity matrix.
[0061] In one embodiment of this application, the semantic alignment score is the dot product between the normalized representation of each visual token and the global normalized representation of the text instruction, wherein the global representation of the text instruction is the pooling result of the text token representation in the sequence dimension or the global representation vector output by the tokenizer.
[0062] In one embodiment of this application, the marginal benefit is obtained by weighted summation of the new diversity contribution and the new alignment contribution, wherein the new diversity contribution is the increase in the diversity utility item after the candidate visual token is added to the selected set, and the new alignment contribution is the semantic alignment score of the candidate visual token.
[0063] In one embodiment of this application, the new diversity contribution is measured by the conditional variance of the candidate visual tokens under the selected set conditions. The conditional variance is obtained by recursively updating the residual covariance matrix through Schur complement. After each round of selection, the residual covariance matrix is updated according to the selected visual tokens for the calculation of conditional variance in the next round.
[0064] In one embodiment of this application, in the weighted fusion of the unified objective function, the weight coefficients of the diversity utility term and the alignment utility term are preset constants, and the values of the weight coefficients are determined according to the compression requirements of the inference task type or deployment scenario.
[0065] The modules in the aforementioned visual token pruning system based on diversity and alignment awareness can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0066] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a visual token pruning method based on diversity and alignment awareness. The display screen can be an LCD screen or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.
[0067] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0068] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0069] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0070] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0071] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0072] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0073] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0074] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A visual token pruning method based on diversity and alignment awareness, characterized in that, The method, applied to a multimodal large language model, which includes a visual encoder, projector, word segmenter, and Transformer decoder, comprises: Obtain the visual token feature sequence output by the visual encoder after encoding the input image, and obtain the text token representation output by the word segmenter after processing the text instruction; Construct a similarity matrix between any two visual tokens in the visual token feature sequence. The similarity matrix is used to characterize the degree of similarity between the visual tokens in the semantic space after the projector mapping. Calculate the semantic alignment score between each visual token and the global representation of the text instruction; A diversity utility term is constructed based on the similarity matrix, and an alignment utility term is constructed based on the semantic alignment score. The diversity utility term and the alignment utility term are then weighted and fused into a unified objective function. Given a visual token budget, with the goal of maximizing the unified objective function, visual tokens are iteratively selected from the visual token feature sequence. In each round, candidate visual tokens that maximize the marginal benefit of the unified objective function are added to the selected set until the number of visual tokens in the selected set reaches the visual token budget. The visual tokens and text tokens from the selected set are input together into the Transformer decoder for inference, generating an inference result.
2. The visual token pruning method based on diversity and alignment perception according to claim 1, characterized in that, The diversity utility term is a logarithmic determinant function based on the principal submatrix formed by the selected visual token indices in the similarity matrix. The value of the logarithmic determinant function increases as the average similarity between visual tokens in the selected set decreases.
3. The visual token pruning method based on diversity and alignment perception according to claim 1, characterized in that, The similarity matrix is constructed as follows: each visual token is mapped to the semantic space of the multimodal large language model through the projector and normalized; the dot product between the normalized visual token representations is calculated; and the dot product value is used as the element at the corresponding position in the similarity matrix.
4. The visual token pruning method based on diversity and alignment perception according to claim 1, characterized in that, The semantic alignment score is the dot product between the normalized representation of each visual token and the global normalized representation of the text instruction, wherein the global representation of the text instruction is either the pooling result of the text token representation in the sequence dimension or the global representation vector output by the tokenizer.
5. The visual token pruning method based on diversity and alignment awareness according to claim 1, characterized in that, The marginal benefit is obtained by weighted summation of the new diversity contribution and the new alignment contribution. The new diversity contribution is the increase in the diversity utility item after the candidate visual token is added to the selected set, and the new alignment contribution is the semantic alignment score of the candidate visual token.
6. The visual token pruning method based on diversity and alignment awareness according to claim 5, characterized in that, The contribution of new diversity is measured by the conditional variance of the candidate visual tokens under the selected set conditions. The conditional variance is obtained by recursively updating the residual covariance matrix through Schur complement. After each round of selection, the residual covariance matrix is updated according to the selected visual tokens for the calculation of conditional variance in the next round.
7. A visual token pruning method based on diversity and alignment perception according to claim 1, characterized in that, In the weighted fusion of the unified objective function, the weight coefficients of the diversity utility term and the alignment utility term are preset constants, and the values of the weight coefficients are determined according to the compression requirements of the inference task type or deployment scenario.
8. A visual token pruning system based on diversity and alignment awareness, characterized in that, The system is applied to a multimodal large language model, which includes a visual encoder, a projector, a word segmenter, and a Transformer decoder. The system comprises: The input acquisition module is used to acquire the visual token feature sequence output by the visual encoder after encoding the input image, and to acquire the text token representation output by the word segmenter after processing the text instruction; A similarity construction module is used to construct a similarity matrix between any two visual tokens in the visual token feature sequence. The similarity matrix is used to characterize the degree of similarity between the visual tokens in the semantic space after the projector mapping. An alignment calculation module is used to calculate the semantic alignment score between each visual token and the global representation of the text instruction; The objective construction module is used to construct a diversity utility term based on the similarity matrix, construct an alignment utility term based on the semantic alignment score, and weight and fuse the diversity utility term and the alignment utility term into a unified objective function. The token selection module is used to iteratively select visual tokens from the visual token feature sequence under a given visual token budget, with the goal of maximizing the unified objective function. In each round, the candidate visual token that maximizes the marginal benefit of the unified objective function is added to the selected set until the number of visual tokens in the selected set reaches the visual token budget. The inference output module is used to input the visual tokens and text tokens from the selected set into the Transformer decoder for inference and generate inference results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.