Visual token compression method and device based on spatiotemporal forest multi-modal large language model

By constructing a spatiotemporal forest structure using the spatiotemporal forest algorithm, the problem of cross-frame redundancy in video tasks for multimodal large language models is solved. This achieves performance preservation and computational efficiency improvement under high compression ratio and is suitable for visual token compression of multimodal large language models.

CN122269034APending Publication Date: 2026-06-23XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610346202.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-20
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing multimodal large language models have cross-frame redundancy that is not modeled in video tasks, resulting in a sharp drop in performance under high compression rates and making them difficult to apply to existing inference systems.

Method used

The spatiotemporal forest algorithm is used to construct a spatiotemporal forest structure. Directed edges between nodes are constructed by combining semantic similarity, spatial distance and temporal order constraints. Global optimization and compression of visual token embedding are performed to form a spatiotemporal tree and spatiotemporal forest. Nodes are pruned and merged to reach the preset number of tokens.

Benefits of technology

Maintaining 94% model performance with 90% visual token compression, reducing attention computation and GPU memory usage, suitable for any multimodal large language model, requiring no training or parameter tuning, and applicable to long videos and resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122269034A_ABST
    Figure CN122269034A_ABST
Patent Text Reader

Abstract

The application discloses a kind of visual Token compression method and device of multi-modal large language model based on space-time forest, comprising: each frame video frame is input into multi-modal large language model, after the visual encoder in multi-modal large language model, the visual Token embedding sequence corresponding to each frame video frame is obtained, from the visual Token embedding sequence corresponding to each frame video frame, several key visual Token embeddings are selected and node feature is constructed;Node feature is input into space-time forest algorithm, each key visual Token embedding is regarded as node, the directed edge between two nodes is constructed based on the joint constraint of semantic similarity, spatial distance and time sequence, node is constructed into space-time tree by directed edge and is organized into space-time forest structure, and space-time forest structure is processed under preset constraint, obtain final node set, the key visual Token embedding corresponding to the node in final node set is regarded as compressed visual Token embedding sequence.The application can reduce the amount of calculation of model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, specifically to a visual token compression method and apparatus based on a multimodal large language model of spatiotemporal forest. Background Technology

[0002] With the development of multimodal large language models, they have achieved significant improvements in image, video, and long video analysis and reasoning. Typical multimodal large language models usually use visual encoders to extract multi-frame image features from the input video, with each frame embedding hundreds to thousands of visual tokens. As the number of input video frames increases, the number of tokens in the video task grows linearly or even approximately quadratically. Existing research attempts to compress the data through token pruning or token merging, but mainstream methods are based on the estimation of the importance of features in a single frame of the image, ignoring the retention of globally invalid tokens caused by cross-frame redundancy, resulting in a sharp decline in performance under high compression ratio scenarios.

[0003] Existing token compression technologies have the following main shortcomings:

[0004] 1. It only processes the importance of tokens within a frame and does not model cross-frame redundancy, thus failing to handle the continuous semantic structure of the video.

[0005] 2. Local alignment of adjacent frames or regions alone cannot avoid the accumulation of global semantic redundancy;

[0006] 3. When the compression ratio is high, the model's inference performance decreases significantly;

[0007] 4. Some methods require retraining or fine-tuning and are difficult to apply to existing reasoning systems. Summary of the Invention

[0008] The purpose of this application is to propose a visual token compression method and apparatus based on a multimodal large language model of spatiotemporal forest to address the aforementioned technical problems.

[0009] In a first aspect, the present invention provides a visual token compression method based on a multimodal large language model of spatiotemporal forest, comprising the following steps:

[0010] Each video frame in the input video is input into the multimodal large language model. After passing through the visual encoder in the multimodal large language model, the visual token embedding sequence corresponding to each video frame is obtained. Several key visual tokens are selected from the visual token embedding sequence corresponding to each video frame and node features are constructed.

[0011] The node features are input into the spatiotemporal forest algorithm. In the spatiotemporal forest algorithm, each key visual token embedding is used as a node. Directed edges between two nodes are constructed based on joint constraints of semantic similarity, spatial distance and temporal order. The nodes are constructed into a spatiotemporal tree and assembled into a spatiotemporal forest structure through the directed edges. The spatiotemporal forest structure is then processed under preset constraints to obtain the final node set. The key visual token embeddings corresponding to the nodes in the final node set are used as a compressed visual token embedding sequence and participate in the calculation process of other modules in the multimodal large language model.

[0012] Preferably, the key visual token embedding is obtained by selecting the visual token embeddings in the visual token embedding sequence using a pruning or merging method, and the node feature is a vector formed by concatenating all the key visual token embeddings.

[0013] As a preferred approach, the processing of the spatiotemporal forest structure includes merging, sorting, and pruning spatiotemporal trees.

[0014] Preferably, directed edges are constructed between two nodes based on joint constraints of semantic similarity, spatial distance, and temporal order. These directed edges are then used to build a spatiotemporal tree and a spatiotemporal forest structure. This spatiotemporal forest structure is then processed under a preset number of tokens to obtain the final node set, specifically including:

[0015] For each node, record its coordinates and timestamp in the corresponding video frame;

[0016] Calculate the semantic similarity between the key visual token embeddings corresponding to any two nodes, and construct a similarity matrix;

[0017] Calculate the distance between the coordinates of any two nodes and construct the spatial distance matrix;

[0018] Construct an adjacency matrix based on the similarity matrix, spatial distance matrix, and timestamp;

[0019] The root node is determined based on the out-degree and in-degree of the elements in the adjacency matrix, and the root set is constructed.

[0020] The sorting matrix is ​​constructed based on the adjacency matrix, similarity matrix, and spatial distance matrix, as shown in the following formula:

[0021] ;

[0022] in, Represents the Hadema product. Indicates the weighting coefficient. Represents the similarity matrix. Represents the spatial distance matrix. Represents the adjacency matrix. Represents the sorting matrix;

[0023] For the k-th node that is not in the root set, select the root node of the column corresponding to the largest value among all elements in the k-th row of the sorting matrix in the root set as the root node to which the k-th node belongs. Count the nodes to which each root node belongs and obtain the corresponding node set. According to the timestamp, connect the current node in the node set corresponding to each root node with the nodes at previous times through directed edges to construct a spatiotemporal tree with different levels. Several spatiotemporal trees form the initial spatiotemporal forest structure.

[0024] The preset constraints include a preset number of tokens. In response to the fact that the number of root nodes in the root set is much greater than the preset number of tokens, the two corresponding spatiotemporal trees in the initial spatiotemporal forest structure are merged according to the semantic similarity between the two root nodes to obtain the merged spatiotemporal forest structure. The tree depth of each spatiotemporal tree in the merged spatiotemporal forest structure is then recalculated.

[0025] Sort all the spatiotemporal trees in the merged spatiotemporal forest structure according to their tree depth from largest to smallest. Then, prune each spatiotemporal tree in the merged spatiotemporal forest structure according to the order of the sorting to obtain the pruned spatiotemporal forest structure. Continue until the number of all nodes in the pruned spatiotemporal forest structure reaches the preset number of tokens. Finally, take the set of all nodes in the pruned spatiotemporal forest structure as the final node set.

[0026] Preferably, an adjacency matrix is ​​constructed based on the similarity matrix, spatial distance matrix, and timestamp, specifically including:

[0027] Traverse any two nodes, denoted as the i-th node and the j-th node. The semantic similarity between the i-th node and the j-th node is the element in the i-th row and j-th column of the similarity matrix. The distance between the i-th node and the j-th node is the element in the i-th row and j-th column of the spatial distance matrix. ,like Greater than or equal to the similarity threshold If the timestamp of the i-th node is less than or equal to the distance threshold, and the timestamp of the i-th node is less than the timestamp of the j-th node, then the element in the i-th row and j-th column of the adjacency matrix is... If it is 1, otherwise the element in the i-th row and j-th column of the adjacency matrix It is 0.

[0028] As a preferred option, the specific process of pruning is as follows:

[0029] First, remove the outermost node of each spatiotemporal tree in the merged spatiotemporal forest structure, and count the number of nodes remaining after removal;

[0030] If the number of remaining nodes after removal does not reach the preset number of Tokens, then remove the second outermost node from each spatiotemporal tree in the merged spatiotemporal forest structure. Repeat this step to remove nodes layer by layer until the number of remaining nodes after removal reaches the preset number of Tokens. At this point, the remaining nodes after removal constitute the pruned spatiotemporal forest structure.

[0031] Secondly, the present invention provides a visual token compression device based on a spatiotemporal forest multimodal large language model, comprising:

[0032] The node feature construction module is configured to input each video frame in the input video into the multimodal large language model, and obtain the visual token embedding sequence corresponding to each video frame through the visual encoder in the multimodal large language model. Several key visual tokens are selected from the visual token embedding sequence corresponding to each video frame and node features are constructed.

[0033] The compression module is configured to input node features into the spatiotemporal forest algorithm. In the spatiotemporal forest algorithm, each key visual token is embedded as a node. Directed edges between two nodes are constructed based on joint constraints of semantic similarity, spatial distance, and temporal order. The nodes are constructed into a spatiotemporal tree and assembled into a spatiotemporal forest structure through the directed edges. The spatiotemporal forest structure is then processed under preset constraints to obtain the final node set. The key visual token embeddings corresponding to the nodes in the final node set are used as the compressed visual token embedding sequence and participate in the calculation process of other modules in the multimodal large language model.

[0034] Thirdly, the present invention provides an electronic device including one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.

[0035] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the implementations of the first aspect.

[0036] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in any of the implementations in the first aspect.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] (1) The visual token compression method based on the spatiotemporal forest multimodal large language model proposed in this invention compresses the key visual token embeddings extracted and selected by the visual encoder through the spatiotemporal forest algorithm. It uses semantic similarity, spatial distance and temporal order to jointly control the connection of directed edges between nodes and constructs a spatiotemporal tree. It further constructs a spatiotemporal forest structure and processes the spatiotemporal forest structure. It captures the continuous semantics of the video through the spatiotemporal tree, avoids information breakage caused by single frame pruning, and effectively improves the cross-frame information utilization rate. It can still maintain more than 94% of the model performance under the condition of pruning 90% of the visual token embeddings, and has a high compression ratio.

[0039] (2) The visual token compression method based on spatiotemporal forest proposed in this invention generates a large number of visual token embeddings. After compression, the final node set is obtained and mapped to a small number of corresponding key visual token embeddings. This can reduce the attention computation and memory usage of the multimodal large language model and improve the inference throughput.

[0040] (3) The visual token compression method based on spatiotemporal forest multimodal large language model proposed in this invention is applicable to any multimodal large language model with video encoder, and does not require training, parameter tuning and specific framework dependency. It has strong portability and versatility, and is also applicable to long videos and resource-constrained devices. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart illustrating the visual token compression method based on a spatiotemporal forest multimodal large language model, which is an embodiment of this application.

[0043] Figure 2 A flowchart illustrating the visual token compression method based on a spatiotemporal forest multimodal large language model, as an embodiment of this application;

[0044] Figure 3 This is a schematic diagram of a visual token compression device based on a spatiotemporal forest multimodal large language model, which is an embodiment of this application.

[0045] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0047] Figure 1 This application illustrates an embodiment of a visual token compression method based on a spatiotemporal forest multimodal large language model, comprising the following steps:

[0048] S1. Input each video frame in the input video into the multimodal large language model. After passing through the visual encoder in the multimodal large language model, obtain the visual token embedding sequence corresponding to each video frame. Select several key visual tokens from the visual token embedding sequence corresponding to each video frame and construct node features.

[0049] In a specific embodiment, the key visual token embedding is obtained by selecting the visual token embeddings in the visual token embedding sequence using a pruning or merging method, and the node feature is a vector formed by concatenating all the key visual token embeddings.

[0050] For details, please refer to Figure 2 This application proposes a visual token compression method based on a spatiotemporal forest-based multimodal large language model. This method uses a spatiotemporal forest algorithm to globally optimize and compress the visual token embeddings generated in the multimodal large language model, obtaining compressed key visual token embeddings to participate in subsequent calculations within the multimodal large language model. Specifically, the key visual token embeddings selected across frames are used as nodes. Directed edges are constructed according to the principles of semantic similarity, spatial proximity, and temporal order. These directed edges and nodes form a spatiotemporal tree, ultimately forming a spatiotemporal forest structure composed of multiple spatiotemporal trees. After construction, the importance of the key visual token embeddings is evaluated and selected based on the role and depth of the nodes in the spatiotemporal tree. Under preset constraints, nodes in the spatiotemporal forest structure are removed, outputting the final node set, which retains the compressed key visual token embeddings. This achieves global optimization and compression of video visual tokens, replacing the original dense visual token embeddings in subsequent inference processes, and shifting the video processing from frame-by-frame compression to cross-frame global modeling.

[0051] First, each video frame from the input video is fed into a multimodal large language model. The visual encoder within the multimodal large language model encodes each video frame, resulting in the final video frame. Embed visual tokens, first embed each video frame... Several key visual token embeddings are selected as nodes from the visual token embeddings. The specific selection method can employ existing pruning or merging methods, which will not be elaborated upon here. The selected key visual token embeddings are then concatenated to form node features. As input to the subsequent spatiotemporal forest algorithm, where, Indicates the number of video frames. This indicates the number of nodes in each video frame. This indicates the dimension in which the key visual token is embedded. It represents the set of real numbers.

[0052] S2. The node features are input into the spatiotemporal forest algorithm. In the spatiotemporal forest algorithm, each key visual token embedding is used as a node. Directed edges between two nodes are constructed based on joint constraints of semantic similarity, spatial distance and temporal order. The nodes are constructed into a spatiotemporal tree and assembled into a spatiotemporal forest structure through the directed edges. The spatiotemporal forest structure is processed under preset constraints to obtain the final node set. The key visual token embeddings corresponding to the nodes in the final node set are used as a compressed visual token embedding sequence and participate in the calculation process of other modules in the multimodal large language model.

[0053] In a specific embodiment, the processing of the spatiotemporal forest structure includes merging, sorting, and pruning spatiotemporal trees.

[0054] In a specific embodiment, a directed edge is constructed between two nodes based on joint constraints of semantic similarity, spatial distance, and temporal order. The nodes are then used to construct a spatiotemporal tree and form a spatiotemporal forest structure. This spatiotemporal forest structure is then processed under a preset number of tokens to obtain the final node set, specifically including:

[0055] For each node, record its coordinates and timestamp in the corresponding video frame;

[0056] Calculate the semantic similarity between the key visual token embeddings corresponding to any two nodes, and construct a similarity matrix;

[0057] Calculate the distance between the coordinates of any two nodes and construct the spatial distance matrix;

[0058] Construct an adjacency matrix based on the similarity matrix, spatial distance matrix, and timestamp;

[0059] The root node is determined based on the out-degree and in-degree of the elements in the adjacency matrix, and the root set is constructed.

[0060] The sorting matrix is ​​constructed based on the adjacency matrix, similarity matrix, and spatial distance matrix, as shown in the following formula:

[0061] ;

[0062] in, Represents the Hadema product. Indicates the weighting coefficient. Represents the similarity matrix. Represents the spatial distance matrix. Represents the adjacency matrix. Represents the sorting matrix;

[0063] For the k-th node that is not in the root set, select the root node of the column corresponding to the largest value among all elements in the k-th row of the sorting matrix in the root set as the root node to which the k-th node belongs. Count the nodes to which each root node belongs and obtain the corresponding node set. According to the timestamp, connect the current node in the node set corresponding to each root node with the nodes at previous times through directed edges to construct a spatiotemporal tree with different levels. Several spatiotemporal trees form the initial spatiotemporal forest structure.

[0064] The preset constraints include a preset number of tokens. In response to the fact that the number of root nodes in the root set is much greater than the preset number of tokens, the two corresponding spatiotemporal trees in the initial spatiotemporal forest structure are merged according to the semantic similarity between the two root nodes to obtain the merged spatiotemporal forest structure. The tree depth of each spatiotemporal tree in the merged spatiotemporal forest structure is then recalculated.

[0065] Sort all the spatiotemporal trees in the merged spatiotemporal forest structure according to their tree depth from largest to smallest. Then, prune each spatiotemporal tree in the merged spatiotemporal forest structure according to the order of the sorting to obtain the pruned spatiotemporal forest structure. Continue until the number of all nodes in the pruned spatiotemporal forest structure reaches the preset number of tokens. Finally, take the set of all nodes in the pruned spatiotemporal forest structure as the final node set.

[0066] In a specific embodiment, an adjacency matrix is ​​constructed based on the similarity matrix, spatial distance matrix, and timestamp, specifically including:

[0067] Traverse any two nodes, denoted as the i-th node and the j-th node. The semantic similarity between the i-th node and the j-th node is the element in the i-th row and j-th column of the similarity matrix. The distance between the i-th node and the j-th node is the element in the i-th row and j-th column of the spatial distance matrix. ,like Greater than or equal to the similarity threshold If the timestamp of the i-th node is less than or equal to the distance threshold, and the timestamp of the i-th node is less than the timestamp of the j-th node, then the element in the i-th row and j-th column of the adjacency matrix is... If it is 1, otherwise the element in the i-th row and j-th column of the adjacency matrix It is 0.

[0068] Record the coordinates of each node in the original frame. and timestamp Then, calculate the global semantic similarity and distance to construct a similarity matrix. Spatial distance matrix This provides a basis for cross-frame connections based on the principles of semantic consistency, spatial proximity, and temporal order.

[0069] In one embodiment, the similarity matrix The elements in the matrix use cosine similarity to calculate the semantic similarity between two nodes, while the spatial distance matrix... The distance between two nodes is calculated using the L2 norm. Further, a similarity threshold is applied. Distance threshold And the order of timestamps will filter out two nodes that may have a directed edge, when and Furthermore, it mandates that the time direction can only be from morning to night. In the case of ) This allows us to retain only potential cross-frame correspondences that satisfy the triple constraints. This indicates whether an edge can be connected between the i-th node and the j-th node, ultimately determined by... Construct an adjacency matrix .

[0070] First, through the neighbor matrix The root node is determined by the in-degree and out-degree of each node, and the root set is constructed. Specifically, the root node has an in-degree of 0 and has outgoing edges, that is, when... , At that time, determine To identify the root node and avoid isolated nodes, a node may belong to multiple root nodes. To determine the root node to which the remaining nodes not in the root set belong, a sorting matrix needs to be constructed. As shown in the following formula: The weighting coefficient This is used to control the trade-off between semantic similarity and spatial distance, and is achieved using a ranking matrix. The size of the elements in the set determines which node is assigned to the most matching root node. That is, for the k-th node that is not in the root set, the node is assigned to the most matching root node. Winning As its root node, where, From sorting matrix The corresponding elements; then all those that satisfy The node is assigned to the first The set of nodes of a spacetime tree Therefore, we can obtain the node set of each spacetime tree. This refers to which root node the k-th node belongs to. It refers to the first The root node of the k-th time-space tree is equal to the root node of the k-th time-space tree. If they are equal, it means that the k-th node belongs to the k-th time-space tree. The root node of the time-space tree.

[0071] In obtaining Afterwards, use To truly connect the nodes into a directed spacetime tree, where, Represents the set of directed edges of a spacetime tree. This represents the node depth. The spatiotemporal forest algorithm traverses time, connecting the node at the current moment to those nodes that satisfy the given conditions. The tree depth is accumulated at nodes from earlier times (i.e., nodes that can be connected). This allows cross-frame persistence to be quantified using tree depth. Tree depth reflects the sequential connections between nodes corresponding to key visual token embeddings in video frames captured at different timestamps. Nodes corresponding to key visual token embeddings in video frames captured at earlier timestamps are closer to the root node, while nodes corresponding to key visual token embeddings in video frames captured at the current timestamp are connected to the nodes corresponding to key visual token embeddings in video frames captured at even earlier timestamps. For example, if the current timestamp is 5s, the nodes corresponding to key visual token embeddings in video frames captured at timestamps 0 to 4s can be connected to the nodes corresponding to key visual token embeddings in video frames captured between 0s and 4s. Therefore, a single spatiotemporal tree can be constructed from a single root node and its constituent nodes, and several spatiotemporal trees can form the initial spatiotemporal forest structure.

[0072] Furthermore, after constructing the initial spatiotemporal forest structure, taking a preset number of tokens as an example, when the number of root nodes in the initial spatiotemporal forest structure is much greater than the preset number of tokens, a greedy algorithm is first used to merge the most similar spatiotemporal trees. That is, based on the semantic similarity between two root nodes, the two most similar spatiotemporal trees are merged. Specifically, the merging process connects the root nodes of the two most similar spatiotemporal trees to the earlier root node, reducing the instability and resource waste caused by too many spatiotemporal trees with a small number of nodes. Then, a unified sorting and pruning process is initiated. In the sorting stage, the tree depth of each spatiotemporal tree in the merged spatiotemporal forest structure is calculated, and then the spatiotemporal trees are sorted in descending order according to their tree depth.

[0073] In a specific embodiment, the pruning process is as follows:

[0074] First, remove the outermost node of each spatiotemporal tree in the merged spatiotemporal forest structure, and count the number of nodes remaining after removal;

[0075] If the number of remaining nodes after removal does not reach the preset number of Tokens, then remove the second outermost node from each spatiotemporal tree in the merged spatiotemporal forest structure. Repeat this step to remove nodes layer by layer until the number of remaining nodes after removal reaches the preset number of Tokens. At this point, the remaining nodes after removal constitute the pruned spatiotemporal forest structure.

[0076] Specifically, during the pruning phase, the spatiotemporal trees with the highest tree depth are pruned first. Then, each spatiotemporal tree is pruned sequentially according to its tree depth. In one example, the spatiotemporal trees are sorted horizontally by tree depth and vertically by the timestamps of their root nodes. The outermost node of each spatiotemporal tree is removed first. If the number of remaining nodes after removing all the outermost nodes does not reach the preset token count, the next outermost node is removed, and so on, until the number of remaining nodes reaches the preset token count. The final set of nodes constitutes the final node set. The key visual token embeddings corresponding to the nodes in the final node set can then participate in other computational processes within the multimodal large language model, thereby reducing the computational cost of redundant visual token embeddings.

[0077] The technical effects of the embodiments of this application are demonstrated below through specific experiments.

[0078] This experiment aims to systematically evaluate the impact of different visual token compression methods on the performance of large multimodal video models under high compression ratios (70% / 80% / 90%), with a focus on verifying:

[0079] (1) In high compression ratio scenarios, do existing image compression or local compression methods experience significant performance degradation?

[0080] (2) Whether the method proposed in the embodiments of this application can maintain stable performance under the condition of cross-frame global modeling;

[0081] (3) Robustness and generalization ability of different methods on various video understanding datasets.

[0082] The experiments in the embodiments of this application, under a unified base model and inference configuration, applied the same compression ratio to various methods such as FastV, VisionZip, GPRune, STTM, and FrameFusion, ensuring a consistent number of input visual token embeddings and guaranteeing fairness in the comparison. Furthermore, experiments were conducted on the NExT-QA, VideoMME, MLVU, LongVideoBench, and MVBench datasets.

[0083] Using the LLaVA-Video 7B model as the multimodal large language model, the experimental procedure is as follows:

[0084] 1. Visual token embedding and extraction are performed on the input video using a visual encoder based on a multimodal large language model;

[0085] 2. Perform token compression at the specified compression ratio;

[0086] 3. Input the compressed token into the LLaVA-Video 7B model;

[0087] 4. Evaluate the accuracy of the multimodal large language model on various benchmarks;

[0088] 5. Calculate the performance retention rate (relative to the unpruned model). Pay special attention to the performance of the multimodal large language model under the extreme compression condition of 90% compression rate (retaining only 10% of the token count).

[0089] The experimental results are shown in Table 1. This experiment systematically verifies that high-ratio visual token compression is an extremely challenging problem in video multimodal tasks; existing methods generally experience significant performance degradation at compression rates above 80%; the method proposed in this application maintains stable performance at compression rates of 70%, 80%, and 90%; and even at an extreme compression rate of 90%, it significantly outperforms all comparative methods. This indicates that the method proposed in this application can maintain high-quality video understanding capabilities with extremely low computational budgets; it is more suitable for long videos and large-scale inference scenarios; and it provides a feasible solution for high-compression inference in video MLLM.

[0090] Table 1

[0091] The visual token compression method for a multimodal large language model based on spatiotemporal forest proposed in this application constructs a spatiotemporal forest structure across frames, which can extract the main key visual token embeddings across the time dimension and avoid redundant and repeated tokens; it uses semantic similarity, spatial distance and temporal order to jointly control the directed edges between two nodes; it performs priority selection based on the depth and structural role (root node, trunk node, leaf node) of the node in the spatiotemporal tree; it can maintain the high performance of the multimodal large language model even with a compression ratio of up to 90%, is compatible with existing multimodal large language models, and requires no training.

[0092] Further reference Figure 4 As an implementation of the methods shown in the above figures, this application provides an embodiment of a visual token compression device based on a spatiotemporal forest multimodal large language model. This device embodiment is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0093] This application provides a visual token compression device based on a spatiotemporal forest multimodal large language model, including:

[0094] The node feature construction module 1 is configured to input each video frame in the input video into the multimodal large language model, and obtain the visual token embedding sequence corresponding to each video frame through the visual encoder in the multimodal large language model. Several key visual tokens are selected from the visual token embedding sequence corresponding to each video frame and node features are constructed.

[0095] Compression module 2 is configured to input node features into the spatiotemporal forest algorithm. In the spatiotemporal forest algorithm, each key visual token is embedded as a node. A directed edge between two nodes is constructed based on the joint constraints of semantic similarity, spatial distance and temporal order. The nodes are constructed into a spatiotemporal tree and assembled into a spatiotemporal forest structure through the directed edge. The spatiotemporal forest structure is processed under preset constraints to obtain the final node set. The key visual token embeddings corresponding to the nodes in the final node set are used as the compressed visual token embedding sequence and participate in the calculation process of other modules in the multimodal large language model.

[0096] Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. For example... Figure 4As shown, the electronic device in this embodiment includes a processor 401 and a memory 402; wherein the memory 402 is used to store computer execution instructions; and the processor 401 is used to execute the computer execution instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.

[0097] Alternatively, the memory 402 can be either standalone or integrated with the processor 401.

[0098] When the memory 402 is set up independently, the electronic device also includes a bus 403 for connecting the memory 402 and the processor 401.

[0099] This invention also provides a computer storage medium storing computer execution instructions, which, when executed by processor 401, implement the above method.

[0100] This invention also provides a computer program product, including a computer program, which, when executed by a processor 401, implements the above-described method.

[0101] In the embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0102] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.

[0103] Furthermore, the functional modules in the various embodiments of this invention can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit formed by the above modules can be implemented in hardware or in the form of hardware plus software functional units.

[0104] The integrated modules implemented as software functional modules described above can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor 401 to execute some steps of the methods of the various embodiments of this application.

[0105] It should be understood that the processor 401 described above can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor, or the processor 401 can be any conventional processor 401. The steps of the method disclosed in this invention can be directly manifested as the hardware processor 401 executing the steps, or as a combination of hardware and software modules within the processor 401 executing the steps.

[0106] The memory 402 may include high-speed RAM memory, and may also include non-volatile memory NVM, such as at least one disk storage device, and may also be a USB flash drive, portable hard drive, read-only memory, disk or optical disc, etc.

[0107] Bus 403 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 403 can be divided into address bus, data bus, control bus, etc. For ease of illustration, the bus 403 in the accompanying drawings of this application is not limited to only one bus 403 or one type of bus 403.

[0108] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.

[0109] An exemplary storage medium is coupled to processor 401, enabling processor 401 to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of processor 401. Processor 401 and storage medium can reside in application-specific integrated circuits (ASICs). Alternatively, processor 401 and storage medium can exist as discrete components in an electronic device or host device.

[0110] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A visual token compression method based on a multimodal large language model using spatiotemporal forest, characterized in that, Includes the following steps: Each video frame in the input video is input into the multimodal large language model. After passing through the visual encoder in the multimodal large language model, the visual token embedding sequence corresponding to each video frame is obtained. Several key visual tokens are selected from the visual token embedding sequence corresponding to each video frame and node features are constructed. The node features are input into the spatiotemporal forest algorithm. In the spatiotemporal forest algorithm, each key visual token embedding is used as a node. A directed edge between two nodes is constructed based on the joint constraints of semantic similarity, spatial distance, and temporal order. The nodes are constructed into a spatiotemporal tree and a spatiotemporal forest structure through the directed edge. The spatiotemporal forest structure is then processed under preset constraints to obtain the final node set. The key visual token embeddings corresponding to the nodes in the final node set are used as a compressed visual token embedding sequence and participate in the calculation process of other modules in the multimodal large language model.

2. The visual token compression method based on a multimodal large language model of spatiotemporal forest according to claim 1, characterized in that, The key visual token embedding is obtained by selecting visual token embeddings from the visual token embedding sequence using pruning or merging methods, and the node feature is a vector formed by concatenating all key visual token embeddings.

3. The visual token compression method based on a multimodal large language model of spatiotemporal forest according to claim 1, characterized in that, The processing of the spatiotemporal forest structure includes merging, sorting, and pruning spatiotemporal trees.

4. The visual token compression method based on a multimodal large language model of spatiotemporal forest according to claim 1, characterized in that, A directed edge is constructed between two nodes based on joint constraints of semantic similarity, spatial distance, and temporal order. The nodes are then used to build a spatiotemporal tree and a spatiotemporal forest structure. This spatiotemporal forest structure is processed under a preset number of tokens to obtain the final node set, specifically including: For each node, record its coordinates and timestamp in the corresponding video frame; Calculate the semantic similarity between the key visual token embeddings corresponding to any two nodes, and construct a similarity matrix; Calculate the distance between the coordinates of any two nodes and construct the spatial distance matrix; Construct an adjacency matrix based on the similarity matrix, spatial distance matrix, and timestamp; The root node is determined based on the out-degree and in-degree of the elements in the adjacency matrix, and the root set is constructed. The sorting matrix is ​​constructed based on the adjacency matrix, similarity matrix, and spatial distance matrix, as shown in the following formula: ; in, Represents the Hadema product. Indicates the weighting coefficient. Represents the similarity matrix. Represents the spatial distance matrix. Represents the adjacency matrix. Represents the sorting matrix; For the k-th node that is not in the root set, select the root node of the column corresponding to the largest value among all elements in the k-th row of the sorting matrix in the root set as the root node to which the k-th node belongs. Count the nodes to which each root node belongs and obtain the corresponding node set. According to the timestamp, connect the current node in the node set corresponding to each root node with the nodes at previous times through directed edges to construct a spatiotemporal tree with different levels. Several spatiotemporal trees form the initial spatiotemporal forest structure. The preset constraint includes a preset number of tokens. In response to determining that the number of root nodes in the root set is much greater than the preset number of tokens, the two corresponding spatiotemporal trees in the initial spatiotemporal forest structure are merged according to the semantic similarity between the two root nodes to obtain a merged spatiotemporal forest structure, and the tree depth of each spatiotemporal tree in the merged spatiotemporal forest structure is recalculated. Sort all the spatiotemporal trees in the merged spatiotemporal forest structure according to their tree depth from largest to smallest. Then, prune each spatiotemporal tree in the merged spatiotemporal forest structure according to the order of the sorting to obtain a pruned spatiotemporal forest structure. Continue until the number of all nodes in the pruned spatiotemporal forest structure reaches a preset number of tokens. Finally, take the set of all nodes in the pruned spatiotemporal forest structure as the final node set.

5. The visual token compression method based on a multimodal large language model of spatiotemporal forest according to claim 4, characterized in that, Constructing an adjacency matrix based on the similarity matrix, spatial distance matrix, and timestamps specifically includes: Traverse any two nodes, denoted as the i-th node and the j-th node. The semantic similarity between the i-th node and the j-th node is the element in the i-th row and j-th column of the similarity matrix. The distance between the i-th node and the j-th node is the element in the i-th row and j-th column of the spatial distance matrix. ,like Greater than or equal to the similarity threshold If the distance threshold is less than or equal to the distance threshold, and the timestamp of the i-th node is less than the timestamp of the j-th node, then the element in the i-th row and j-th column of the adjacency matrix... If it is 1, otherwise the element in the i-th row and j-th column of the adjacency matrix is ​​1. It is 0.

6. The visual token compression method based on a multimodal large language model of spatiotemporal forest according to claim 2, characterized in that, The specific process of pruning is as follows: First, remove the outermost node of each spatiotemporal tree in the merged spatiotemporal forest structure, and count the number of nodes remaining after removal; If the number of remaining nodes after removal does not reach the preset number of Tokens, then remove the second outermost node from each spatiotemporal tree in the merged spatiotemporal forest structure. Repeat this step to remove nodes layer by layer until the number of remaining nodes after removal reaches the preset number of Tokens. At this point, the remaining nodes after removal constitute the pruned spatiotemporal forest structure.

7. A visual token compression device based on a spatiotemporal forest multimodal large language model, characterized in that, include: The node feature construction module is configured to input each video frame in the input video into a multimodal large language model, and obtain the visual token embedding sequence corresponding to each video frame through the visual encoder in the multimodal large language model. Then, select several key visual tokens from the visual token embedding sequence corresponding to each video frame and construct node features. The compression module is configured to input the node features into the spatiotemporal forest algorithm. In the spatiotemporal forest algorithm, each key visual token embedding is used as a node. A directed edge between two nodes is constructed based on the joint constraints of semantic similarity, spatial distance, and temporal order. The nodes are constructed into a spatiotemporal tree and a spatiotemporal forest structure through the directed edge. The spatiotemporal forest structure is processed under preset constraints to obtain a final node set. The key visual token embeddings corresponding to the nodes in the final node set are used as a compressed visual token embedding sequence and participate in the calculation process of other modules in the multimodal large language model.

8. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.