Model inference method, apparatus, equipment and media based on key-value cache compression

CN122220386BActive Publication Date: 2026-09-01SHENZHEN RES INST OF BIG DATA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610686432.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-09-01
Estimated Expiration
2046-05-19

AI Technical Summary

Technical Problem

然而,这种方式虽然能够直接减少缓存规模,但可能导致关键信息被过度剪枝,或保留过多冗余,从而难以在显著提升模型推理效率的同时维持原有的推理精度

Benefits of technology

[0015]本申请实施例通过获取文本查询序列与对应的图像帧序列,并通过预训练的视觉大语言模型中的每个注意力层,基于文本查询序列与图像帧序列,计算文本查询序列对应的每个文本令牌对应的文本键值对,以及图像帧序列对应的每个视觉令牌对应的视觉键值对;基于每个注意力层针对图像帧序列的语义提取粒度,从视觉大语言模型包含的多个注意力层中确定压缩边界层;从多个注意力层中确定位于压缩边界层前序的多个第一注意力层,基于每个第一注意力层与压缩边界层之间的层间距离关系,计算每个第一注意力层对应的压缩比例;在每个第一注意力层中,按照压缩比例对每个视觉令牌对应的视觉键值对进行剪枝,得到视觉键值对缓存;从多个注意力层中确定位于压缩边界层后序的多个第二注意力层,并在每个第二注意力层中,对每个视觉令牌对应的所有视觉键值对进行剪枝;基于每个注意力层对应的文本键值对,确定每个第一注意力层对应的第一文本键值对,以及每个第二注意力层对应的第二文本键值对,并依次基于每个第一注意力层对应的第一文本键值对和视觉键值对缓存,以及每个第二注意力层对应的第二文本键值对进行推理,生成文本查询序列对应的推理结果。以此,能够针对不同网络层采用差异化的分层压缩策略。具体来说,由于浅层网络通常关注局部细节、深层网络聚焦语义信息,本申请根据层间距离关系为不同层分配非统一的压缩比例,使得关键信息密集的层保留更多视觉键值对,而冗余较多的层则进行更激进压缩,从而避免了采用相同固定比例剪枝所导致的关键信息过度剪枝或冗余保留过多的问题。综上,本申请能够在显著提升模型推理效率的同时维持原有的推理精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122220386B_ABST
    Figure CN122220386B_ABST
Patent Text Reader

Abstract

This application provides a model inference method, apparatus, device, and medium based on key-value cache compression. The method includes: acquiring a text query sequence and an image frame sequence; calculating the text key-value pair for each text token and the visual key-value pair for each visual token through each attention layer in a visual large language model; determining a compression boundary layer based on the semantic extraction granularity of each attention layer for the image frame sequence; determining the compression ratio of each first attention layer preceding the compression boundary layer; pruning the visual key-value pairs corresponding to each visual token in the first attention layer according to the compression ratio to obtain a visual key-value pair cache; pruning all visual key-value pairs for each visual token in the second attention layer following the compression boundary layer; and performing inference sequentially based on the first text key-value pairs and visual key-value pair caches of the first attention layer and the second text key-value pairs of the second attention layer to generate the corresponding inference result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a model reasoning method, apparatus, device and medium based on key-value caching compression. Background Technology

[0002] Key-value cache compression is a key technique used to reduce storage and computational overhead in the autoregressive generation of large language models based on the Transformer architecture. Specifically, when the model generates each new token, it stores the key and value corresponding to that token in a cache for use in attention calculations for subsequent tokens. Key-value cache compression aims to transform the ever-growing original key-value cache into a more compact representation through strategies such as filtering, merging, or recoding. This reduces memory usage and latency associated with attention calculations while maintaining model generation quality as much as possible.

[0003] In related technologies, a unified compression strategy is generally used to process the key-value cache in the visual large language model, and this is used as the basis for model inference. Specifically, for the key-value cache of all historical visual tokens, pruning is performed in each attention module layer according to the same and fixed retention ratio to discard token representations deemed unimportant. However, while this approach can directly reduce the cache size, it may lead to excessive pruning of key information or retention of too much redundancy, making it difficult to maintain the original inference accuracy while significantly improving model inference efficiency. Summary of the Invention

[0004] This application proposes a model inference method, apparatus, device, and medium based on key-value caching compression, which can significantly improve model inference efficiency while maintaining the original inference accuracy.

[0005] To achieve the above objectives, a first aspect of this application proposes a model inference method based on key-value caching compression, the method comprising: Obtain the text query sequence and the corresponding image frame sequence, and through each attention layer in the pre-trained visual large language model, calculate the text key-value pair corresponding to each text token of the text query sequence and the visual key-value pair corresponding to each visual token of the image frame sequence based on the text query sequence and the image frame sequence. Based on the semantic extraction granularity of each attention layer for the image frame sequence, a compression boundary layer is determined from the multiple attention layers contained in the visual large language model; From the plurality of attention layers, determine a plurality of first attention layers that precede the compression boundary layer, and calculate the compression ratio corresponding to each first attention layer based on the interlayer distance relationship between each first attention layer and the compression boundary layer; In each first attention layer, the visual key-value pairs corresponding to each visual token are pruned according to the compression ratio to obtain a visual key-value pair cache. From the plurality of attention layers, determine a plurality of second attention layers located after the compressed boundary layer, and in each second attention layer, prune all visual key-value pairs corresponding to each visual token; Based on the text key-value pairs corresponding to each attention layer, the first text key-value pairs corresponding to each first attention layer and the second text key-value pairs corresponding to each second attention layer are determined. Then, reasoning is performed sequentially based on the first text key-value pairs and the visual key-value pairs cached for each first attention layer, and the second text key-value pairs corresponding to each second attention layer, to generate the reasoning result corresponding to the text query sequence.

[0006] Accordingly, a second aspect of this application proposes a model inference apparatus based on key-value cache compression, the apparatus comprising: The acquisition module is used to acquire the text query sequence and the corresponding image frame sequence, and through each attention layer in the pre-trained visual large language model, calculate the text key-value pair corresponding to each text token of the text query sequence and the visual key-value pair corresponding to each visual token of the image frame sequence based on the text query sequence and the image frame sequence. The first determining module is used to determine the compression boundary layer from the multiple attention layers contained in the visual large language model based on the semantic extraction granularity of each attention layer for the image frame sequence. The calculation module is used to determine a plurality of first attention layers located before the compression boundary layer from the plurality of attention layers, and to calculate the compression ratio corresponding to each first attention layer based on the interlayer distance relationship between each first attention layer and the compression boundary layer. The pruning module is used to prune the visual key-value pairs corresponding to each visual token in each first attention layer according to the compression ratio to obtain a visual key-value pair cache. The second determining module is used to determine a plurality of second attention layers located after the compressed boundary layer from the plurality of attention layers, and in each second attention layer, to prune all visual key-value pairs corresponding to each visual token; The generation module is used to determine the first text key-value pair corresponding to each first attention layer and the second text key-value pair corresponding to each second attention layer based on the text key-value pair corresponding to each attention layer, and to perform inference sequentially based on the first text key-value pair and the visual key-value pair cache corresponding to each first attention layer and the second text key-value pair corresponding to each second attention layer to generate the inference result corresponding to the text query sequence.

[0007] In some embodiments, the acquisition module is further configured to: Obtain the text query sequence and its corresponding multiple initial image frames; The text query sequence is encoded using a pre-defined lightweight visual language model to obtain a text feature vector, and each initial image frame is encoded to obtain an image feature vector. Based on the first similarity between the text feature vector and each image feature vector, multiple target image frames are selected from the multiple initial image frames; Based on the multiple target image frames, an image frame sequence is constructed.

[0008] In some embodiments, the acquisition module is further configured to: The initial image frame corresponding to the image feature vector with the highest similarity to the text feature vector is determined as the first target image frame, and the first target image frame is added to the image frame set; For each initial image frame, calculate the second similarity between the image feature vector and the target image feature vector of each target image frame in the image frame set, and select the maximum similarity among the second similarities with each target image frame in the image frame set; For each initial image frame, a first weighted value is calculated based on the product of a preset first balance coefficient and the first similarity, and a second weighted value is calculated based on the product of a preset second balance coefficient and the maximum similarity. The comprehensive score of each initial image frame is calculated based on the difference between the first weighted value and the second weighted value, wherein the first balance coefficient is greater than the second balance coefficient. Based on the comprehensive score corresponding to each initial image frame, the next target image frame is selected from the plurality of initial image frames, the next target image frame is added to the image frame set, and based on the next target image frame, the next target image frame is determined from the remaining initial image frames and added to the image frame set; The step of determining the next target image frame from the remaining initial image frames and adding it to the image frame set based on the next target image frame is repeated until the total number of target image frames in the image frame set reaches a preset threshold, thus obtaining multiple target image frames.

[0009] In some embodiments, the computing module is further configured to: Determine the first layer number corresponding to the compressed boundary layer, and the second layer number corresponding to each first attention layer, wherein the type of the first attention layer includes the compressed boundary layer; Based on the difference between the first layer number corresponding to the compressed boundary layer and the preset reference value, the first offset is determined, and based on the square of the first offset, the target first offset is obtained. For each first attention layer, a second offset is determined based on the difference between the second layer number and the preset reference value, and a target second offset is obtained based on the square of the second offset. Based on the first target offset and the second target offset corresponding to each first attention layer, intermediate parameters are obtained; Based on the difference between the preset reference value and the intermediate parameter corresponding to each first attention layer, the compression ratio corresponding to each first attention layer is calculated.

[0010] In some embodiments, the pruning module is further used for: In each first attention layer, the attention score corresponding to each visual token is obtained, wherein the first attention layer includes a compression boundary layer; Obtain the total number of visual tokens corresponding to each first attention layer, and calculate the number of retained tokens corresponding to the current first attention layer based on the total number of visual tokens and the compression ratio. Based on the size relationship between the attention scores of multiple visual tokens corresponding to the first attention layer, a target visual token is selected from the multiple visual tokens with the same number of tokens to be retained; Based on the target visual token, the visual key-value pairs corresponding to the remaining visual tokens are pruned to obtain the pruning result. Based on the pruning results and the target visual key-value pairs corresponding to the target visual token, a visual key-value pair cache for each first attention layer is obtained.

[0011] In some implementations, the target visual key-value pair includes a first visual key vector and a first visual value vector, and the pruning module is further configured to: Based on the pruning results, determine the visual value vector to be processed for each other visual token; For each other visual token corresponding to the visual value vector to be processed, calculate the value similarity between it and the first visual value vector corresponding to each target visual token. Based on the value similarity between each first visual value vector and each visual value vector to be processed, each visual value vector to be processed is fused into each first visual value vector to obtain the target first visual value vector corresponding to each target visual token. In each first attention layer, the first visual key vector and the first visual value vector corresponding to each target visual token are cached to obtain the visual key-value pair cache of each first attention layer.

[0012] In some implementations, the second determining module is further configured to: Obtain the second visual key vector and second visual value vector corresponding to all visual tokens in each second attention layer; Prune the second visual key vector and the second visual value vector corresponding to all visual tokens in each second attention layer.

[0013] Accordingly, a third aspect of the embodiments of this application proposes a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the model inference method based on key-value cache compression according to any one of the embodiments of the first aspect of this application.

[0014] Accordingly, a fourth aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the model inference method based on key-value cache compression according to any one of the embodiments of the first aspect of this application.

[0015] This application embodiment obtains a text query sequence and its corresponding image frame sequence, and through each attention layer in a pre-trained visual large language model, calculates the text key-value pair corresponding to each text token in the text query sequence and the visual key-value pair corresponding to each visual token in the image frame sequence based on the text query sequence and the image frame sequence; based on the semantic extraction granularity of each attention layer for the image frame sequence, it determines a compression boundary layer from the multiple attention layers included in the visual large language model; it determines multiple first attention layers located before the compression boundary layer from the multiple attention layers, and calculates the compression ratio corresponding to each first attention layer based on the inter-layer distance relationship between each first attention layer and the compression boundary layer; in each In each first attention layer, visual key-value pairs corresponding to each visual token are pruned according to a compression ratio to obtain a visual key-value pair cache. Multiple second attention layers are determined from the multiple attention layers, following the compression boundary layer. In each second attention layer, all visual key-value pairs corresponding to each visual token are pruned. Based on the text key-value pairs corresponding to each attention layer, the first text key-value pairs corresponding to each first attention layer and the second text key-value pairs corresponding to each second attention layer are determined. Reasoning is then performed sequentially based on the first text key-value pairs and visual key-value pair caches corresponding to each first attention layer, as well as the second text key-value pairs corresponding to each second attention layer, to generate the reasoning result corresponding to the text query sequence. This allows for differentiated layered compression strategies for different network layers. Specifically, since shallow networks typically focus on local details and deep networks focus on semantic information, this application allocates non-uniform compression ratios to different layers based on inter-layer distance relationships. This allows layers with dense key information to retain more visual key-value pairs, while layers with more redundancy undergo more aggressive compression, thus avoiding the problems of over-pruning of key information or excessive retention of redundancy caused by using the same fixed pruning ratio. In summary, this application can significantly improve the model's inference efficiency while maintaining the original inference accuracy. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the architecture of the model inference system based on key-value caching compression provided in the embodiments of this application; Figure 2 This is a flowchart of the model inference method based on key-value cache compression provided in the embodiments of this application; Figure 3 This is a flowchart of the overall model inference method based on key-value caching compression provided in the embodiments of this application; Figure 4 This is a schematic diagram of the functional modules of the model inference device based on key-value cache compression provided in the embodiments of this application; Figure 5 This is a schematic diagram of the hardware structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0018] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0020] Key-value cache compression is a key technique used to reduce storage and computational overhead in the autoregressive generation of large language models based on the Transformer architecture. Specifically, when the model generates each new token, it stores the key and value corresponding to that token in a cache for use in attention calculations for subsequent tokens. Key-value cache compression aims to transform the ever-growing original key-value cache into a more compact representation through strategies such as filtering, merging, or recoding. This reduces memory usage and latency associated with attention calculations while maintaining model generation quality as much as possible.

[0021] In related technologies, a unified compression strategy is generally used to process the key-value cache in the visual large language model, and this is used as the basis for model inference. Specifically, for the key-value cache of all historical visual tokens, pruning is performed in each attention module layer according to the same and fixed retention ratio to discard token representations deemed unimportant. However, while this approach can directly reduce the cache size, it may lead to excessive pruning of key information or retention of too much redundancy, making it difficult to maintain the original inference accuracy while significantly improving model inference efficiency.

[0022] Based on this, the embodiments of this application provide a model inference method, apparatus, device and medium based on key-value cache compression, which can significantly improve the model inference efficiency while maintaining the original inference accuracy.

[0023] The model inference method, apparatus, device and medium based on key-value cache compression provided in this application are specifically described through the following embodiments. First, the model inference system based on key-value cache compression in this application is described.

[0024] Please refer to Figure 1 In some implementations, embodiments of this application provide a model inference system based on key-value caching compression, including a terminal 11 and a server 12.

[0025] In some implementations, terminal 11 can be used to acquire the text query sequence and corresponding image frame sequence input by the user, and to present the final reasoning result to the user. For example, it can be a hardware device such as a smartphone, tablet, personal computer, smart camera, or augmented reality glasses. Terminal 11 can send text and image / video data to server 12 through its configured communication module and user interface, and receive and display the reasoning result returned by server 12, thereby realizing the functions of data acquisition and result presentation.

[0026] Furthermore, the server-side component 12 can be used to run pre-trained visual large language models. For example, it can be a standalone GPU server, cloud server cluster, or edge computing node. Through its configured model inference engine and key-value cache compression module, the server-side component 12 performs functions such as determining the compression boundary layer, calculating the compression ratio based on inter-layer distance relationships, hierarchical pruning of visual key-value pairs, and accelerating autoregressive inference, thereby completing the entire computation process from text and image frame sequences to inference results.

[0027] Specifically, terminal 11 can encapsulate the text query sequence and image frame sequence collected by the user into a request message and send it to server 12; server 12 runs a key-value caching compression method, generates an inference result, and returns it to terminal 11 in the form of a response message; terminal 11 receives and parses the result, and finally presents it to the user. In this way, efficient visual large language model inference can be achieved on resource-constrained terminal devices, while maintaining model accuracy by utilizing the powerful computing capabilities of the server.

[0028] The model inference method based on key-value cache compression in this application can be illustrated by the following embodiments.

[0029] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user will be obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent will the necessary user-related data for the normal operation of the embodiments of this application be obtained.

[0030] In this application embodiment, the description will focus on a model inference device based on key-value cache compression, which can be integrated into a computer device. See also Figure 2 , Figure 2 The flowchart illustrates the steps of the key-value cache compression-based model inference method provided in this embodiment. Taking the key-value cache compression-based model inference device specifically integrated into a terminal or server as an example, the specific process when the processor on the terminal or server executes the program instructions corresponding to the key-value cache compression-based model inference method is as follows: Step 101: Obtain the text query sequence and the corresponding image frame sequence, and through each attention layer in the pre-trained visual large language model, calculate the text key-value pair corresponding to each text token in the text query sequence and the visual key-value pair corresponding to each visual token in the image frame sequence based on the text query sequence and the image frame sequence.

[0031] In some implementations, in order to accurately eliminate the large amount of redundant computation and storage overhead caused by visual modalities during the reasoning process of visual large language models, and to fundamentally solve the problem of the explosive growth of key-value cache size caused by the number of visual tokens being far greater than that of text tokens, the parallel input text query sequence and image frame sequence can be pre-computed in each attention layer of the model to explicitly decouple and generate their respective text key-value pairs and visual key-value pairs. This provides a clear data foundation for subsequent selective compression and pruning of visual key-value pairs based on hierarchical semantic differences, and achieves the initial decoupling of reasoning efficiency and accuracy.

[0032] The text query sequence can be a piece of natural language text input by the user to query image or video content, such as "What are the people in the video doing?" or "Is there a cat in the picture?", which can be used to guide the visual big language model to focus on specific semantic information in the visual input.

[0033] The image frame sequence can be multiple image frames arranged in chronological order corresponding to the text query scenario. For example, it can be several frames of images obtained by uniformly sampling or filtering from an input video.

[0034] Among them, the visual large language model can be a deep learning model that can process visual and language modal inputs in parallel, such as the Qwen2.5-VL series models or multimodal large models. The visual large language model contains multiple sequential attention layers.

[0035] The attention layer can be a basic neural network layer that performs self-attention or cross-attention calculations within the visual large language model, and each attention layer independently maintains its own key-value cache.

[0036] Among them, a text token can be the smallest processing unit obtained after segmenting or sub-segmenting a text query sequence, such as a word, a character, or a sub-word fragment, which can be used to represent the semantic content of the text query in a discretized form.

[0037] In this context, the text key-value pair can be a set of key vectors and value vectors obtained by linear transformation mapping of the hidden representation corresponding to the text token in each attention layer. For example, the key is used to calculate the attention weight with the query, and the value is used to obtain the output by weighted summation. This can be used to characterize the semantic features of the text token in the current attention layer.

[0038] The visual token can be an embedding vector representation of multiple image blocks obtained by dividing each frame of an image into blocks or by convolutional encoding. For example, it can be a token sequence generated by dividing a 224x224 image into 196 16x16 visual blocks, which can be used to represent the local visual content of an image or video frame in a discretized form.

[0039] In this context, visual key-value pairs can be a set of key vectors and value vectors obtained by linear transformation mapping of the hidden representation corresponding to the visual token in each attention layer. Their scale is much larger than that of text key-value pairs, and they can be used to characterize the visual semantic features of the visual token in the current attention layer.

[0040] For example, a text query sequence in natural language form provided by a user through an input interface can be received, such as "What are the people in the video doing?"; at the same time, visual input data corresponding to the query can be obtained. The visual input data can be a single image, multiple images, or a complete video file read from local storage, network, or real-time acquisition device. By uniformly or keyframe sampling the video file, multiple initial image frames arranged in chronological order are obtained, and these initial image frames are directly used as the image frame sequence.

[0041] Further, after completing the preparation of input data, the text query sequence and the image frame sequence can be jointly input into a pre-trained large vision-language model (e.g., the Qwen2.5-VL-3B-Instruct model) to perform the first forward propagation calculation. It should be noted that in actual inference deployment, the actual input of the model usually integrates three parts: a fixed system cache (whose corresponding key-value cache is derived from system prompts, such as "You are a helpful assistant." and other instructions for setting the model's behavior and role), a variable prompt cache (corresponding to the text query sequence input by the user this time), and a vision cache (corresponding to the image frame sequence input this time). These three parts of caches can be concatenated in order to form a complete input sequence.

[0042] Specifically, in the Prefill stage, the model can process this complete input sequence in parallel. Specifically in each attention layer, for each element (token) in the input sequence, the model can map its hidden state to Query, Key and Value respectively through an independent linear transformation matrix. For the text part therein, each text token (e.g., the three tokens ["Please", "describe", "weather"] obtained after tokenizing "please describe the weather") will generate corresponding text query vectors, text key vectors and text value vectors after mapping. The model organizes the text key vectors and text value vectors into text key-value pairs, and uses the text query vectors for the attention calculation of the current step.

[0043] Further, the text key-value pairs can be stored in different caches according to their sources: the key-value pairs corresponding to fixed system prompts are stored in the system cache, and the key-value pairs corresponding to the user's current query are stored in the prompt cache. The text key-value pairs in these two caches will be completely retained in subsequent inference without any pruning, so as to ensure the model's ability to follow instructions and fully understand the text context.

[0044] Meanwhile, in the forward propagation of the same Prefill stage, for the input image frame sequence, each image frame can first be divided into multiple vision patches. For example, an image frame with a resolution of 224x224 can be divided into 196 16x16 vision patches. After passing through a linear embedding layer, each vision patch can be converted into an initial vision token. In each internal attention layer of the model, the hidden state of each vision token is also mapped to its corresponding vision query vector, vision key vector and vision value vector respectively through an independent linear transformation matrix. The model organizes the vision key vectors and vision value vectors into vision key-value pairs, and uses the vision query vectors to calculate the attention weight between the current vision token and all tokens (including text and vision) in the sequence.

[0045] Furthermore, these visual key-value pairs can be temporarily stored in an initialized visual key-value cache. The number of these visual key-value pairs is far greater than the number of text tokens, and they are the main source of storage and computational overhead during inference. At this point, the model has generated complete queries, keys, and values ​​for all tokens in the current entire input sequence (system part, text query part, image frame part) in each attention layer, providing the necessary data foundation for subsequent autoregressive generation and the hierarchical compression operation of this application.

[0046] The above methods can provide clear operational objects and data basis for subsequent visual token pruning strategies based on hierarchical compression boundaries, thereby laying a data foundation for significantly compressing visual key-value cache, reducing memory usage, and accelerating the inference process while maintaining model accuracy.

[0047] In some implementations, to fundamentally reduce the size of subsequent visual key-value caching and inference latency, a pre-built lightweight visual language model can be used to jointly semantically encode the text query and each initial image frame. Based on the first similarity between the text feature vector and the image feature vector, target image frames that are most relevant to the query semantics and have low redundancy are selected. This allows for the construction of a compact, high-semantic-density image frame sequence with extremely low computational cost, significantly reducing the total number of visual tokens that the main model needs to process. For example, step 101, "obtaining the text query sequence and the corresponding image frame sequence," may include: (101.1) Obtain the text query sequence and the corresponding multiple initial image frames; (101.2) The text query sequence is encoded using a pre-defined lightweight visual language model to obtain a text feature vector, and each initial image frame is encoded to obtain an image feature vector; (101.3) Based on the first similarity between the text feature vector and each image feature vector, select multiple target image frames from multiple initial image frames; (101.4) Construct an image frame sequence based on multiple target image frames.

[0048] The initial image frame can be any frame in the original image frame sequence of the input video without any semantic filtering.

[0049] Among them, the lightweight visual language model can be a multimodal model with a parameter scale much smaller than the main visual large language model and the ability to jointly encode vision and language. For example, it can be a contrastive language-image pre-trained model or its encoder part, which can be used to quickly map images and text to feature vectors in the same semantic space and evaluate the semantic relevance of the two with extremely low computational overhead.

[0050] Among them, the text feature vector can be a dense vector representation of fixed dimensions obtained by encoding the text query sequence through a text encoder of a lightweight visual language model. It can be used to represent the core semantic content of the user query in the semantic space and serve as a query benchmark for filtering keyframes.

[0051] The image feature vector can be a dense vector representation with the same dimension as the text feature vector, obtained by encoding each initial image frame through an image encoder of a lightweight visual language model. For example, it can be a normalized feature vector output by the image encoder of a contrastive language-image pre-trained model, which can be used to represent the visual semantic content of each frame of the image in the semantic space so as to compare similarity with the text feature vector.

[0052] The first similarity can be a semantic similarity measure between a text feature vector and an image feature vector, such as cosine similarity or dot product similarity, which can be used to quantify the degree of semantic matching between the current image frame and the text query.

[0053] The target image frame can be a subset of image frames selected from multiple initial image frames based on a first similarity and optionally combined with diversity optimization criteria (such as the maximum marginal relevance algorithm), which are highly relevant to the semantics of the text query and have low redundancy among each other. This subset can be used to construct a compact and information-dense sequence of image frames.

[0054] In some implementations, the text query sequence can be a natural language question or instruction input by the user through any human-computer interaction interface, such as "What sport are the people in the video doing?" or "Please describe the weather conditions of the scene in the picture." Multiple initial image frames can be a sequence of raw frames sampled from a complete video file at fixed time intervals (e.g., 1 frame per second) or by a keyframe extraction algorithm (e.g., based on scene transition detection); alternatively, multiple independently input single-frame images can be received directly. For example, for a 30-second video at 30fps, uniform sampling can yield 30 initial image frames. Before filtering, these initial image frames are typically numerous (e.g., tens to hundreds of frames), and directly inputting them into the main model would cause an explosive growth in visual tokens.

[0055] Furthermore, to quickly assess the semantic relevance between the initial image frames and the text query with low computational cost, a pre-defined lightweight visual language model is invoked to encode the text query sequence. This lightweight model (e.g., the text encoder of CLIP (Contrastive Language-Image Pre-training)) has the advantages of small parameter size and fast inference speed, and its encoding process does not depend on the main model. Specifically, the text query sequence can be segmented and embedded before being fed into the text encoder of the lightweight model. After undergoing multiple Transformer transformations, a dense vector of fixed dimensions, i.e., the text feature vector, is output. For example, for the text query "a person is playing ball," the CLIP text encoder can output a text feature vector with a dimension of 512. This text feature vector is normalized to a unit hypersphere for efficient cosine similarity calculation with the image feature vector. This text feature vector only needs to be calculated once in the entire filtering process and represents the core semantics of the user query.

[0056] Specifically, a lightweight visual language model can be used to encode each initial image frame separately, obtaining a corresponding image feature vector. Specifically, each initial image frame (e.g., 224×224 resolution) can be input into a lightweight model's image encoder (e.g., CLIP's Vision Transformer or a ResNet-based encoder). After image patch segmentation, linear embedding, and multi-layer feature extraction, a dense vector with the same dimension as the text feature vector is output—the image feature vector. For the aforementioned 30 uniformly sampled initial image frames, 30 image feature vectors can be generated in parallel or sequentially, each vector also having a dimension of 512, and also undergoing normalization. These image feature vectors encode the visual semantic content of each frame with extremely low computational overhead (compared to running a large visual language model), providing a foundation for subsequent similarity calculations.

[0057] Furthermore, after obtaining the text feature vector and all image feature vectors, a first similarity between each initial image frame and the text query can be calculated, and multiple target image frames can be selected accordingly. In some implementations, the first similarity can be measured using cosine similarity, or other measurement methods can be used. This application does not limit the specific similarity measurement method, as long as it does not depart from the concept of this application. For example, 30 similarity values ​​[0.12, 0.05, 0.78, ..., 0.23] can be calculated.

[0058] In some implementations, to ensure relevance while also considering diversity among the selected frames (avoiding the selection of frames with highly repetitive content), the Maximum Marginal Relevance (MMR) algorithm can be used for iterative selection. In the first round, the frame with the highest first similarity (e.g., frame 3, similarity 0.78) is directly selected as the first target image frame. In subsequent rounds, for each candidate frame that has not yet been selected, a second similarity (also using cosine similarity) between it and the image feature vector of each frame in the selected frame set can be calculated. The maximum value is taken as a redundancy penalty term, and then a comprehensive score is calculated: Comprehensive score = First balance coefficient × First similarity between the frame and the text. (1 The first balance coefficient is multiplied by the maximum second similarity between this frame and any frame in the selected frame set. The candidate frame with the highest comprehensive score is selected as the next target image frame. This process is repeated until the number of selected target image frames reaches a preset threshold (e.g., preset to 5 frames based on the processing capacity of the main model or task requirements). In this way, 5 target image frames that are both highly relevant to the query semantics and significantly different from each other can be obtained.

[0059] Finally, all the target image frames obtained through iterative filtering (e.g., in timestamp order from the original video) are organized into a compact image frame sequence. The length of this sequence (number of target frames) is much smaller than the number of initial image frames (e.g., 5 frames is much smaller than 30 frames), but it retains the core visual content most relevant to the text query. This image frame sequence will replace the original multiple initial image frames as the visual input to the main visual large language model for subsequent visual key-value pair generation and inference.

[0060] This construction method can significantly reduce the number of visual tokens entering the main model from the source, significantly reduce the size of subsequent visual key-value cache and the overhead of attention computation, while the model can still obtain enough visual context information to answer queries due to diversity constraints, thus achieving a good balance between inference efficiency and accuracy.

[0061] For example, besides using a fixed preset threshold (e.g., selecting 5 frames) to terminate the MMR iteration selection, dynamic features of the video content can be introduced to adaptively determine the number of target frames. Specifically, after acquiring multiple initial image frames, the optical flow or pixel difference between adjacent frames is first calculated to obtain an inter-frame motion amplitude sequence. When the overall motion amplitude of the video is large (e.g., a fast-moving sports scene), the preset threshold is increased (e.g., from 5 frames to 10 frames) to retain more temporal details; when the overall motion amplitude of the video is small (e.g., a static scene or talk show), the preset threshold is decreased (e.g., reduced to 3 frames). In this way, the length of the constructed image frame sequence can be matched with the information density of the video content, avoiding redundancy caused by excessive retention in static scenes and avoiding the loss of key information in fast-moving scenes, thereby further improving the efficiency and robustness of subsequent inference.

[0062] By using the above methods, the number of visual tokens entering the main model can be significantly reduced from the source, effectively removing redundant frames (such as repeating backgrounds, irrelevant objects, etc.) in the original long video sequence. This provides a low-redundancy, high-density input basis for the subsequent compression processing of visual key-value caching, while avoiding the ineffective computation and memory waste caused by the main model when processing a large number of irrelevant or repeating frames.

[0063] In some implementations, to simultaneously consider the relevance to the text query and the diversity between frames when selecting target image frames, and to avoid selecting a bunch of target frames with similar but redundant content, thereby constructing a compact image frame sequence that closely revolves around the query semantics and fully covers different visual content, the first round can select the frame most relevant to the text feature vector as the first target frame. In subsequent rounds, for each candidate frame, the second similarity between it and each frame in the selected frame set is calculated, and the maximum value is taken as a redundancy penalty term. Then, a weighted difference operation is performed between this penalty term and the first similarity to obtain a comprehensive score. The frame with the highest score is repeatedly selected and added to the set until the number reaches a preset threshold, so as to achieve a dynamic balance and joint optimization of relevance and diversity. For example, (101.3) may include: (101.3.1) The initial image frame corresponding to the image feature vector with the highest similarity to the text feature vector is determined as the first target image frame, and the first target image frame is added to the image frame set; (101.3.2) For each initial image frame, calculate the second similarity between the image feature vector and the target image feature vector of each target image frame in the image frame set, and select the maximum similarity among the second similarities between the initial image frame and the target image frame in the image frame set; (101.3.3) For each initial image frame, a first weighted value is calculated based on the product of a preset first balance coefficient and a first similarity, and a second weighted value is calculated based on the product of a preset second balance coefficient and the maximum similarity. The comprehensive score of each initial image frame is calculated based on the difference between the first weighted value and the second weighted value, wherein the first balance coefficient is greater than the second balance coefficient. (101.3.4) Based on the comprehensive score corresponding to each initial image frame, select the next target image frame from multiple initial image frames, add the next target image frame to the image frame set, and determine the next target image frame from the remaining initial image frames based on the next target image frame and add it to the image frame set; (101.3.5) Repeat the step of determining the next target image frame from the remaining initial image frames and adding it to the image frame set based on the next target image frame until the total number of target image frames in the image frame set reaches a preset threshold, thus obtaining multiple target image frames.

[0064] The target image frame can be an image frame that is highly relevant to the text query and has low redundancy among itself, selected iteratively from the initial image frames. Specifically, it can be the frame most relevant to the text selected in the first round and the frame with the highest comprehensive score in subsequent rounds.

[0065] The image frame set can be a dynamically growing container used to store selected image frames as targets, such as an initially empty list that is appended to each selected frame in each round. It can be used to record the selected frames so that the visual redundancy between candidate frames and selected frames can be calculated later.

[0066] The target image feature vector can be the image feature vector corresponding to each selected target image frame in the image frame set. It can be used to calculate the second similarity with the image feature vector of the candidate frame to measure the degree of visual content similarity between the two frames.

[0067] The second similarity can be a similarity metric between the image feature vector of a candidate initial image frame and the target image feature vector of a target image frame in the image frame set, such as cosine similarity or dot product similarity. It can be used to quantify the degree of visual content redundancy between two frames and serve as a basis for diversity penalty.

[0068] The maximum similarity can be the maximum value obtained after calculating the second similarity between a candidate initial image frame and each target image frame in the image frame set. For example, the similarity between the candidate frame and the most similar frame in the selected set can be used to characterize the maximum redundancy between the candidate frame and the current selected frame set, and can be incorporated into the comprehensive score as a penalty term.

[0069] The first balance coefficient can be a preset hyperparameter used to adjust the weight of the first similarity (relevance item) in the overall score. Its value is greater than the second balance coefficient. For example, it can be set to 0.7, which can be used to control the contribution of text relevance to frame selection.

[0070] The first weighted value can be the product of the first balance coefficient and the first similarity, and can be used to represent the weighted relevance score between the candidate frame and the text query.

[0071] The second balancing coefficient can be a preset hyperparameter used to adjust the weight of maximum similarity (diversity penalty) in the overall score. Its value is less than that of the first balancing coefficient, for example, it can be set to 0.3. It can be used to control the contribution intensity of visual redundancy penalty to frame selection.

[0072] The second weighting value can be the product of the second balance coefficient and the maximum similarity, and can be used to represent the weighted redundancy penalty value between the candidate frame and the selected frame set.

[0073] The comprehensive score can be a value calculated for each candidate initial image frame by subtracting the second weighted value from the first weighted value (i.e., the difference between the first weighted value and the second weighted value). It can be used to comprehensively evaluate the balance between relevance and diversity of candidate frames. The higher the score, the more worthy the frame is to be selected.

[0074] The preset threshold can be an upper limit for the total number of target image frames. For example, it can be the number of frames to be retained, K, which can be preset according to computing resources or task requirements. When the number of frames in the image frame set reaches this threshold, the iterative selection stops. This can be used to control the balance between the compactness and information density of the final image frame sequence.

[0075] In some implementations, to select the first frame most relevant to the text query from multiple initial image frames as the basis for subsequent iterations, the initial image frame corresponding to the image feature vector with the highest first similarity to the text feature vector can be determined as the first target image frame. Specifically, after calculating the first similarity (e.g., the cosine similarity between the text feature vector and the image feature vector) of each initial image frame, the maximum value is found through sorting or comparison operations. For example, assuming there are 30 initial image frames, and their first similarities calculated by a lightweight visual language model are [0.12, 0.05, 0.78, 0.23, ..., 0.34], with the third frame having the highest similarity of 0.78, this third frame is determined as the first target image frame. It is understood that the image frame most similar to the user query is most likely to contain the key visual information needed to answer the query, and using it as a seed frame can ensure that the final constructed image frame sequence has high semantic relevance.

[0076] Furthermore, to record the selected target image frames for evaluating redundancy between candidate frames and selected frames in subsequent iterations, a dynamically growing set of image frames needs to be maintained. After determining the first target image frame, its frame identifier (e.g., frame index or timestamp) and its corresponding image feature vector are immediately stored in an initially empty set of image frames, S. For example, S = {Frame 3}, while simultaneously saving the image feature vector of that frame. This set is progressively expanded in each iteration to serve as a benchmark for calculating the diversity penalty term (i.e., the maximum similarity between candidate frames and selected frames). By maintaining this set, subsequent rounds of selection can explicitly avoid selecting frames that are too visually similar to selected frames, thus achieving a balance between relevance and diversity.

[0077] After selecting the first frame, for each initial image frame that has not yet been selected (i.e., a candidate frame), the visual similarity between it and each target image frame in the image frame set can be calculated, which is called the second similarity. The second similarity can also be measured using cosine similarity to quantify the degree of redundancy in visual content between the two frames.

[0078] Specifically, for each candidate frame i, its image feature vector is: For each selected frame j in the image frame set S (feature vector is...) The second similarity can be calculated. Since all feature vectors have been normalized, the calculation can be simplified to a dot product. For example, in one iteration, the image frame set S contains {frame 3, frame 8}, and for the candidate frame 12, Sim( , )=0.85 and Sim( , =0.42. This allows us to quantify the degree of visual overlap between the candidate frame and each selected frame.

[0079] For example, to represent the maximum redundancy of a candidate frame with the entire set of selected frames using a single numerical value, the maximum value can be selected from the second similarity scores between the candidate frame and all selected frames in the set. The maximum similarity score represents the level of redundancy between the candidate frame and the most similar frame in the set of selected frames. Continuing with the previous example, candidate frame 12 has a similarity of 0.85 with frame 3 and 0.42 with frame 8, so the maximum value is 0.85. The larger this value, the closer the candidate frame is visually to a particular frame in the set; selecting it would introduce less new information (i.e., low diversity). Conversely, a small maximum similarity score indicates that the candidate frame is significantly different from all selected frames, providing a higher visual content gain.

[0080] Furthermore, to simultaneously evaluate the relevance (the degree of matching with the text query) and diversity (the degree of difference from the selected frames) of candidate frames, a first weighted value can be obtained based on the product of a preset first balance coefficient and a first similarity, and a second weighted value can be obtained based on the product of a preset second balance coefficient and the maximum similarity. The difference between the first weighted value and the second weighted value is used as the comprehensive score. The first balance coefficient is greater than the second balance coefficient to reflect the emphasis on relevance.

[0081] In some implementations, the overall score can be determined using the following formula: ; in, This represents the overall score of candidate frame i; This is the first balance coefficient, with a value greater than 0.5, for example, 0.7; The first similarity; This is the second balance coefficient, for example, 0.3; The maximum similarity mentioned above. The first weighted value is... The second weighted value is For example, candidate frame A has a first similarity of 0.9 and a maximum similarity of 0.8, so its overall score is 0.7 × 0.9 - 0.3 × 0.8 = 0.63 - 0.24 = 0.39; candidate frame B has a first similarity of 0.8 and a maximum similarity of 0.2, so its score is 0.7 × 0.8 - 0.3 × 0.2 = 0.56 - 0.06 = 0.50. Although B has a slightly lower relevance, it is preferred because of its greater contribution to diversity and thus has a higher overall score.

[0082] Furthermore, based on the comprehensive score corresponding to each initial image frame, the frame with the highest comprehensive score can be selected from all currently unselected candidate frames as the next target image frame, and this frame can be added to the image frame set. After being added, this frame becomes part of the selected frames, which will affect the diversity calculation in subsequent rounds. Then, based on the updated image frame set, the next target image frame can be determined from the remaining initial image frames using the same method. For example, after the first round of iteration, the image frame set S = {frame 3}; the comprehensive score of all remaining frames is calculated, and assuming that frame 8 has the highest score, it is selected as the second target frame and added to S, at which point S = {frame 3, frame 8}; then, based on S = {frame 3, frame 8}, the maximum similarity and comprehensive score of the remaining frames are calculated, and the third target frame is selected. This process is repeated, and each step uses the latest selected frame set to dynamically evaluate the redundancy penalty term of the remaining frames.

[0083] Furthermore, the above-mentioned step of "determining the next target image frame from the remaining initial image frames based on the next target image frame" can be repeated. That is, in each round, the comprehensive score of all remaining candidate frames is recalculated and the highest scorer is selected to be added to the image frame set until the total number of target image frames in the image frame set reaches a preset threshold.

[0084] For example, the preset threshold can be pre-set based on the needs of downstream tasks, the processing power of the main model, or the length of the video, such as 5 or 8 frames. The iteration terminates when the number of frames in set S equals the threshold; at this point, all frames in set S are the selected target image frames. These target image frames maintain a high relevance to the text query while having low visual redundancy, representing the most crucial and diverse visual content of the original video with a very small number of frames. Subsequently, these target image frames can be used to construct an image frame sequence, which is then input into the main visual large language model for subsequent inference acceleration processing.

[0085] For example, the value of the first balancing coefficient α can be dynamically adjusted based on the first similarity distribution of the candidate frame set. Specifically, at the beginning of each iteration, the variance or entropy of the first similarity of all remaining candidate frames is calculated. When the similarity distribution is concentrated (i.e., the relevance of each frame to the text is not significantly different), the value of α is appropriately reduced (e.g., from 0.7 to 0.5) to increase the weight of the diversity penalty, thus encouraging the selection of more diverse frames. When the similarity distribution is dispersed (with a few frames having extremely high relevance), the value of α is increased to prioritize relevance. In this way, excessive loss of diversity when relevance is generally high, or forced selection of irrelevant frames when relevance is generally low, can be avoided, making the screening strategy more adaptable to different text-content matching scenarios.

[0086] By employing the above methods, candidate frames that are visually too similar to the selected frames can be explicitly penalized in each iteration, while candidate frames that are highly relevant to the text are retained. This effectively avoids the problem of high redundancy in the target frame sequence caused by traditional sorting based solely on relevance. This ensures that the selected set of target image frames, while semantically closely revolving around the user query, covers different visual content in the video to the greatest extent possible. Consequently, it provides the subsequent main model with a high-quality image frame sequence that has higher information density and lower redundancy, further reducing the storage and computational burden of the visual key-value cache from the source, and improving the inference accuracy and efficiency of the model in long video understanding tasks.

[0087] Step 102: Based on the semantic extraction granularity of each attention layer for the image frame sequence, determine the compression boundary layer from the multiple attention layers contained in the visual large language model.

[0088] In some implementations, in order to overcome the shortcomings of the prior art that adopts a uniform compression strategy for all attention layers and ignores the semantic granularity differences of visual information between deep and shallow layers, and to achieve layered differentiated key-value cache compression, a compression boundary layer can be dynamically determined by analyzing the degree of abstraction of visual features output by each attention layer for the image frame sequence, that is, the characteristic of semantic extraction granularity gradually evolving from local details in the shallow layer to global semantics in the deep layer. This layer serves as the boundary to distinguish between the shallow fine preservation area and the deep large compression area, thereby providing a reasonable structural basis for the subsequent layer-by-layer allocation of asymmetric compression budget.

[0089] The semantic extraction granularity can be the degree of semantic abstraction represented by the visual features output by each attention layer when processing image frame sequences. For example, the features of shallow attention layers mostly correspond to local details such as edges and textures (fine granularity), while the features of deep attention layers mostly correspond to global semantic information such as objects and scenes (coarse granularity). This can be used to measure the importance and compressibility of the visual information of this layer and guide the differentiated allocation of the layer compression budget.

[0090] The compression boundary layer can be a specific layer pre-determined or dynamically determined from the multiple attention layers contained in the visual large language model, based on the semantic extraction granularity change pattern of each layer (such as the inflection point from fine granularity to coarse granularity). For example, the Nth layer (such as the 2 / 3 point of the total number of layers or the layer selected based on attention score statistics). This layer and the layers before it (pre-sequence layers) adopt a progressive compression strategy, while the layers after it (post-sequence layers) adopt complete pruning or truncation. This can be used to divide the different processing stages of key-value cache compression, and realize the layered optimization of shallow fine preservation and deep compression.

[0091] In some implementations, to achieve differentiated compression for different attention layers, a compression boundary layer can be determined from the multiple attention layers included in the visual large language model, based on the semantic extraction granularity of the visual features output by each attention layer for the image frame sequence. Specifically, in the visual large language model, shallow attention layers (such as layers 1 to 8) mainly extract fine-grained local visual features such as edges, textures, and colors, with a lower degree of semantic abstraction; while deep attention layers (such as layers 20 to 32) gradually focus on coarse-grained global semantic features such as objects, scenes, and relationships. To retain crucial details in the shallow layers while compressing redundant abstract representations in the deep layers, a compression boundary layer can be set according to the evolution of semantic extraction granularity from fine to coarse. For example, in a 32-layer model (numbered 1 to 32), the location where the semantic extraction granularity changes significantly (such as layer 12) can be determined as the compression boundary layer by analyzing the differences in attention weight distribution or feature maps of each layer.

[0092] For example, the selection principle for the compression boundary layer can be: the attention layer before the boundary layer (the first attention layer) still retains certain visual details and adopts a progressive compression strategy; in the attention layer after the boundary layer (the second attention layer), the visual representation has been highly abstracted and no longer depends on the fine-grained details of specific visual tokens. Therefore, completely pruning these visual key-value pairs will not significantly affect the model's understanding of the overall scene, thereby maximizing the compression benefits with almost no loss of accuracy.

[0093] In some implementations, the number of the compressed boundary layer This can be predetermined in the following ways: ; in This is a preset boundary scaling factor (e.g., 0.375). This represents the total number of layers in the model. This is the floor function. For example, when... =32、 When =0.375, =12, meaning the 12th layer is used as the compression boundary layer. This coefficient can be adjusted between 0.3 and 0.5 depending on the model structure and task requirements to balance the preservation of detail with the compression gains.

[0094] Furthermore, in specific experiments and deployments, to achieve the optimal trade-off between inference speed and inference performance, a compressed boundary layer can be set based on the proportion of the total number of layers. Experiments have shown that the closer the truncated layer is to the bottom layer (i.e., the smaller the compressed boundary layer number), the more attention layers are assigned to subsequent second attention layers and visual key-value pairs are completely truncated, thus resulting in faster model inference speed, but inference accuracy (such as ROUGE-L or precision) will decrease accordingly. Conversely, the closer the truncated layer is to the top layer, the more layers of visual information are retained, resulting in higher inference accuracy but limited speedup. To balance these two aspects, this application conducted systematic experiments on multiple visual language models (such as Qwen2.5-VL-3B / 32B) and video understanding benchmarks (such as NExTQA and EgoSchema). The experiments show that when the truncated layer is set to 0.75 times the total number of layers (i.e., at 75% depth), significant inference speedup (e.g., a speedup ratio of 2.35 times) can be achieved while maintaining almost no loss in model accuracy. For example, in a 32-layer model, the cutoff layer is set to the 24th layer (32 × 0.75 = 24), that is, layers 1 to 24 are the first attention layer, and layers 25 to 32 are the second attention layer.

[0095] In some implementations, if the total storage limit for the key-value cache is set more generously (e.g., allowing more visual tokens to be retained), the truncation layer can be adjusted towards higher levels (e.g., set to 0.8 times the total number of layers) to retain more visual information. Conversely, if storage resources are extremely limited or higher inference speed is required, the truncation layer can be adjusted towards lower levels (e.g., set to 0.6 times the total number of layers) to sacrifice a small amount of precision for a greater speedup. The specific adjustments can be made according to the actual situation to flexibly adapt to different deployment scenarios and hardware constraints.

[0096] In some implementations, the compression boundary layer can be dynamically determined based on the real-time attention distribution of the model under a given input. Specifically, after completing the key-value pair calculation for all tokens in the Prefill phase, for each attention layer, the entropy value of the attention score of all visual tokens in the token dimension can be calculated: ; Where K is the total number of visual tokens in this layer. The normalized attention weights for the k-th visual token in layer l (satisfying) The smaller the entropy value, the more focused the attention is on a few key tokens, resulting in higher visual information redundancy in that layer, making it more suitable for being classified as a compression region. By setting an entropy threshold, the first layer with an entropy value below that threshold is identified as the compression boundary layer. This allows the selection of the compression boundary layer to be correlated with the specific input content. For simple scenarios (focused attention), an earlier boundary layer is used for aggressive compression, while for complex scenarios (distracted attention), a later boundary layer is used to retain more information, thereby further improving the adaptive balance between inference efficiency and accuracy.

[0097] In some implementations, the compression boundary layer can also be determined by calculating the similarity of visual attention patterns between adjacent attention layers. Specifically, for each layer, the cosine similarity between the attention distribution vector of all visual tokens in that layer and the corresponding attention distribution vector in the next layer is calculated. When the similarity drops sharply (i.e., the attention patterns of adjacent layers change significantly), this location can be considered a semantic transition point in the flow of visual information and can be used as the compression boundary layer. For example, if the attention distribution similarity drops sharply from 0.85 to 0.45 from layer 12 to layer 13, then layer 12 is determined as the compression boundary layer. By utilizing the structural features of the information flow within the model, the inflection point of the evolution of visual semantics from fine-grained to coarse-grained can be captured more accurately, thereby achieving a hierarchical compression partition that is more in line with the characteristics of the model itself.

[0098] By using the above methods, the natural evolution of semantic extraction granularity from details to the global among attention layers can be used as a basis for division. In this way, the compression strategy can be precisely aligned with the information flow characteristics within the model, so that visual key-value pairs carrying fine-grained local details in shallow layers are fully preserved, while highly abstract redundant representations in deep layers are effectively pruned. This provides a reasonable and dynamic structural boundary for subsequent parabolic hierarchical compression budget allocation and pruning and fusion operations for different layers, maximizing compression benefits while maintaining the model's ability to perceive key visual information.

[0099] Step 103: Determine multiple first attention layers located before the compression boundary layer from multiple attention layers, and calculate the compression ratio corresponding to each first attention layer based on the interlayer distance relationship between each first attention layer and the compression boundary layer.

[0100] In some implementations, in order to achieve differentiated compression of visual information in different network layers and to match the compression intensity with the natural evolution of the semantic extraction granularity of each layer (shallow layers retain more details, and deep layers retain less abstract information), multiple first attention layers preceding the compression boundary layer can be determined. Based on the interlayer distance relationship between each first attention layer and the compression boundary layer (e.g., the closer the distance, the greater the compression ratio), a parabolic dynamic allocation strategy is used to calculate the compression ratio corresponding to each first attention layer, so that different layers obtain differentiated compression intensities that are adapted to the importance of their information.

[0101] The first attention layer can be a collective term for multiple attention layers located before the compression boundary layer (including the compression boundary layer itself) in a visual large language model. For example, in a model with a total of 32 layers, if the compression boundary layer is the 12th layer, then layers 1 to 12 are all first attention layers, which can be used as objects to perform progressive key-value caching compression. The compression strength gradually increases with the number of layers.

[0102] The interlayer distance relationship can be a quantitative representation of the difference in network depth between each first attention layer and the compression boundary layer. For example, the interlayer distance between the 5th layer and the boundary layer (the 12th layer) is 7 layers. This distance can be used to determine the compression ratio corresponding to the layer. Generally, the smaller the distance (i.e., the closer to the boundary layer), the larger the compression ratio, and the larger the distance, the smaller the compression ratio. This can be used to achieve differentiated compression budget allocation consistent with the semantic granularity evolution trend.

[0103] The compression ratio can be the degree to which visual key-value pairs need to be pruned in each first attention layer or the proportion of visual tokens to be retained. For example, a compression ratio of 0.6 means that 60% of the original number of visual tokens is retained. This ratio is dynamically calculated based on the inter-layer distance relationship (such as a parabolic function) and can be used to guide the specific number of visual key-value pairs to be pruned in each first attention layer, thereby achieving hierarchical adaptive compression intensity control.

[0104] In some implementations, after determining the compression boundary layer, all attention layers preceding the compression boundary layer (including the compression boundary layer itself) can be selected from all attention layers of the visual large language model and designated as multiple first attention layers. For example, if the total number of layers in the model is 32 and the index of the compression boundary layer is h (e.g., h=12), then the first attention layers are layers 1, 2, ..., 12. These first attention layers undertake the progressive extraction task from fine-grained local features to medium-grained semantics. They require differentiating compression ratios (i.e., the proportion of visual tokens retained) based on the depth difference between each layer and the compression boundary layer to achieve a progressive strategy of less compression in shallow layers and more compression in deep layers.

[0105] Furthermore, to dynamically allocate the compression ratio based on the interlayer distance between each first attention layer and the compression boundary layer, this application employs a parabolic compression budget allocation function. Specifically, for the L-th layer (L is the index of the first attention layer, and 1≤L≤h), its corresponding compression ratio (i.e., the proportion of visual tokens retained by this layer to the total number of original visual tokens) is... It can be calculated using the following formula: ; in, This represents the proportion of visual tokens that need to be retained in the Lth layer; L is the index of the current first attention layer; h is the index of the compression boundary layer (e.g., h=12); the constant 1 in the formula indicates that the retention ratio is 100% in the shallowest layer (L=1); the negative quadratic term makes the retention ratio decrease smoothly with increasing layer depth.

[0106] The above formula is based on the property that a parabola opens downwards: when L=1, =0, therefore =1, meaning all visual tokens in layer 1 are retained; when L=h... With the denominator Dividing by 1 / 2, therefore =0.5, meaning that 50% of the visual tokens are retained at the compression boundary layer; the retention ratio of the intermediate layers decreases with layer depth according to a quadratic curve. Through this parabolic function, the compression ratio of each first attention layer decreases smoothly with increasing layer depth according to a quadratic curve. In this way, layered differential compression that matches the evolution of visual semantic extraction granularity from fine to coarse can be achieved.

[0107] By employing the above methods, shallow layers (far from the boundary layer) can achieve a smaller compression ratio to preserve fine-grained local visual details, while deep layers (closer to the boundary layer) can achieve a larger compression ratio to aggressively remove highly abstract redundant information. This provides a scientific quantitative basis for performing precise pruning operations in each first attention layer that match semantic importance, maximizing the compression of GPU memory and computational overhead while maintaining the model's perceptual integrity of key visual features to the greatest extent possible.

[0108] In some implementations, to ensure that the compression ratio of each first attention layer varies with the network depth difference between it and the compression boundary layer in a quadratic curve (the compression ratio of layers closer to the boundary layer is larger, and the compression ratio of layers farther away is smaller), a smooth, non-linear compression intensity distribution function that highly matches the semantic granularity evolution trend can be established by calculating the compression ratio corresponding to each first attention layer. For example, step 103, "calculating the compression ratio corresponding to each first attention layer based on the inter-layer distance relationship between each first attention layer and the compression boundary layer," may include: (103.1) Determine the first layer number corresponding to the compressed boundary layer and the second layer number corresponding to each first attention layer, wherein the type of the first attention layer includes the compressed boundary layer; (103.2) Based on the difference between the first layer number corresponding to the compressed boundary layer and the preset reference value, the first offset is determined, and the target first offset is obtained based on the square of the first offset; (103.3) For each first attention layer, the second offset is determined based on the difference between the second layer number and the preset reference value, and the target second offset is obtained based on the square of the second offset; (103.4) Based on the first offset of the target and the second offset of the target corresponding to each first attention layer, intermediate parameters are obtained; (103.5) Based on the difference between the preset reference value and the intermediate parameter corresponding to each first attention layer, calculate the compression ratio corresponding to each first attention layer.

[0109] The first layer number can be the position number of the compressed boundary layer in all attention layers of the visual large language model, such as the 12th layer of a 32-layer model (if the numbering starts from 1, the first layer number is 12).

[0110] The second layer number can be the position number of each first attention layer in all attention layers of the visual large language model. For example, the corresponding numbers 1 to 12 for each of the 1st, 2nd, ..., 12th layers can be used to represent the depth position of the layer in the network.

[0111] The preset reference value can be a fixed value set in advance, such as 1.

[0112] The first offset can be the difference between the first layer number of the compressed boundary layer and the preset reference value, for example, 12-33=-21. Its absolute value reflects the distance between the compressed boundary layer and the reference point.

[0113] The target first offset can be the square of the first offset, which can be used to eliminate the influence of positive and negative signs and amplify the distance difference, serving as the numerator of the parabolic function.

[0114] The second offset can be the difference between the second layer number of each first attention layer and a preset reference value. For example, 5-33=-28 for the 5th layer, which can be used to characterize the distance between the layer and the reference point.

[0115] The target second offset can be the square of the second offset.

[0116] The intermediate parameter can be the ratio obtained by dividing the first offset of the target by the second offset of the target of the current first attention layer, for example, 441 / 784≈0.5625. It can be used to reflect the ratio of the squared distance between the compression boundary layer and the reference point, thereby controlling the numerical trend of the compression ratio.

[0117] For example, the first layer number is denoted as h, which is the position number of the compressed boundary layer in all attention layers of the model (e.g., for a model with a total of 32 layers and the compressed boundary layer being the 12th layer, h=12). The first attention layer includes the compressed boundary layer itself and all layers preceding it, totaling h layers. The second layer number corresponding to each first attention layer is denoted as L, where L=1,2,…,h, corresponding to the 1st, 2nd,…, and 12th layers, respectively. These numbers are used to quantify the depth distance between each layer and a preset reference point, providing a basis for subsequent offset calculations.

[0118] Furthermore, the preset reference value can be set to 1, then the first offset is h. 1. The first offset of the target is This square operation amplifies the depth difference, providing non-linear weights for the subsequent parabolic compression ratio allocation.

[0119] For example, for each first attention layer, a second offset can be determined based on the difference between its second layer index L and the same preset reference value (i.e., 1), and its squared value can be calculated to obtain the target second offset. The second offset is... The second offset of the target is For example, for layer 6 (L=6), the second offset is 5, and the target second offset is 25; for layer 12 (L=12), the second offset is 11, and the target second offset is 121. This target second offset reflects the depth of the current layer relative to a preset reference point and will be used in conjunction with the target first offset to construct intermediate parameters.

[0120] Furthermore, the intermediate parameter can be defined as the ratio of the second target offset to twice the first target offset, that is: ; The coefficient 2 in the denominator controls the compression ratio at the compressible boundary layer. This intermediate parameter increases with depth, reflecting the cumulative compressive intensity from shallow to deep layers.

[0121] Then, based on the difference between the preset reference value (i.e., 1) and the intermediate parameters corresponding to each first attention layer, the compression ratio corresponding to each first attention layer can be calculated. This refers to the proportion of visual tokens that this layer needs to retain relative to the total number of original visual tokens. The specific process is as follows: ; The compression ratio of the intermediate layers decreases smoothly with increasing L following a quadratic curve. Each first attention layer obtains a differentiated compression ratio that matches its depth and semantic extraction granularity, providing a quantitative basis for subsequent precise pruning.

[0122] In some implementations, a distance remapping related to the rate of change of attention patterns can be introduced to replace the linear distance (L) in the formula. 1) Define the cosine similarity between the attention distribution vectors of two adjacent layers. And construct the cumulative difference distance Then the formula Replace with Therefore, for layers where attention patterns change drastically (rapid semantic evolution), The growth is faster and the compression ratio decreases more rapidly, thus more accurately matching the actual abstraction process of visual information and further improving the balance between compression efficiency and inference accuracy.

[0123] By using the above methods, the compression ratio can increase non-linearly with increasing layer depth (the compression ratio increases sharply when it is close to the boundary layer and gradually increases when it is far from the boundary layer). This allows for the allocation of differentiated compression intensity to each first attention layer that precisely matches the level of abstraction of visual information. This avoids the problems of excessive compression in shallow layers or insufficient compression in deep layers caused by linear compression. It ensures that the fine-grained visual details in shallow layers are fully preserved while achieving the maximum compression of redundant semantic representations in deep layers.

[0124] Step 104: In each first attention layer, prune the visual key-value pairs corresponding to each visual token according to the compression ratio to obtain a visual key-value pair cache.

[0125] In some implementations, in order to retain the key semantic information in visual key-value pairs after pruning and compression, and to avoid a significant drop in model performance due to the direct discarding of unimportant tokens, visual key-value pairs can be selectively retained or discarded in each first attention layer based on a pre-calculated compression ratio and importance indicators such as the attention score corresponding to each visual token (e.g., retaining key-value pairs corresponding to visual tokens with higher attention scores). This results in a compact visual key-value pair cache that still retains the core visual information, thereby maximizing the maintenance of the model's generation quality while significantly reducing storage and computational overhead.

[0126] The visual key-value pair cache can be a compact storage structure that retains only the key vector and value vector corresponding to the key visual token after pruning according to the compression ratio in each first attention layer.

[0127] In some implementations, in each first attention layer (e.g., the Lth layer, 1≤L≤h, where h is the compression boundary layer number), the compression ratio is calculated. (That is, the proportion of the number of visual tokens that need to be retained in this layer to the total number of original visual tokens), and prune the visual key-value pairs corresponding to each visual token to obtain the compressed visual key-value pair cache.

[0128] Specifically, first, the attention scores of all visual tokens in this layer are obtained (e.g., the sum or maximum value of the attention scores corresponding to each visual token in the attention weights calculated through the Prefill stage). Then, the visual tokens are sorted from high to low according to their attention scores, and the number to be retained is calculated. ,in This represents the total number of visual tokens in this layer. This indicates rounding down. The highest attention score is retained. The visual key-value pairs (including key vectors and value vectors) corresponding to each visual token are processed, while the visual key-value pairs corresponding to the remaining visual tokens are pruned (i.e., discarded). For example, assuming the total number of original visual tokens in a certain first attention layer is 1568, the compression ratio of this layer is... =0.8967 (corresponding to approximately 89.7%), then the number to be retained. = =1406, meaning 1406 key-value pairs corresponding to visual tokens are retained, and 162 are pruned. The retained key-value pairs are organized into a visual key-value pair cache for this layer. The size of this cache is much smaller than the original full cache, but it carries the most important visual semantic information of this layer. Subsequent attention calculations only need to use this compressed cache, thereby significantly reducing memory usage and computational complexity.

[0129] Furthermore, after obtaining the compressed visual key-value pair cache for each first attention layer, it can be sequentially concatenated with the system cache (key-value pairs corresponding to fixed system prompts) and the prompt word cache (key-value pairs corresponding to text query sequences) to form a complete key-value cache for subsequent autoregressive generation. It should be noted that the above pruning process only applies to visual key-value pairs; text key-value pairs (including the system cache and prompt word cache) remain intact to ensure the model's ability to follow instructions and its coherent understanding of text context. For example, for a model with 12 first attention layers, the above pruning operation is performed independently on each layer, resulting in 12 compressed visual key-value pair caches. These caches can be directly reused in each step of autoregressive generation, avoiding repeated computation of historical visual representations. Thus, the storage and computational overhead of the visual key-value cache can be significantly reduced without significantly affecting model accuracy.

[0130] By using the above methods, the size of the visual key-value cache can be compressed from linearly increasing with the sequence length to a constant range that is much smaller than the original size. This significantly reduces the memory usage and attention computation overhead during model inference, thus providing a feasible storage and computational foundation for efficiently completing attention computation and autoregressive generation of long sequences on the compressed small-scale cache.

[0131] In some implementations, the importance of each visual token can be quantified by obtaining its attention score at that layer. The number of tokens to be retained at the current layer can be calculated according to the compression ratio, and the corresponding number of visual tokens with the highest attention scores can be selected as target visual tokens. Simultaneously, the visual key-value pairs corresponding to the remaining tokens are pruned, thereby organizing the compressed target key-value pairs into a lightweight visual key-value pair cache. This allows the limited cache capacity to prioritize the visual semantic information that contributes most to the current generation task. For example, step 104 may include: (104.1) In each first attention layer, obtain the attention score corresponding to each visual token, wherein the first attention layer includes a compression boundary layer; (104.2) Obtain the total number of visual tokens corresponding to each first attention layer, and calculate the number of retained tokens corresponding to the current first attention layer based on the total number of visual tokens and the compression ratio; (104.3) Based on the size relationship between the attention scores of multiple visual tokens in the first attention layer, select the target visual token from the multiple visual tokens with the same number of tokens to be retained; (104.4) Based on the target visual token, the visual key-value pairs corresponding to the remaining visual tokens are pruned to obtain the pruning result; (104.5) Based on the pruning results and the target visual key-value pairs corresponding to the target visual token, the visual key-value pairs cache of each first attention layer is obtained.

[0132] The attention score can be the attention weight between the visual token and the current query in each first attention layer, or a quantitative indicator based on attention statistics, which can be used to measure the importance of each visual token to the current generation task.

[0133] The total number of visual tokens can be the total number of visual tokens generated after visual encoding of the image frame sequence in the current first attention layer. For example, if there are N input frames and M visual blocks are generated per frame, the total number is N×M.

[0134] The number of visual tokens to be retained can be the target number of visual tokens that need to be retained in the current first attention layer, calculated based on the compression ratio and the total number of visual tokens.

[0135] The target visual token can be the visual token selected in the current first attention layer according to the attention score from high to low, and the number of tokens is equal to the number to be retained. It can be used to indicate that its corresponding visual key-value pair is retained, while the key-value pairs corresponding to the other tokens are pruned.

[0136] Here, the target visual key-value pair can be a set of key-value pairs consisting of the key vector and value vector corresponding to each target visual token in the current first attention layer.

[0137] The pruning result can be a processing record obtained after discarding visual key-value pairs of visual tokens other than the target visual token (i.e., non-target visual tokens) in the current first attention layer. For example, an index list that identifies which key-value pairs have been removed or a direct vacancy status can be used to indicate that these redundant key-value pairs are no longer contained in the cache.

[0138] In some implementations, in each first attention layer (including the compression boundary layer, i.e., the Lth layer, where 1≤L≤h, and h is the compression boundary layer number), the attention score corresponding to each visual token in that layer is first obtained. The attention score can be obtained by the attention weight matrix calculated when the forward propagation is completed in the Prefill stage.

[0139] Specifically, for this attention layer, the multi-head self-attention mechanism calculates the similarity between each query token (typically the token to be generated or the input token) and all key tokens, and then normalizes it using softmax to obtain the attention weights. For visual tokens, their corresponding attention weights can be aggregated (e.g., the sum, maximum, or average of the attention weights of the visual token across all attention heads) as their attention score. For example, in layer 6 of the model, there are 1568 visual tokens. In the first step of the autoregression, the attention score for each visual token is calculated as the average of the attention weights of that token across all attention heads and all query positions. These scores reflect the importance of each visual token to the generation task in the current layer; a higher score indicates that the information carried by the token is more crucial.

[0140] Obtain the total number of visual tokens corresponding to each first attention layer. And based on the compression ratio of this layer Calculate the number of items retained corresponding to the current first attention layer. The number to be retained can be determined by the following formula: ; For example, assuming a compressed boundary layer h=12, for the 6th layer (L=6), the calculation is as follows: ≈0.8967; if the total number of visual tokens in this layer =1568, then =1406, meaning this layer needs to retain 1406 key-value pairs corresponding to visual tokens. This number will be used as the basis for selecting target visual tokens later.

[0141] Furthermore, based on the relationship between the attention scores of multiple visual tokens in the first attention layer, a number of target visual tokens, equal to the reserved number, can be selected from all visual tokens. Specifically, all visual tokens in this layer can be sorted from high to low according to their attention scores, and the top-ranked tokens can be selected. One visual token is selected as the target visual token. For example, following the previous example, among the 1568 visual tokens in layer 6, the visual tokens ranked 1st to 1406th in attention score are identified as target visual tokens, while the remaining 162 visual tokens are other visual tokens to be pruned. This selection strategy ensures that the most important visual information is preferentially retained in the cache, thereby maintaining high inference accuracy even after compression.

[0142] For example, based on the selected target visual token, pruning is performed on the visual key-value pairs corresponding to the remaining visual tokens in the current layer (i.e., visual tokens not selected as the target visual token), yielding the pruning result. The pruning operation may include removing the key and value vectors corresponding to these other visual tokens from the visual key-value cache of that layer, so they no longer participate in the subsequent attention calculation generated by autoregression. The pruning result can be an index list or mask indicating which key-value pairs corresponding to which visual tokens are retained and which are discarded. For example, for layer 6, the key-value pairs corresponding to the visual tokens ranked last 162 times in attention score are all removed, and the pruning result records the indices of these 162 tokens.

[0143] Furthermore, a visual key-value pair cache for each first attention layer can be constructed based on the pruning results and the target visual key-value pairs corresponding to the target visual tokens. Specifically, the key vector and value vector (i.e., the target visual key-value pair) corresponding to each target visual token can be extracted, and these vectors can be stored in the visual key-value pair cache of that layer in their original order or a reorganized order. The size of this cache is much smaller than the original full visual cache, and it contains only the most important visual information.

[0144] It should be noted that in actual deployment, the model's complete input cache consists of three parts: a system cache (corresponding to fixed system prompts), a prompt word cache (corresponding to the text query sequence), and a visual key-value pair cache (i.e., the compressed result in this application). These three are concatenated in sequence to form a complete key-value cache, which is used in each step of the subsequent autoregressive generation.

[0145] In this way, each first attention layer obtains a compact and information-dense visual key-value pair cache, which significantly reduces storage and computational overhead while maximizing the model's ability to perceive key visual content.

[0146] In some implementations, the similarity between the visual value vector to be processed for each token to be discarded and the first visual value vector for each token to be retained can be calculated. Based on this similarity, the semantic information of the value vector to be processed can be weighted and fused into the first visual value vector. This allows some information of the pruned tokens to be fed back into the retained cache structure, maximizing the compression of the cache size while maintaining the integrity of the overall semantic context. This avoids irreversible information loss caused by directly discarding the semantic information carried by the value vectors corresponding to unimportant visual tokens. For example, the target visual key-value pair may include a first visual key vector and a first visual value vector, and (104.5) may include: (104.5.1) Based on the pruning results, determine the visual value vector to be processed corresponding to each other visual token; (104.5.2) For each other visual token corresponding to the visual value vector to be processed, calculate the value similarity between it and the first visual value vector corresponding to each target visual token; (104.5.3) Based on the value similarity between each first visual value vector and each visual value vector to be processed, each visual value vector to be processed is fused into each first visual value vector to obtain the target first visual value vector corresponding to each target visual token; (104.5.4) In each first attention layer, the first visual key vector and the first visual value vector corresponding to each target visual token are cached to obtain the visual key-value pair cache of each first attention layer.

[0147] The first visual key vector can be the key vector corresponding to the target visual token in the current first attention layer. For example, it can be a floating-point vector with dimension d, which is directly retained after pruning without being merged and updated.

[0148] The first visual value vector can be the original value vector corresponding to the target visual token in the current first attention layer, such as a floating-point vector with dimension d, which is used to multiply with the attention weight in the attention calculation to obtain the output. This vector will be fused with information from other discarded tokens and updated to the target first visual value vector.

[0149] Other visual tokens can be all remaining visual tokens that were not selected as the target visual token in the current first attention layer. That is, the set of tokens with low attention scores that should be pruned according to the compression ratio. Their corresponding key-value pairs (especially value vectors) will be processed instead of being completely discarded.

[0150] The visual value vector to be processed can be the original value vector corresponding to each other visual token in the current first attention layer, such as the value vector of a visual token to be pruned. This vector will not be stored separately in the cache, but will be fused into the target first visual value vector of the target visual token as an information source.

[0151] Value similarity can be a similarity measure between a visual value vector to be processed and a first visual value vector. For example, it can be the cosine similarity or the similarity based on the inner product between the two. It is used to measure the similarity between the two value vectors in the semantic space and serves as a weight coefficient during weighted fusion, so that the discarded value vector that is more similar to the retained value vector contributes more information.

[0152] The target first visual value vector can be an updated value vector obtained by weighted fusion of the original first visual value vector of each target visual token with all visual value vectors to be processed. For example, it can be a vector obtained by weighted summation after similarity normalization. This vector incorporates some semantic information of the pruned tokens and is finally stored in the visual key-value pair cache for subsequent inference.

[0153] In some implementations, the visual value vector to be processed for each discarded visual token can be determined based on the pruning results (i.e., an index list indicating which visual tokens were discarded). Specifically, after the target visual tokens are selected, the key-value pairs corresponding to the remaining visual tokens are pruned. For these pruned tokens, this application does not directly discard their value vectors, but extracts them as visual value vectors to be processed for subsequent information fusion. For example, suppose there are 1406 target visual tokens and 162 other visual tokens in layer 6. For these 162 pruned tokens, their value vectors in the current attention layer are obtained, with each vector having a dimension of d (e.g., d=4096). These visual value vectors to be processed carry part of the semantic information of the pruned tokens and will be fed back to the retained cache through weighted fusion.

[0154] Furthermore, for each other visual token's corresponding visual value vector to be processed, the value similarity between it and the first visual value vector corresponding to each target visual token can be calculated to measure the similarity between the two value vectors in the semantic space. For example, cosine similarity can be used as a metric. Let a visual value vector to be processed be... The first visual value vector of a target visual token is Then the formula for calculating the value similarity s is: ; Understandably, to simplify calculations, all value vectors are usually pre-normalized (making the magnitude 1), at which point similarity can be simplified to a dot product. These similarity values ​​can be used to characterize the semantic association between the pruned tokens and each retained token; the higher the similarity, the closer their information is, and the greater the weight should be given during fusion.

[0155] Specifically, for the j-th target visual token, its target first visual value vector It can be calculated using the following formula: ; in, M is the original first visual value vector of the j-th target visual token; M is the total number of other visual tokens (pruned tokens); N is the total number of target visual tokens (i.e., the number retained). ); Let be the value similarity between the i-th visual value vector to be processed and the j-th first visual value vector; The normalized fusion weight represents the proportion of information from the i-th pruned value vector that is allocated to the j-th retained value vector.

[0156] Therefore, each pruned value vector distributes its semantic information proportionally to each retained value vector based on its similarity to each retained value vector. The retained value vector with higher similarity receives more fusion information, so that the semantic contribution of the pruned token can be preserved without occupying additional storage space in the cache.

[0157] Furthermore, in each first attention layer, the first visual key vector corresponding to each target visual token can be ( ) and the target first visual value vector calculated above ( The visual key-value pair cache of this layer is cached. Specifically, for each retained target visual token, its key vector remains unchanged, while its value vector is updated to a new value vector that incorporates the information of the pruned token. These key-value pairs are stored sequentially in the visual cache of this layer, while the system cache and prompt word cache remain unchanged. Finally, the three caches are concatenated and used for each step of subsequent autoregressive inference. Through this asymmetric compression strategy, the size of the visual cache can be significantly reduced while the semantic information of the pruned token can be preserved to the greatest extent through weighted fusion of value vectors, thus achieving a better balance between inference efficiency and model accuracy.

[0158] Step 105: Determine multiple second attention layers located after the compressed boundary layer from multiple attention layers, and in each second attention layer, prune all visual key-value pairs corresponding to each visual token.

[0159] In some implementations, in order to completely eliminate the redundant storage and computation burden caused by the highly abstract representation of visual information in deep networks, and to make full use of the characteristics that visual information has converged to global semantics and is insensitive to local details in deep attention layers, all attention layers located after the compressed boundary layer can be identified as second attention layers, and full pruning (i.e. complete removal) can be performed directly on all visual key-value pairs corresponding to visual tokens in each second attention layer to reduce the size of deep visual cache to the greatest extent.

[0160] The second attention layer can be a collective term for all attention layers in a visual large language model that are located after the compression boundary layer (excluding the compression boundary layer itself). For example, in a model with a total of 32 layers, if the compression boundary layer is the 12th layer, then layers 13 to 32 are all second attention layers. The visual representations in these deep networks are highly abstract and globalized, and are not sensitive to fine-grained details of visual tokens. They can be used to adopt an aggressive full pruning strategy, that is, to directly discard all visual key-value pairs corresponding to visual tokens in this layer and no longer retain any visual cache, so as to maximize the compression effect.

[0161] Understandably, the second attention layer is located in a deep region of the network, where its visual representation has undergone multiple non-linear transformations, abstracting from fine-grained features such as shallow edges and textures to global coarse-grained information such as objects, scenes, and semantic relationships. At this stage, the contribution of individual details of visual tokens to the final generation task is significantly reduced, and the model relies more on overall semantic understanding. Therefore, these second attention layers are suitable for an aggressive compression strategy, i.e., completely pruning all visual key-value pairs, to maximize memory savings and computational acceleration without significantly affecting inference accuracy.

[0162] Furthermore, within each second attention layer, full pruning can be performed on all visual key-value pairs (including visual key vectors and visual value vectors) corresponding to all visual tokens in that layer. This means discarding all visual key-value pairs and eliminating any visual cache. Specifically, for each layer in the second attention layer, the visual cache of that layer is directly cleared. During subsequent autoregressive generation, the attention computation of that layer will only rely on the system cache (corresponding to fixed system prompts) and the prompt word cache (corresponding to the text query sequence), no longer involving any visual tokens. After performing full pruning, all these key-value pairs are removed, significantly reducing the memory usage and attention computation complexity of that layer.

[0163] It's important to note that since the second attention layer doesn't use visual caching at all, the model's system cache and prompt word cache are fully preserved to ensure the model's ability to follow instructions and its coherent understanding of text context. By progressively compressing and fusing key visual information in the first attention layer, and completely pruning the visual cache in the second attention layer, the model achieves extreme compression of visual key-value cache while maintaining overall inference accuracy, significantly improving the efficiency of long-context inference.

[0164] By employing the above methods, the model can avoid performing attention operations on visual tokens during deep computation, significantly reducing the computational load and memory usage in deep networks. This, combined with the progressive compression in the first attention layer, forms a layered collaborative mechanism that preserves fine details in the shallow layer and completely prunes in the deep layer. This maximizes the overall inference speed and minimizes memory overhead while ensuring the model's ability to perceive key visual information.

[0165] In some implementations, to completely eliminate the storage and computational redundancy caused by visual key-value caching in deep networks, the second visual key vector and second visual value vector corresponding to all visual tokens in each second attention layer can be obtained and full pruning (i.e., all discarded) can be performed to achieve zero caching of the deep network's visual cache, thereby minimizing the memory overhead and attention computation complexity of the deep autoregressive stage. For example, step 105, "pruning all visual key-value pairs corresponding to each visual token in each second attention layer," can include: (105.1) Obtain the second visual key vector and the second visual value vector corresponding to all visual tokens in each second attention layer; (105.2) Prune the second visual key vector and the second visual value vector corresponding to all visual tokens in each second attention layer.

[0166] The second visual key vector can be a key vector obtained by linearly transforming the hidden representation of each visual token in the second attention layer.

[0167] The second visual value vector can be a value vector obtained by linearly transforming the hidden representation of each visual token in the second attention layer.

[0168] For example, after determining the compression boundary layer (e.g., layer 24), for each subsequent second attention layer (e.g., layers 25 to 32), the second visual key vectors and second visual value vectors corresponding to all visual tokens in that layer are first obtained. These vectors are obtained by mapping the hidden representation of each visual token through independent linear transformation matrices during the forward propagation in the Prefill phase and are temporarily stored in the visual buffer of that layer. For example, assuming a second attention layer contains K = 1568 visual tokens, the set of second visual key vectors is... The second visual value vector set is Each vector has the same dimension as the model's hidden dimension (e.g., 4096 dimensions). These vectors should have participated in the attention calculation at each time step in the subsequent autoregressive generation, but because the second attention layer is located deep in the network, its visual representation has undergone multiple nonlinear transformations, abstracting from shallow fine-grained features such as edges and textures to global semantic information, and the dependence on the individual details of each visual token has been greatly reduced.

[0169] Furthermore, a complete pruning process is performed on the second visual key vector and second visual value vector corresponding to all visual tokens in each second attention layer. This means that all visual key-value pairs are discarded directly, and the visual cache of that layer is cleared. Specifically, all visual key vectors and value vectors can be removed from the storage structure, releasing the corresponding GPU memory space, and the cache index of that layer is updated. This ensures that in each subsequent autoregressive generation step, the attention calculation of that layer is based solely on the system cache (corresponding to fixed system prompts) and the prompt word cache (corresponding to the text query sequence), without involving any visual tokens.

[0170] By using the above methods, the storage and computational redundancy of visual key-value caching can be completely eliminated in deep networks, significantly accelerating autoregressive generation and reducing peak memory usage with almost no loss of overall model inference accuracy.

[0171] Step 106: Based on the text key-value pairs corresponding to each attention layer, determine the first text key-value pairs corresponding to each first attention layer and the second text key-value pairs corresponding to each second attention layer, and perform inference sequentially based on the first text key-value pairs and visual key-value pairs cached for each first attention layer and the second text key-value pairs corresponding to each second attention layer to generate the inference results corresponding to the text query sequence.

[0172] In some implementations, the first text key-value pairs and the second text key-value pairs corresponding to the first attention layer and the second attention layer can be extracted from the complete text key-value pairs corresponding to each attention layer, and the compressed visual key-value pairs can be cached and used only in conjunction with the first text key-value pairs of the first attention layer for attention calculation (the second attention layer only uses text key-value pairs). In this way, the autoregressive forward propagation of each layer can be completed in a hierarchical fusion manner, and finally semantically accurate and reasoning-efficient text query results can be generated.

[0173] The first text key-value pair can be a set of key vectors and value vectors obtained by linear transformation of all text tokens corresponding to the text query sequence in each first attention layer. This key-value pair is completely preserved during compression without any pruning and can be used in the first attention layer to perform cross-modal attention calculation together with the compressed visual key-value pair cache to fuse text and visual semantics.

[0174] The second text key-value pair can be the set of key vectors and value vectors obtained by linear transformation of all text tokens corresponding to the text query sequence in each second attention layer. Since all visual key-value pairs in the second attention layer have been pruned, the text key-value pair will be used separately for self-attention or attention calculation within the text, and no longer involves visual information. It can be used to continue to advance the autoregressive generation of the text sequence in the deep network.

[0175] The inference result can be the final output sequence generated by the visual big language model through token-by-token autoregression, with the participation of compressed visual key-value pair cache and complete text key-value pairs. For example, the descriptive text "a person is playing basketball" output for video understanding task or the answer "yes" output for visual question answering task can be used to respond to user text queries and provide accurate semantic answers.

[0176] In some implementations, the first text key-value pair may include two parts: a system cache (corresponding to a fixed system cue word, such as "You are a helpful assistant.") and a cue word cache (corresponding to a text query sequence). Regardless of the attention layer, these two parts of the text key-value pair remain intact and do not participate in any pruning operations. Therefore, for each first attention layer (e.g., layers 1 to 24), its first text key-value pair is the union of the complete system cache and the cue word cache for that layer. For example, in a 32-layer model, each layer from layer 1 to layer 24 has the same set of text key-value pairs (assuming a total of 50 text tokens, each layer's text key-value pair contains 50 key vectors and 50 value vectors). These first text key-value pairs will be used in conjunction with the compressed visual key-value pair cache of this layer to achieve cross-modal attention computation.

[0177] Furthermore, for each second attention layer (e.g., layers 25 to 32), its corresponding second text key-value pairs are also the union of the complete system cache and the prompt word cache for that layer, completely identical in content to the first text key-value pairs. The difference is that in the second attention layer, the visual key-value pair cache has been completely pruned (i.e., all visual key-value pairs are discarded). Therefore, the second attention layer only retains the second text key-value pairs and no longer contains any visual information. For example, in layer 28, the text key-value pairs still contain the keys and values ​​of 50 text tokens, while the visual cache is empty.

[0178] Furthermore, in the autoregressive generation process, forward propagation can be performed sequentially through each attention layer. For each first attention layer, the first text key-value pair of this layer is concatenated sequentially with the compressed and fused visual key-value pair cache to form a complete key-value cache. This cache is then used for attention calculation with the query vector of the current step, outputting the updated hidden representation. For each second attention layer, since the visual key-value pair cache is empty, attention calculation is performed only using the second text key-value pairs, i.e., self-attention is only performed between text tokens. This process is repeated layer by layer until the calculation of the last attention layer is completed. The probability distribution of the next token is obtained through the output layer, a new token is sampled and generated, and it is appended to the input sequence. The above process is repeated to generate subsequent tokens sequentially until an end symbol is encountered or the preset maximum length is reached. The final output complete token sequence is the inference result corresponding to the text query sequence. For example, for a video question answering task, the input text query is "What is the person in the video doing?". After the above layered inference, the model generates the inference result "A person is playing basketball". In this way, while maintaining the complete understanding of text semantics, redundant calculation of visual tokens in deep networks can be significantly reduced, achieving significant inference acceleration and memory saving.

[0179] This application embodiment obtains a text query sequence and its corresponding image frame sequence, and through each attention layer in a pre-trained visual large language model, calculates the text key-value pair corresponding to each text token in the text query sequence and the visual key-value pair corresponding to each visual token in the image frame sequence based on the text query sequence and the image frame sequence; based on the semantic extraction granularity of each attention layer for the image frame sequence, it determines a compression boundary layer from the multiple attention layers included in the visual large language model; it determines multiple first attention layers located before the compression boundary layer from the multiple attention layers, and calculates the compression ratio corresponding to each first attention layer based on the inter-layer distance relationship between each first attention layer and the compression boundary layer; in each In each first attention layer, visual key-value pairs corresponding to each visual token are pruned according to a compression ratio to obtain a visual key-value pair cache. Multiple second attention layers are determined from the multiple attention layers, following the compression boundary layer. In each second attention layer, all visual key-value pairs corresponding to each visual token are pruned. Based on the text key-value pairs corresponding to each attention layer, the first text key-value pairs corresponding to each first attention layer and the second text key-value pairs corresponding to each second attention layer are determined. Reasoning is then performed sequentially based on the first text key-value pairs and visual key-value pair caches corresponding to each first attention layer, as well as the second text key-value pairs corresponding to each second attention layer, to generate the reasoning result corresponding to the text query sequence. This allows for differentiated layered compression strategies for different network layers. Specifically, since shallow networks typically focus on local details and deep networks focus on semantic information, this application allocates non-uniform compression ratios to different layers based on inter-layer distance relationships. This allows layers with dense key information to retain more visual key-value pairs, while layers with more redundancy undergo more aggressive compression, thus avoiding the problems of over-pruning of key information or excessive retention of redundancy caused by using the same fixed pruning ratio. In summary, this application can significantly improve the model's inference efficiency while maintaining the original inference accuracy.

[0180] Please see Figure 3 In some implementations, combined with Figure 3This paper introduces the overall process of this application. The model inference based on key-value cache compression can be divided into two collaborative stages: coarse-grained keyframe filtering and fine-grained key-value cache compression. First, in the coarse-grained keyframe filtering stage, the system receives a text query sequence (such as "What is the main purpose of that glass of water in this video?") and several corresponding initial image frames. A small visual language model (e.g., a CLIP encoder) encodes the text and image frames separately, and based on the first similarity between the text feature vector and the image feature vector, a maximum marginal relevance (MMR) strategy is used to filter out semantically relevant and low-redundancy keyframes, constructing a compact image frame sequence. This stage reduces the number of visual tokens entering the main model from the source, significantly reducing the computational burden of subsequent processing.

[0181] Furthermore, the selected keyframes and original prompts can be input together into the visual large language model to perform prefill forward propagation. During the calculation of each attention layer in the model, the system automatically saves the attention score corresponding to each visual token. These scores reflect the importance of different visual tokens to the current generation task. Then, these attention scores are used to guide the KV pruning operation: based on a pre-determined hierarchical compression boundary (dividing the model layers into a first attention layer and a second attention layer), the first attention layer retains the key-value pairs of the target visual tokens with the highest attention scores according to a parabolic compression ratio, and performs similarity-weighted asymmetric fusion on the value vectors of the pruned tokens; for the second attention layer, all visual key-value pairs are directly pruned. Finally, a compressed key-value cache is obtained.

[0182] Furthermore, this compressed key-value cache, along with the complete text key-value pairs (including system cache and prompt word cache), is used for subsequent autoregressive generative inference. At each time step, the model sequentially calculates attention through a first attention layer (utilizing both text key-value pairs and compressed visual key-value pairs) and a second attention layer (utilizing only text key-value pairs) and outputs the next token, until a complete inference result is generated. The entire process achieves efficient inference from visual input to the final answer, significantly reducing memory usage and computational latency while maintaining accuracy.

[0183] Please see Figure 4 This application also provides a model inference apparatus based on key-value cache compression, which can implement the above-mentioned model inference method based on key-value cache compression. The model inference apparatus based on key-value cache compression includes: The acquisition module 41 is used to acquire the text query sequence and the corresponding image frame sequence, and through each attention layer in the pre-trained visual big language model, calculate the text key-value pair corresponding to each text token in the text query sequence and the visual key-value pair corresponding to each visual token in the image frame sequence based on the text query sequence and the image frame sequence. The first determining module 42 is used to determine the compression boundary layer from the multiple attention layers contained in the visual large language model based on the semantic extraction granularity of each attention layer for the image frame sequence. The calculation module 43 is used to determine multiple first attention layers located before the compression boundary layer from multiple attention layers, and calculate the compression ratio corresponding to each first attention layer based on the interlayer distance relationship between each first attention layer and the compression boundary layer. Pruning module 44 is used to prune the visual key-value pairs corresponding to each visual token according to the compression ratio in each first attention layer to obtain a visual key-value pair cache. The second determining module 45 is used to determine multiple second attention layers located after the compressed boundary layer from multiple attention layers, and in each second attention layer, to prune all visual key-value pairs corresponding to each visual token. The generation module 46 is used to determine the first text key-value pair corresponding to each first attention layer and the second text key-value pair corresponding to each second attention layer based on the text key-value pair corresponding to each attention layer, and to perform inference based on the first text key-value pair and visual key-value pair cache corresponding to each first attention layer and the second text key-value pair corresponding to each second attention layer in sequence, and to generate the inference result corresponding to the text query sequence.

[0184] The specific implementation of the model inference device based on key-value cache compression is basically the same as the specific embodiment of the model inference method based on key-value cache compression described above, and will not be repeated here. Subject to meeting the requirements of the embodiments of this application, the model inference device based on key-value cache compression may also be equipped with other functional modules to implement the model inference method based on key-value cache compression in the above embodiments.

[0185] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned model inference method based on key-value cache compression. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0186] Please see Figure 5 , Figure 5 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes: The processor 51 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 52 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 52 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 52 and called and executed by the processor 51 to execute the model inference method based on key-value cache compression of the embodiments of this application. Input / output interface 53 is used to implement information input and output; The communication interface 54 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 55 transmits information between various components of the device (e.g., processor 51, memory 52, input / output interface 53, and communication interface 54); The processor 51, memory 52, input / output interface 53, and communication interface 54 are connected to each other within the device via bus 55.

[0187] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described model inference method based on key-value cache compression.

[0188] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0189] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0190] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0191] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0192] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0193] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0194] It should be understood that in this application, "at least one" and "several" refer to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0195] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0196] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0197] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0198] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0199] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A model inference method based on key-value cache compression, characterized in that, The method includes: Obtain the text query sequence and the corresponding image frame sequence, and through each attention layer in the pre-trained visual large language model, calculate the text key-value pair corresponding to each text token of the text query sequence and the visual key-value pair corresponding to each visual token of the image frame sequence based on the text query sequence and the image frame sequence. The step of obtaining the text query sequence and the corresponding image frame sequence includes: obtaining the text query sequence and multiple initial image frames; encoding the text query sequence using a preset lightweight visual language model to obtain a text feature vector, and encoding each initial image frame to obtain an image feature vector; selecting multiple target image frames from the multiple initial image frames based on a first similarity between the text feature vector and each image feature vector; and constructing an image frame sequence based on the multiple target image frames. The step of selecting multiple target image frames from the multiple initial image frames based on the first similarity between the text feature vector and each image feature vector includes: determining the initial image frame corresponding to the image feature vector with the highest first similarity to the text feature vector as the first target image frame, and adding the first target image frame to the image frame set; calculating a second similarity between the image feature vector corresponding to each initial image frame and the target image feature vector of each target image frame in the image frame set, and selecting the maximum similarity among the second similarities with each target image frame in the image frame set; for each initial image frame, calculating a first weighted value based on the product of a preset first balance coefficient and the first similarity, and calculating a weighted value based on the product of a preset second balance coefficient and the first similarity. The product of the maximum similarity values ​​is used to calculate a second weighted value. Based on the difference between the first weighted value and the second weighted value, a comprehensive score is calculated for each initial image frame, wherein the first balance coefficient is greater than the second balance coefficient. Based on the comprehensive score corresponding to each initial image frame, the next target image frame is selected from the plurality of initial image frames and added to the image frame set. Based on the next target image frame, the next target image frame is determined from the remaining initial image frames and added to the image frame set. The step of determining the next target image frame from the remaining initial image frames and adding it to the image frame set based on the next target image frame is repeated until the total number of target image frames in the image frame set reaches a preset threshold, resulting in a plurality of target image frames. Based on the semantic extraction granularity of each attention layer for the image frame sequence, a compression boundary layer is determined from the multiple attention layers contained in the visual large language model; The compression boundary layer and the attention layer preceding the compression boundary layer are selected from the plurality of attention layers to form a plurality of first attention layers. Based on the interlayer distance relationship between each first attention layer and the compression boundary layer, the compression ratio corresponding to each first attention layer is calculated. In each first attention layer, the visual key-value pairs corresponding to each visual token are pruned according to the compression ratio to obtain a visual key-value pair cache. From the plurality of attention layers, determine a plurality of second attention layers located after the compressed boundary layer, and in each second attention layer, prune all visual key-value pairs corresponding to each visual token; Based on the text key-value pairs corresponding to each attention layer, the first text key-value pairs corresponding to each first attention layer and the second text key-value pairs corresponding to each second attention layer are determined. Then, reasoning is performed sequentially based on the first text key-value pairs and the visual key-value pairs cached for each first attention layer, and the second text key-value pairs corresponding to each second attention layer, to generate the reasoning result corresponding to the text query sequence.

2. The model inference method based on key-value cache compression according to claim 1, characterized in that, The calculation of the compression ratio corresponding to each first attention layer based on the interlayer distance relationship between each first attention layer and the compression boundary layer includes: Determine the first layer number corresponding to the compressed boundary layer, and the second layer number corresponding to each first attention layer, wherein the type of the first attention layer includes the compressed boundary layer; Based on the difference between the first layer number corresponding to the compressed boundary layer and the preset reference value, the first offset is determined, and the target first offset is obtained based on the square of the first offset. For each first attention layer, a second offset is determined based on the difference between the second layer number and the preset reference value, and a target second offset is obtained based on the square of the second offset. Based on the first target offset and the second target offset corresponding to each first attention layer, intermediate parameters are obtained; Based on the difference between the preset reference value and the intermediate parameter corresponding to each first attention layer, the compression ratio corresponding to each first attention layer is calculated.

3. The method of claim 1, wherein, In each first attention layer, the visual key-value pairs corresponding to each visual token are pruned according to the compression ratio to obtain a visual key-value pair cache, including: In each first attention layer, the attention score corresponding to each visual token is obtained, wherein the first attention layer includes a compression boundary layer; Obtain the total number of visual tokens corresponding to each first attention layer, and calculate the number of retained tokens corresponding to the current first attention layer based on the total number of visual tokens and the compression ratio. Based on the size relationship between the attention scores of multiple visual tokens corresponding to the first attention layer, a target visual token is selected from the multiple visual tokens with the same number of tokens to be retained; Based on the target visual token, the visual key-value pairs corresponding to the remaining visual tokens are pruned to obtain the pruning result. Based on the pruning results and the target visual key-value pairs corresponding to the target visual token, a visual key-value pair cache for each first attention layer is obtained.

4. The model inference method based on key-value cache compression according to claim 3, characterized in that, The target visual key-value pair includes a first visual key vector and a first visual value vector. The step of obtaining a visual key-value pair cache for each first attention layer based on the pruning result and the target visual key-value pair corresponding to the target visual token includes: Based on the pruning results, determine the visual value vector to be processed for each other visual token; For each other visual token corresponding to the visual value vector to be processed, calculate the value similarity between it and the first visual value vector corresponding to each target visual token. Based on the value similarity between each first visual value vector and each visual value vector to be processed, each visual value vector to be processed is fused into each first visual value vector to obtain the target first visual value vector corresponding to each target visual token. In each first attention layer, the first visual key vector and the first visual value vector corresponding to each target visual token are cached to obtain the visual key-value pair cache of each first attention layer.

5. The method of claim 1, wherein, In each second attention layer, pruning is performed on all visual key-value pairs corresponding to each visual token, including: Obtain the second visual key vector and second visual value vector corresponding to all visual tokens in each second attention layer; Prune the second visual key vector and the second visual value vector corresponding to all visual tokens in each second attention layer.

6. An apparatus for model inference based on key-value cache compression, the apparatus comprising: The device includes: An acquisition module is used to acquire a text query sequence and its corresponding image frame sequence, and through each attention layer in a pre-trained visual large language model, calculate the text key-value pair corresponding to each text token in the text query sequence and the visual key-value pair corresponding to each visual token in the image frame sequence based on the text query sequence and the image frame sequence; wherein, acquiring the text query sequence and its corresponding image frame sequence includes: acquiring the text query sequence and its corresponding multiple initial image frames; encoding the text query sequence to obtain a text feature vector and encoding each initial image frame to obtain an image feature vector through a preset lightweight visual language model; selecting multiple target image frames from the multiple initial image frames based on a first similarity between the text feature vector and each image feature vector; constructing an image frame sequence based on the multiple target image frames; wherein, selecting multiple target image frames from the multiple initial image frames based on the first similarity between the text feature vector and each image feature vector includes: determining the initial image frame corresponding to the image feature vector with the highest first similarity to the text feature vector as the first target image frame, and adding the first target image frame to... An image frame set is used. For each initial image frame, a second similarity is calculated between its corresponding image feature vector and the target image feature vector of each target image frame in the image frame set. The maximum similarity among the second similarities is selected. For each initial image frame, a first weighted value is calculated based on the product of a preset first balance coefficient and the first similarity. A second weighted value is calculated based on the product of a preset second balance coefficient and the maximum similarity. A comprehensive score is calculated based on the difference between the first weighted value and the second weighted value, wherein the first balance coefficient is greater than the second balance coefficient. Based on the comprehensive score corresponding to each initial image frame, a next target image frame is selected from the plurality of initial image frames and added to the image frame set. Based on the next target image frame, a next target image frame is determined from the remaining initial image frames and added to the image frame set. The step of determining the next target image frame from the remaining initial image frames and adding it to the image frame set based on the next target image frame is repeated until the total number of target image frames in the image frame set reaches a preset threshold, resulting in a plurality of target image frames. The first determining module is used to determine the compression boundary layer from the multiple attention layers contained in the visual large language model based on the semantic extraction granularity of each attention layer for the image frame sequence. The calculation module is used to select the compression boundary layer and the attention layer preceding the compression boundary layer from the plurality of attention layers to determine a plurality of first attention layers, and calculate the compression ratio corresponding to each first attention layer based on the interlayer distance relationship between each first attention layer and the compression boundary layer. The pruning module is used to prune the visual key-value pairs corresponding to each visual token in each first attention layer according to the compression ratio, so as to obtain a visual key-value pair cache. The second determining module is used to determine a plurality of second attention layers located after the compressed boundary layer from the plurality of attention layers, and in each second attention layer, to prune all visual key-value pairs corresponding to each visual token; The generation module is used to determine the first text key-value pair corresponding to each first attention layer and the second text key-value pair corresponding to each second attention layer based on the text key-value pair corresponding to each attention layer, and to perform inference sequentially based on the first text key-value pair and the visual key-value pair cache corresponding to each first attention layer and the second text key-value pair corresponding to each second attention layer to generate the inference result corresponding to the text query sequence.

7. A computer device, comprising: The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the model inference method based on key-value cache compression as described in any one of claims 1 to 5.

8. A computer readable storage medium, the storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by the processor, it implements the model inference method based on key-value cache compression as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Large model reasoning method and device, data processing method and device, equipment and storage medium

    CN121279425A

  • A method and architecture for an interactive two-way data communication network

    EP0779759A2