Visual language large model understanding method and device, computer device and storage medium
By allocating differentiated dynamic compression rates and target temporal index sets in a large visual language model, the cache of long videos is compressed, which solves the problem of redundant visual features occupying context space and improves the effect of long video understanding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG LAB
- Filing Date
- 2026-04-15
- Publication Date
- 2026-08-04
AI Technical Summary
Existing large-scale visual language models suffer from insufficient understanding when processing long videos due to redundant visual features crowding out the context space.
By acquiring long videos and text prompts, and utilizing a pre-trained visual language large model, we determine the splicing features and multi-head attention score matrix. Based on the maximum context window length and global compression rate, we assign differentiated dynamic compression rates to each layer of the attention network, and compress the cache through a target temporal index set.
It improves the performance of large visual language models in long video understanding, reduces the interference of redundant features on decision-making, and improves the model's understanding accuracy and efficiency.
Smart Images

Figure CN122049784B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to methods, apparatus, computer devices, and storage media for understanding large visual language models. Background Technology
[0002] Visual language large models, as a multimodal technology integrating perception and understanding capabilities, have shown broad application prospects in fields such as video surveillance, human-computer interaction, and autonomous driving. However, due to limitations in the length of the model context window, existing visual language large models still face significant challenges when processing long videos. Therefore, an efficient visual language large model understanding technique will directly affect the understanding performance of visual language large models.
[0003] Existing large-scale visual language understanding techniques typically employ compression methods based on fixed window truncation or keyframe sampling. Their drawback is that a large number of redundant or low-quality visual features encroach on the limited context space, resulting in insufficient understanding performance of the large-scale visual language model.
[0004] There is currently no effective solution to the problem that redundant visual features in related technologies crowd out the context space, resulting in insufficient understanding of large visual language models. Summary of the Invention
[0005] This embodiment provides a method, apparatus, computer device, and storage medium for understanding large visual language models, in order to solve the problem in related technologies where redundant visual features crowd out the context space, resulting in insufficient understanding of large visual language models.
[0006] Firstly, this embodiment provides a method for understanding large-scale visual language models, including:
[0007] Obtain the long video and the text prompt corresponding to the long video;
[0008] The long video and the text prompt are input into a pre-trained visual language large model; the visual language large model includes a large language model;
[0009] Based on the video features corresponding to the long video and the text features corresponding to the text prompt, the splicing features are determined; and the splicing features are input into the large language model.
[0010] Based on the splicing features, determine the multi-head attention score matrix of each layer of the attention network in the large language model; determine the global compression rate based on the maximum context window length of the large language model, the text features, and the video features;
[0011] Based on the multi-head attention score matrix and the global compression ratio, the response strength score of each layer of the attention network is determined;
[0012] Based on the response intensity score, a differentiated dynamic compression ratio is assigned to each layer of the attention network;
[0013] Based on the multi-head attention score matrix, a target temporal index set is determined; and based on the target temporal index set and the dynamic compression ratio, the cache of each layer of the attention network is compressed.
[0014] Based on the aforementioned visual language model, a response is generated for the text prompt.
[0015] In some embodiments, determining the splicing features based on the video features corresponding to the long video and the text features corresponding to the text prompt includes:
[0016] The visual encoder of the large visual language model extracts the corresponding video features from the long video; and assigns temporal indexes to the video features in chronological order.
[0017] The text encoder of the visual language big data model converts the text prompt into the corresponding text features.
[0018] The video features and the text features are concatenated to obtain the concatenated features.
[0019] In some embodiments, the total number of tokens in the splicing feature is less than or equal to the maximum context window length.
[0020] In some embodiments, assigning differentiated dynamic compression ratios to each layer of the attention network based on the response intensity score includes:
[0021] Based on the response intensity scores, the attention networks in each layer are grouped to obtain the grouping results;
[0022] Based on the grouping results, the dynamic compression ratio of the attention network in each layer is determined.
[0023] In some embodiments, determining the target temporal index set based on the multi-head attention score matrix includes:
[0024] Based on the multi-head attention score matrix, determine the intensity mean ranking result, intensity variance ranking result, and intensity frequency ranking result;
[0025] The target time series index set is determined based on the intensity mean sorting result, the intensity variance sorting result, and the intensity frequency sorting result.
[0026] In some embodiments, determining the target time-series index set based on the intensity mean sorting result, the intensity variance sorting result, and the intensity frequency sorting result includes:
[0027] The inverse mixed sort score is determined based on the sorting results of the intensity mean, the intensity variance, and the intensity frequency.
[0028] The target time-series index set is selected based on the inverse mixed sort score.
[0029] In some embodiments, the expression for the reciprocal mixed sort score is:
[0030] ;
[0031] In the formula, The smoothing weighting coefficients set for experience, The sorting results of the intensity mean The ranking results of the intensity variance, The intensity frequencies are sorted.
[0032] Secondly, this embodiment provides a visual language large model understanding device, including: an acquisition module, a processing module, and a compression module;
[0033] The acquisition module is used to acquire the long video and the text prompts corresponding to the long video;
[0034] The processing module is used to input the long video and the text prompt into a pre-trained visual language large model; the visual language large model includes a large language model; determine splicing features based on the video features corresponding to the long video and the text features corresponding to the text prompt; input the splicing features into the large language model; determine the multi-head attention score matrix of each layer of the attention network in the large language model based on the splicing features; determine the global compression rate based on the maximum context window length of the large language model, the text features, and the video features; and determine the response intensity score of each layer of the attention network based on the multi-head attention score matrix and the global compression rate.
[0035] The compression module is used to assign differentiated dynamic compression rates to each layer of the attention network based on the response intensity score; determine a target temporal index set based on the multi-head attention score matrix; compress the cache of each layer of the attention network based on the target temporal index set and the dynamic compression rate; and generate an answer to the text prompt based on the visual language big model.
[0036] Thirdly, this embodiment provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the visual language large model understanding method described in the first aspect above.
[0037] Fourthly, this embodiment provides a readable storage medium storing a computer program that, when executed by a processor, implements the visual language large model understanding method described in the first aspect above.
[0038] Compared with related technologies, the visual language large-scale model understanding method, apparatus, computer device, and storage medium provided in this embodiment acquire a long video and corresponding text prompts; input the long video and text prompts into a pre-trained visual language large-scale model; the visual language large-scale model includes a large language model; determine splicing features based on the video features corresponding to the long video and the text features corresponding to the text prompts; input the splicing features into the large language model; determine the multi-head attention score matrix of each layer of attention network in the large language model based on the splicing features; determine the global compression ratio based on the maximum context window length, text features, and video features of the large language model; determine the response intensity score of each layer of attention network based on the multi-head attention score matrix and the global compression ratio; assign differentiated dynamic compression ratios to each layer of attention network based on the response intensity scores; determine the target temporal index set based on the multi-head attention score matrix; and compress the cache of each layer of attention network based on the target temporal index set and the dynamic compression ratio; and generate a response to the text prompts based on the visual language large-scale model. By assigning differentiated dynamic compression rates to different layers based on the differences in response intensity of each layer of the attention network, the problem of redundant visual features crowding out the context space and causing insufficient understanding of large visual language models in related technologies is solved, thereby improving the understanding effect of large visual language models.
[0039] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description
[0040] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0041] Figure 1 This is a hardware structure block diagram of a terminal device for a visual language large model understanding method provided in an embodiment of this application;
[0042] Figure 2This is a flowchart of a visual language large-scale model understanding method provided in an embodiment of this application;
[0043] Figure 3 This is a schematic diagram of long video capture provided in an embodiment of this application;
[0044] Figure 4 This is a schematic diagram of dynamic compression ratio allocation provided in an embodiment of this application;
[0045] Figure 5 This is a schematic diagram of the structure of a large visual language model provided in an embodiment of this application;
[0046] Figure 6 This is a schematic diagram of an embodiment of the understanding task provided in this application;
[0047] Figure 7 This is a flowchart illustrating a visual language large-scale model understanding method provided in an embodiment of this application;
[0048] Figure 8 This is a structural block diagram of a visual language large-scale model understanding device provided in an embodiment of this application.
[0049] In the diagram: 102, processor; 104, memory; 106, transmission device; 108, input / output device; 210, acquisition module; 220, processing module; 230, compression module. Detailed Implementation
[0050] To better understand the purpose, technical solution, and advantages of this application, the application is described and explained below in conjunction with the accompanying drawings and embodiments.
[0051] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific order of objects.
[0052] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. For example, it can run on a terminal. Figure 1 This is a hardware structure block diagram of the terminal for the visual language large-scale model understanding method in this embodiment. For example... Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 and a memory 104 for storing data are also included. The processor 102 may be, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA). The terminal may also include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that… Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown are illustrated.
[0053] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the visual language large model understanding method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0054] The transmission device 106 is used to receive or send data via a network. This network includes a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 can be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0055] This embodiment provides a method for understanding large-scale visual language models. Figure 2 This is a flowchart of the visual language large-scale model understanding method in this embodiment, such as... Figure 2 As shown, the process includes the following steps:
[0056] Step S210: Obtain the long video and the text prompts corresponding to the long video.
[0057] Specifically, long videos can be acquired by connecting to a video capture device to obtain a real-time video stream; they can also be obtained by reading pre-recorded historical video files from local or cloud storage devices; or they can be retrieved from a streaming media server via a network interface. There are no restrictions on the method of acquiring long videos. Text prompts corresponding to long videos can be obtained through direct user input; they can also be automatically generated through predefined video understanding tasks; or they can be obtained from external systems via API calls. There are no restrictions on the method of obtaining text prompts.
[0058] Long videos consist of multiple consecutive RGB images, and issues such as image jitter and blurriness should be avoided as much as possible during the acquisition process. Text prompts are the starting signal for the video understanding task and should be semantically complete and unambiguous. The acquisition time of text prompts should be after the acquisition time of the long video.
[0059] Step S220: Input the long video and text prompts into the pre-trained visual language large model; the visual language large model includes a large language model; determine the splicing features based on the video features corresponding to the long video and the text features corresponding to the text prompts; input the splicing features into the large language model; determine the multi-head attention score matrix of each layer of attention network in the large language model based on the splicing features; determine the global compression rate based on the maximum context window length, text features, and video features of the large language model; determine the response intensity score of each layer of attention network based on the multi-head attention score matrix and the global compression rate.
[0060] Specifically, text features are concatenated with video features prioritized by time sequence. That is, video features are split into multiple segments according to the time order of video frames, and the segments with earlier time sequences are concatenated with text features first. The concatenated features are then input into a large language model. During the pre-filling stage, the multi-head attention score matrix between text features and video features in each layer of the attention network is calculated sequentially. The global compression rate of video features is calculated based on the maximum context window length of the large language model, the length of the input text features, and the total length of all video features. Then, based on the multi-head attention score matrix and the global compression rate, the response intensity score of each layer of the attention network is calculated.
[0061] Step S230: Assign differentiated dynamic compression rates to each layer of the attention network based on the response intensity score; determine the target temporal index set based on the multi-head attention score matrix; compress the cache of each layer of the attention network based on the target temporal index set and the dynamic compression rate; and generate answers for text prompts based on the visual language big model.
[0062] Specifically, a differentiated dynamic compression ratio is assigned to each layer of the attention network based on the response intensity score. This can be achieved by using a proportional allocation method after sorting by response intensity score; a grouping allocation strategy based on nearest neighbor clustering; or a dynamic prediction model based on meta-learning to generate the optimal compression ratio for each layer in real time based on the input video and text features. There are no restrictions on the allocation method of the dynamic compression ratio. The target temporal index set is determined based on the multi-head attention score matrix. This can be achieved by using a selection method based on the mean of attention scores; a stability screening method based on variance or frequency; or a mixed-dimensional sorting method. There are no restrictions on the screening method of the target temporal index set. Finally, based on the target temporal index set and the corresponding dynamic compression ratio, the cache of each layer of the attention network is filtered and compressed layer by layer, and all video feature segments are iteratively processed until the compression and cache update of all video features are completed. After the pre-filling stage, the decoding stage of the large language model is initiated, generating an accurate answer to the text prompt by predicting the next token.
[0063] Through the above steps, the long video is first segmented and concatenated with text prompts. While ensuring it doesn't exceed the context window, the multi-head attention score matrix between text and video features, as well as the global compression rate of the video features, are calculated. The response intensity score of each layer of the attention network is then calculated using the multi-head attention score matrix and the global compression rate. Based on the response intensity score, a differentiated compression rate is assigned to each layer. Then, according to the target temporal index set determined by the multi-head attention score matrix and the dynamic compression rate, the cache of each layer of the attention network is compressed, allowing more features to be retained in layers more critical to the task. After iteratively processing all video segments, the final decoding stage generates the answer based on the compressed, high-quality cache. This reduces the interference of redundant features on model decisions, effectively improving the understanding performance of large visual language models, and is suitable for practical applications such as long video surveillance and human-computer interaction. It solves the problem in related technologies where redundant visual features crowd out the context space, leading to insufficient understanding performance of large visual language models.
[0064] The above steps are explained in detail below:
[0065] In some embodiments, step S220, determining the splicing features based on the video features corresponding to the long video and the text features corresponding to the text prompts, includes the following steps:
[0066] The visual encoder of the large visual language model extracts corresponding video features from long videos and assigns temporal indices to the video features according to time order.
[0067] The text encoder of the visual language big data model converts the text prompts into corresponding text features;
[0068] The video features and text features are concatenated to obtain the concatenated features.
[0069] Specifically, a large visual language model typically consists of three parts: a visual encoder, a text encoder, and a large language model. For example... Figure 3 As shown, a long video is acquired using an RGB video camera, and video features are extracted using a visual encoder. It includes three dimensions. This represents the number of image frames in a long video sequence. This represents the number of tokens representing the spatial dimensions of each image's features. This represents the length of the feature for each token.
[0070] The text prompt is converted into corresponding text features using a text encoder. It contains two dimensions. This represents the number of tokens after text feature encoding. This represents the length of the feature for each token.
[0071] Text features are concatenated with time-series-prioritized video features and input into the large language model of the visual language model. Video features... Break it down into chronological order of time. Video feature segments:
[0072] ;
[0073] in, For the segmented video features, The length of the split video segment is given by N, where N represents the number of tokens in the spatial dimension of each frame's image features, and C represents the feature length of each token. The length of the last video feature segment can be less than [a certain value]. .
[0074] Temporal priority specifically refers to the priority of concatenation with text features. Let the video features that need to be spliced according to time order be denoted as... The concatenated feature is obtained by concatenating it with the text features, and the expression for the concatenated feature is:
[0075] ;
[0076] in, For text features, The video features are split, and len represents the number of tokens after text feature encoding. N represents the length of the split video segment, N represents the number of tokens in the spatial dimension of each frame's image features, and C represents the feature length of each token.
[0077] Another method for obtaining stitched features is to perform spatial importance analysis on video features to determine the importance score of each spatial region in each frame. Based on the spatial importance score, video features corresponding to spatial regions with importance higher than a preset threshold are selected from each frame, resulting in the selected video features. These selected video features are then stitched together with text features to obtain the stitched features. The importance score can be determined by calculating the gradient magnitude, texture complexity, or similarity to pre-trained visual priors for each spatial region in each frame. No restrictions are placed on the method for obtaining stitched features.
[0078] In this embodiment, long video features are divided into time-series blocks and preferentially concatenated with text features, thereby preserving the temporal continuity and integrity of video features. This provides a structured input basis for subsequent block-by-block iterative processing and dynamic compression, ensuring that the model can efficiently utilize the limited context space to focus on key information.
[0079] In some of these embodiments, the total number of tokens for the splicing feature is less than or equal to the maximum context window length.
[0080] Specifically, to ensure that the total number of tokens in the concatenated features does not exceed the maximum context window length of the large language model, the constraint expression that needs to be satisfied is:
[0081] ;
[0082] in, This represents the number of tokens after text feature encoding. The length of the split video segment. This represents the number of tokens in the spatial dimension of each image feature, and win represents the maximum length of the context window for the large language model.
[0083] This embodiment provides clear constraints for the block processing of video features, ensuring that the feature length of each input to the large language model is within the context window that the model can process, thus avoiding truncation or inference failure due to excessively long features.
[0084] In some embodiments, the step S230 of assigning differentiated dynamic compression rates to each layer of the attention network based on the response intensity score includes the following steps:
[0085] Based on the response intensity scores, the attention networks of each layer are grouped to obtain the grouping results;
[0086] Based on the grouping results, the dynamic compression ratio of each layer of the attention network is determined.
[0087] Specifically, firstly, based on the concatenation features, the multi-head attention score matrix of each layer of the attention network is calculated. This can be done sequentially according to the Transformer implementation of the large language model. The expression for the multi-head attention score matrix of a layer is:
[0088] ;
[0089] in, represents the number of heads in the multi-head attention score matrix, and len represents the number of tokens after text feature encoding. The length of the split video segment. This represents the number of tokens representing the spatial dimensions of each image's features.
[0090] Secondly, the global compression ratio of the video features is calculated. The global compression ratio is calculated based on the maximum context window length of the large language model, the text feature length, and the total length of all video features. The formula for calculating the global compression ratio is:
[0091] ;
[0092] in, The maximum context window length for a large language model This represents the number of tokens after text feature encoding. This represents the number of image frames in a long video sequence. This represents the number of tokens representing the spatial dimensions of each image's features.
[0093] Then, based on the multi-head attention score matrix and the global compression ratio, the response strength score of each layer of the attention network is calculated. For example, for a large language model... Layer network, then the first Layer response strength score The calculation formula is:
[0094] ;
[0095] in, This is an indicator function that is set to 1 when the condition is met and 0 when the condition is not met. In response to the intensity index, For the first The average response intensity of the layered video to the text prompts. The calculation formula is:
[0096] ;
[0097] in, This indicates that an element-wise strategy is used to compare the size of each column of features and obtain the Kth value arranged from largest to smallest as the final output value. The calculation formula is:
[0098] ;
[0099] The KV Cache (Key-Value Cache) uses a dynamic compression ratio allocation strategy, employing a top-neighbor network layer approach. Clustering methods divide the adjacent network layers of a large language model into... Groups, given Response strength score of layer attention network The specific steps for grouping and dynamic compression ratio allocation are as follows:
[0100] right Sort by value in descending order and filter Top set of element indices The expression that satisfies this condition is:
[0101] ;
[0102] Initialize grouping ;
[0103] Traverse all ,calculate ,Will Add distance The latest group When the distance is the same as multiple groups, select to assign to A smaller group;
[0104] Obtain the adjacency network of the large language model There are several groups, and the expression for grouping is:
[0105] ;
[0106] Finally, as Figure 4 As shown, based on the grouping results, a dynamic compression ratio is assigned to each layer of the attention network. The formula for calculating the dynamic compression ratio is:
[0107] ;
[0108] Among them, the Layer allocation dynamic compression ratio The dynamic compression ratio of the group corresponding to this layer Same means that different attention layers in the same group are assigned the same compression rate.
[0109] The dynamic compression ratio can also be assigned based on the response intensity scores of each attention network layer, ranking the layers to obtain a ranking result; based on the ranking result, each layer is divided into multiple continuous intervals; and based on the statistical characteristics of the response intensity scores within each interval, a differentiated dynamic compression ratio is assigned to each interval. There are no restrictions on the method of assigning the dynamic compression ratio.
[0110] This embodiment achieves refined and differentiated allocation of compression rates for each layer of the attention network in a large language model. While maintaining the overall compression ratio, it significantly improves the information retention quality of key layers, thereby effectively enhancing the model's understanding accuracy while reducing computational and caching overhead.
[0111] In some embodiments, determining the target temporal index set based on the multi-head attention score matrix in step S230 includes the following steps:
[0112] Based on the multi-head attention score matrix, determine the intensity mean ranking result, intensity variance ranking result, and intensity frequency ranking result;
[0113] The target time series index set is determined based on the intensity mean sorting result, intensity variance sorting result, and intensity frequency sorting result.
[0114] Specifically, the target temporal index set can be selected based on the intensity mean sorting result, intensity variance sorting result, and intensity frequency sorting result; it can also be based on the entropy value of the attention score to measure the information content of the features. The lower the entropy value, the more concentrated and discriminative the response is, thus prioritizing the retention of low-entropy features; it can also introduce a cross-layer consistency index to calculate the variance or correlation coefficient of the same video token in the attention scores of different layers, and select features with stable cross-layer responses that are not prone to drastic changes; or it can combine gradient information to evaluate its contribution to the final prediction by calculating the gradient magnitude of the model loss on the video token, and prioritizing the retention of features with larger gradients; there are no restrictions on the method of determining the target temporal index set.
[0115] For example, such as Figure 5 As shown, the visual encoder first extracts video features from the long video, while the text encoder simultaneously converts the text prompts in the long video into text features. These two types of features are combined to form concatenated features, which are then input into the large language model. In the pre-filling stage, a dynamic allocation strategy for cache compression ratios assigns differentiated compression ratios to each layer. Combined with the inverse cache compression module, high-quality video features are selected and retained based on the attention score matrix. Then, based on these indices, the K-cache and V-cache of the large language model are filtered, compressed, and updated, filtering out redundant video feature caches. Finally, the large language model enters the decoding and prediction stage. Based on the processed features and cache information, it generates and outputs accurate answers to the text prompts, such as recognizing the action type "clapping, taking fries, taking a sandwich," thus completing the entire process of long video understanding. Figure 6 As shown, in a long video of 293 seconds, through dynamic caching, compression and updating, the human action information during periods such as "clapping", "taking fries", and "taking a sandwich" can be accurately located, assisting the visual language model in making accurate and robust understanding and analysis.
[0116] This embodiment provides diverse evaluation dimensions and flexible selection space for determining the target time-series index set, enabling cache compression to adaptively select the most suitable screening indicators or combinations based on specific task requirements, model characteristics, or data features. This maximizes compression efficiency while preserving key information, significantly improving the adaptability and robustness of the method in different application scenarios.
[0117] In some embodiments, the target time-series index set is determined based on the intensity mean sorting result, intensity variance sorting result, and intensity frequency sorting result, including the following steps:
[0118] The reciprocal mixed ranking score is determined based on the ranking results of intensity mean, intensity variance, and intensity frequency.
[0119] The target time-series index set is selected based on the inverse mixed sort score.
[0120] Specifically, such as Figure 7 As shown, in the pre-filling stage, the compression ratio of each layer is calculated and allocated using the KV Cache dynamic compression ratio. Then, the video KV Cache is reduced based on the inverse cache compression module. An inverse hybrid sorting method is used to reorder and filter the temporal indices of video features, resulting in a target temporal index set of high-quality video features. Target time series index set The expression is:
[0121] ;
[0122] in, To select Top indivual The video token time-series index corresponding to the numerical value. For the first The inverse mixed sort score of the layer, Assign a dynamic compression ratio to the l-th layer.
[0123] Based on the selected high-quality video token index The cache of each layer of the attention network is compressed. The expression for video feature compression is:
[0124] ;
[0125] ;
[0126] in, For the existing K cache in layer l, For the existing V cache in layer l, The current task to be processed K-caching of partial video features, The current task to be processed V-caching of partial video features. Thus, the first... Key-Value Cache (KV) Compressing and Update in Layered Attention Networks.
[0127] The target time-series index set can also be obtained by filtering it based on the weighted product fusion score. Through weighted product fusion, the fusion score will only be significantly higher when the video token performs well across all three dimensions. There are no restrictions on the method used to obtain the target time-series index set.
[0128] This embodiment combines inverse hybrid sorting with dynamic compression ratio to achieve efficient and accurate compression of the KV cache. Inverse hybrid sorting considers three dimensions—mean intensity, first-degree variance, and intensity frequency—to select high-quality and stable temporal indices for video tokens. This helps reduce the volatility of subsequent KV cache compression, thereby preserving the most important information for the task.
[0129] In some of these embodiments, the expression for the reciprocal mixed sort score is:
[0130] ;
[0131] In the formula, The smoothing weighting coefficients set for experience, The results are sorted by the mean intensity. The results of the intensity variance ranking The results are sorted by intensity frequency.
[0132] Specifically, the intensity mean ranking results, intensity variance ranking results, and intensity frequency ranking results can be obtained in several ways, one of which is by calculation using mathematical formulas. If mathematical formulas are used for calculation, then... The calculation formula is:
[0133] ;
[0134] The calculation formula is:
[0135] ;
[0136] The calculation formula is:
[0137] ;
[0138] in, This function returns indices in descending numerical order, where len represents the number of tokens after text feature encoding. The number of heads in the multi-head attention score matrix. Let be the multi-head attention score matrix of the l-th layer. Here is the frequency indication function, and its expression is:
[0139] ;
[0140] in, This query retrieves the third-quarter value of the attention matrix and returns the corresponding value.
[0141] This embodiment comprehensively evaluates and ranks video features from three key dimensions: intensity, stability, and frequency, avoiding the screening bias that may be caused by a single indicator.
[0142] This embodiment also provides a visual language large-scale model understanding device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as described above. The terms "module," "unit," "subunit," etc., used below can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0143] Figure 8 This is a structural block diagram of the visual language large-scale model understanding device in this embodiment, as shown below. Figure 8 As shown, the device includes: an acquisition module 210, a processing module 220, and a compression module 230;
[0144] The acquisition module 210 is used to acquire the long video and the text prompts corresponding to the long video;
[0145] Processing module 220 is used to input long videos and text prompts into a pre-trained visual language large model; the visual language large model includes a large language model; determine splicing features based on the video features corresponding to the long video and the text features corresponding to the text prompts; input the splicing features into the large language model; determine the multi-head attention score matrix of each layer of attention network in the large language model based on the splicing features; determine the global compression rate based on the maximum context window length, text features, and video features of the large language model; and determine the response intensity score of each layer of attention network based on the multi-head attention score matrix and the global compression rate.
[0146] Compression module 230 is used to assign differentiated dynamic compression rates to each layer of the attention network based on the response intensity score; determine the target temporal index set based on the multi-head attention score matrix; and compress the cache of each layer of the attention network based on the target temporal index set and the dynamic compression rate; and generate answers for text prompts based on the visual language big model.
[0147] The aforementioned device solves the problem in related technologies where redundant visual features crowd out the context space, leading to insufficient understanding of large visual language models, and improves the understanding effect of large visual language models.
[0148] In some embodiments, the processing module 220 is further configured to extract corresponding video features from the long video using a visual encoder of a large visual language model; and to assign temporal indexes to the video features in chronological order.
[0149] The text encoder of the visual language big data model converts the text prompts into corresponding text features;
[0150] The video features and text features are concatenated to obtain the concatenated features.
[0151] In some of these embodiments, the total number of tokens for the splicing feature is less than or equal to the maximum context window length.
[0152] In some embodiments, the compression module 230 is further configured to group the attention network layers according to the response intensity scores to obtain grouping results;
[0153] Based on the grouping results, the dynamic compression ratio of each layer of the attention network is determined.
[0154] In some embodiments, the compression module 230 is further configured to determine the intensity mean ranking result, the intensity variance ranking result, and the intensity frequency ranking result based on the multi-head attention score matrix.
[0155] The target time series index set is determined based on the intensity mean sorting result, intensity variance sorting result, and intensity frequency sorting result.
[0156] In some embodiments, the compression module 230 is further configured to determine the reciprocal mixed sort score based on the intensity mean sorting result, the intensity variance sorting result, and the intensity frequency sorting result;
[0157] The target time-series index set is selected based on the inverse mixed sort score.
[0158] In some of these embodiments, the expression for the reciprocal mixed sort score is:
[0159] ;
[0160] In the formula, The smoothing weighting coefficients set for experience, The results are sorted by the mean intensity. The results of the intensity variance ranking The results are sorted by intensity frequency.
[0161] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.
[0162] This embodiment also provides a computer device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0163] Optionally, the computer device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0164] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0165] S1, retrieve the long video and the corresponding text prompt.
[0166] S2, input the long video and text prompts into the pre-trained visual language large model; the visual language large model includes a large language model; determine the splicing features based on the video features corresponding to the long video and the text features corresponding to the text prompts; input the splicing features into the large language model; determine the multi-head attention score matrix of each layer of attention network in the large language model based on the splicing features; determine the global compression rate based on the maximum context window length, text features, and video features of the large language model; determine the response intensity score of each layer of attention network based on the multi-head attention score matrix and the global compression rate.
[0167] S3. Based on the response intensity score, assign differentiated dynamic compression rates to each layer of the attention network; determine the target temporal index set based on the multi-head attention score matrix; and compress the cache of each layer of the attention network based on the target temporal index set and the dynamic compression rate; and generate answers for text prompts based on the visual language big model.
[0168] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated in this embodiment.
[0169] Furthermore, in conjunction with the visual language large model understanding method provided in the above embodiments, this embodiment can also provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the visual language large model understanding methods described in the above embodiments.
[0170] It should be noted that all information and data involved in this application are authorized by the user or fully authorized by all parties and will be used legally.
[0171] It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. All other embodiments derived by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0172] Obviously, the accompanying drawings are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application.
[0173] The term "embodiment" in this application refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply that it is mutually exclusive with or independent of other embodiments. It will be clearly or implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0174] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.
Claims
1. A visual language large model understanding method, characterized in that, include: Obtain the long video and the text prompt corresponding to the long video; The long video and the text prompt are input into a pre-trained visual language large model; The visual language big model includes a big language model; Based on the video features corresponding to the long video and the text features corresponding to the text prompt, the splicing features are determined; and the splicing features are input into the large language model. Based on the splicing features, determine the multi-head attention score matrix of each layer of the attention network in the large language model; determine the global compression rate based on the maximum context window length of the large language model, the text features, and the video features; Based on the multi-head attention score matrix and the global compression ratio, the response strength score of each layer of the attention network is determined; Based on the response intensity score, a differentiated dynamic compression ratio is assigned to each layer of the attention network; Based on the multi-head attention score matrix, determine the target time-series index set; And based on the target temporal index set and the dynamic compression rate, the cache of each layer of the attention network is compressed; The step of determining the target temporal index set based on the multi-head attention score matrix includes: Based on the multi-head attention score matrix, determine the intensity mean ranking result, intensity variance ranking result, and intensity frequency ranking result; The target time series index set is determined based on the intensity mean sorting result, the intensity variance sorting result, and the intensity frequency sorting result; Based on the aforementioned visual language model, a response is generated for the text prompt.
2. The visual language large model understanding method of claim 1, wherein, The step of determining the splicing features based on the video features corresponding to the long video and the text features corresponding to the text prompt includes: The visual encoder of the large visual language model extracts the corresponding video features from the long video; and assigns temporal indexes to the video features according to the time sequence. The text encoder of the visual language big data model converts the text prompt into the corresponding text features. The video features and the text features are concatenated to obtain the concatenated features.
3. The visual language large model understanding method of claim 1, wherein, The total number of tokens in the splicing feature is less than or equal to the maximum window length of the context.
4. The visual language large model understanding method of claim 1, wherein, The step of assigning differentiated dynamic compression ratios to each layer of the attention network based on the response intensity score includes: Based on the response intensity score, the attention network layers are grouped to obtain the grouping results; Based on the grouping results, the dynamic compression ratio of the attention network in each layer is determined.
5. The visual language large-scale model understanding method according to claim 1, characterized in that, The step of determining the target time-series index set based on the intensity mean sorting result, the intensity variance sorting result, and the intensity frequency sorting result includes: The inverse mixed sort score is determined based on the sorting results of the intensity mean, the intensity variance, and the intensity frequency. The target time-series index set is selected based on the inverse mixed sort score.
6. The visual language large-scale model understanding method according to claim 5, characterized in that, The expression for the reciprocal mixed sort score is: ; In the formula, The smoothing weighting coefficients set for experience, The sorting results of the intensity mean, The ranking results of the intensity variance, The intensity frequencies are sorted.
7. A visual language large-scale model understanding device, characterized in that, include: Acquisition module, processing module, and compression module; The acquisition module is used to acquire the long video and the text prompts corresponding to the long video; The processing module is used to input the long video and the text prompt into a pre-trained visual language large model; the visual language large model includes a large language model; determine splicing features based on the video features corresponding to the long video and the text features corresponding to the text prompt; input the splicing features into the large language model; determine the multi-head attention score matrix of each layer of the attention network in the large language model based on the splicing features; determine the global compression rate based on the maximum context window length of the large language model, the text features, and the video features; and determine the response intensity score of each layer of the attention network based on the multi-head attention score matrix and the global compression rate. The compression module is used to assign a differentiated dynamic compression rate to each layer of the attention network based on the response intensity score. Based on the multi-head attention score matrix, determine the target time-series index set; And based on the target temporal index set and the dynamic compression rate, the cache of each layer of the attention network is compressed; The step of determining the target temporal index set based on the multi-head attention score matrix includes: Based on the multi-head attention score matrix, determine the intensity mean ranking result, intensity variance ranking result, and intensity frequency ranking result; The target time series index set is determined based on the intensity mean sorting result, the intensity variance sorting result, and the intensity frequency sorting result; Based on the aforementioned visual language model, a response is generated for the text prompt.
8. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the steps of the visual language large model understanding method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the visual language large model understanding method as described in any one of claims 1 to 6.