Visual token pruning method, vision-language model, device, storage medium, and program product

WO2026200257A1PCT designated stage Publication Date: 2026-10-01CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/075158
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-01-27
Publication Date
2026-10-01

Smart Images

  • Figure CN2026075158_01102026_PF_FP_ABST
    Figure CN2026075158_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a visual token pruning method, a vision-language model, a device, a storage medium, and a program product. In the embodiments of the present disclosure, for initial attention scores of a plurality of visual tokens in an input token sequence, attribute information of an image is introduced to correct the initial attention scores, wherein the attribute information reflects the importance of the image during attention score correction; the initial attention scores are corrected by means of the attribute information, in order to correct biases arising from positional differences among a plurality of visual tokens during the allocation of the initial attention scores, such that target attention scores obtained after correction more accurately indicate attention levels of different regions in the image; thus, during pruning processing, important components of each image can be retained, so as to cache important keys and values of each image, thereby substantially preserving the integrity of visual information.
Need to check novelty before this filing date? Find Prior Art

Description

Visual tagging pruning methods and visual language models, devices, storage media, and program products

[0001] This disclosure claims priority to Chinese Patent Application No. 2025103764458, filed on March 27, 2025, entitled "Visual Marker Pruning Method and Visual Language Model, Device, Storage Medium and Program Product", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of cloud computing technology, and in particular to a visual tagging pruning method, visual language model, device, storage medium, and program product. Background Technology

[0003] With the development of artificial intelligence technology, neural network models based on the self-attention mechanism have been widely used. In the self-attention mechanism, the token is the basic unit of model input and output. Each token generates three vectors: Query: used to calculate the attention score with the keys of other tokens; Key: used to calculate the attention score with the queries of other tokens; Value: a weighted sum based on the attention scores to obtain an intermediate representation of the current token, which is used to generate a new token. The new token serves as new input to continue generating the next token until a terminator appears.

[0004] To improve the model's inference speed, a KV Cache technique was employed. KV Cache is an optimization technique used in self-attention layers. It caches the key and value values ​​used in attention score calculations during token generation and reuses these historical token key and value values ​​in subsequent token generation without recalculation, thus improving computational efficiency.

[0005] If the key-value (KV) values ​​of all tokens in the cache sequence consume a large amount of GPU memory, it will increase inference time. Therefore, the KV cache pruning technique was developed, which selects important KV values ​​for caching based on attention scores; where a higher attention score indicates a more important KV value. However, in self-attention mechanisms, the self-attention matrix is ​​implemented using a lower triangular matrix. This leads to significant pruning of KV values ​​for later tokens, easily causing a loss of contextual information, especially in multi-image or video scenarios, resulting in catastrophic loss of visual information. Summary of the Invention

[0006] This disclosure provides a visual marker pruning method, visual language model, device, storage medium, and program product to improve the rate of update and iteration of control parameters during rubber mixing, thereby improving the quality of the rubber compound.

[0007] This disclosure provides a visual tag pruning method, comprising: obtaining an input tag sequence of the current self-attention sublayer, the input tag sequence including multiple visual tags, the multiple visual tags being semantic descriptions of multiple image regions generated by structured decomposition of at least two images, with each image generating at least two visual tags; obtaining multiple initial key values ​​and multiple initial attention scores corresponding to the multiple visual tags; correcting the multiple initial attention scores according to the attribute information of the at least two images to obtain multiple target attention scores, the attribute information reflecting the importance of the at least two images in the attention score correction process; and pruning the multiple initial key values ​​according to the multiple target attention scores to obtain target key values ​​to be cached.

[0008] This disclosure provides a visual language model, comprising: a plurality of decoder layers connected in sequence, wherein the output tag sequence of the preceding decoder layer is the input tag sequence of the following decoder layer; the input tag sequence includes a plurality of visual tags, wherein the plurality of visual tags are semantic descriptions of a plurality of image regions generated by structured decomposition of at least two images, and each image generates at least one visual tag; the plurality of decoder layers are configured to, in a pre-filling stage, implement the steps of various methods provided in this disclosure to obtain a target key value; and in a decoding stage, perform attention calculation based on the cached target key value to obtain new tag information.

[0009] This disclosure also provides an electronic device, including a memory and a processor. The memory is used to store a computer program, and the processor is coupled to the memory to execute the computer program for implementing the steps in the various methods provided in this disclosure.

[0010] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the processor to perform the steps in the methods described above.

[0011] This disclosure also provides a computer program product, which includes a computer program / instructions that, when executed by a processor, enable the processor to perform the steps described in the method embodiments above.

[0012] In this embodiment of the disclosure, for the initial attention scores of multiple visual markers in the input marker sequence, image attribute information is introduced to correct the initial attention scores. The attribute information reflects the importance of the image in the attention score correction process. By correcting the initial attention scores with attribute information, the deviation caused by the different positions of multiple visual markers when assigning initial attention scores is corrected. The corrected target attention score more accurately indicates the attention level of different regions in the image. Furthermore, during the pruning process, important components in each image can be preserved, and important key values ​​in each image can be cached, thus protecting the integrity of visual information to a greater extent. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation of the disclosure. In the drawings:

[0014] Figure 1a is a schematic diagram of the structure of a self-attention matrix provided in an exemplary embodiment of the present disclosure;

[0015] Figure 1b is a schematic diagram of a visual tagging pruning method provided in an exemplary embodiment of this disclosure;

[0016] Figure 2 is a schematic diagram of a non-critical key-value recycling process provided in another exemplary embodiment of this disclosure;

[0017] Figure 3a is a schematic diagram of the architecture of a visual language model provided in another exemplary embodiment of this disclosure;

[0018] Figure 3b is a schematic diagram of another modified result provided by yet another exemplary embodiment of this disclosure;

[0019] Figure 4 is a schematic diagram of the structure of an electronic device provided in another exemplary embodiment of this disclosure. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this disclosure clearer, the technical solutions of this disclosure will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.

[0021] It should be noted that, in the cases involving user information in the embodiments of this disclosure, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this disclosure (including but not limited to language models or large models) comply with relevant laws and standards.

[0022] In the self-attention mechanism, in order to achieve the autoregressive property and ensure that the model only accesses previous tokens and not future tokens when generating the current token, a lower triangular matrix is ​​used as the self-attention matrix to mask the position of future tokens.

[0023] Figure 1a shows an example of a lower triangular matrix. In this example, the local regions marked in each column are used to generate the token for the image corresponding to that column. For example, U1 represents the token for the first image, which includes the tokens of the local regions marked 1, 2, 3, and 4 in the first column of Figure 1a. Similarly, U2, U3, and U4 represent the tokens corresponding to the second, third, and fourth images, respectively, and each includes the token of the local region of its respective column.

[0024] When the self-attention matrix is ​​implemented using a lower triangular matrix, tokens positioned earlier in the hierarchy, such as U4, receive higher attention scores, while tokens positioned later receive lower attention scores. Furthermore, when using key-value cache pruning techniques, tokens positioned later in the hierarchy suffer from lower attention scores, leading to significant pruning of their key-value values. This can result in a loss of contextual information, especially in multi-image or video scenarios where key-value values ​​in later images are heavily pruned, potentially causing catastrophic loss of visual information.

[0025] To address the aforementioned technical problems, in this embodiment of the disclosure, image attribute information is introduced to correct the initial attention scores of multiple visual markers in the input marker sequence. The attribute information reflects the importance of the image in the attention score correction process. By correcting the initial attention scores using attribute information, the deviation caused by the different positions of multiple visual markers when assigning initial attention scores is corrected. The corrected target attention score more accurately indicates the attention level of different regions in the image. Furthermore, during pruning, important components in each image can be preserved, and important key values ​​in each image can be cached, thus protecting the integrity of visual information to a greater extent.

[0026] The technical solutions provided by the embodiments of this disclosure are described in detail below with reference to the accompanying drawings.

[0027] Figure 1b is a schematic diagram of a visual tagging pruning method provided in an exemplary embodiment of this disclosure. As shown in Figure 1b, the method includes:

[0028] S101: Obtain the input label sequence of the current self-attention sublayer. The input label sequence includes multiple visual labels. The multiple visual labels are semantic descriptions of multiple image regions generated by the structured decomposition of at least two images. Each image generates at least two visual labels.

[0029] S102: Obtain multiple initial key values ​​and multiple initial attention scores corresponding to multiple visual tags;

[0030] S103: Based on the attribute information of at least two images, correct multiple initial attention scores to obtain multiple target attention scores. The attribute information reflects the importance of the image in the attention score correction process.

[0031] S104: Based on multiple target attention scores, prune multiple initial key values ​​to obtain the target key values ​​that need to be cached.

[0032] In this embodiment, the model architecture for applying the above-described visual pruning method is not limited. Any model based on the self-attention mechanism can be applied to this embodiment, and the model "based on the self-attention mechanism" is referred to as a "self-attention model". The self-attention model may include multiple self-attention sub-layers. When processing the input labeled sequence, the self-attention sub-layers can capture the dependencies between different positions in the input labeled sequence, better understand the contextual information in the input sequence, and generate a more accurate input representation.

[0033] In this embodiment, the pruning method can be applied to any of multiple self-attention sub-layers. The self-attention sub-layer to which the above method is currently applied is referred to as the current self-attention sub-layer. The following description uses the application of the above method to the current self-attention sub-layer as an example.

[0034] In this embodiment, an input token sequence is obtained for the current self-attention sublayer. In some embodiments, two images are segmented into multiple patches, and each patch is then processed through linear embedding and positional encoding to form an input token sequence, which includes multiple tokens arranged in order.

[0035] In this embodiment, the input tag sequence includes multiple visual tags, which can be implemented as the Token described above. Multiple visual tags are semantic descriptions of multiple image regions generated by structured decomposition of at least two images. Structured decomposition refers to decomposing at least two images according to the order between them to generate multiple visual tags. Each image generates at least two visual tags, and each visual tag corresponds to a local region of a given image.

[0036] In this embodiment, if the current self-attention sublayer is the first self-attention sublayer in the self-attention model, the input label sequence is generated by preprocessing image regions in each image in the order of at least two images. Each image region is preprocessed to obtain one visual label, and each image generates at least two visual labels. The multiple visual labels generated by at least two images form the input label sequence. If the current self-attention sublayer is not the first self-attention sublayer, the input label sequence of the non-first self-attention sublayer is the output label sequence of the previous self-attention sublayer. The number and order of visual labels in this output label sequence are the same as those in the input label sequence of the first self-attention sublayer.

[0037] In this embodiment, multiple initial key values ​​and multiple initial attention scores corresponding to multiple visual tags are obtained. Optionally, based on the weight matrix of the current self-attention sub-layer, contextual information perception processing is performed on the multiple visual tags to obtain multiple initial key values ​​and multiple initial attention scores corresponding to the multiple visual tags. The weight matrix is ​​used to capture the dependencies between multiple visual tags in the input tag sequence; these weight matrices are learned during training. In a model with multiple self-attention sub-layers, each self-attention sub-layer has its own weight matrix to capture richer information representations in the input tag sequence; the initial key values ​​generated by different self-attention sub-layers reflect different levels of semantic information.

[0038] In this process, each visual tag in the input tag sequence is linearly transformed with the corresponding weight matrix in the current self-attention sublayer to obtain three corresponding vectors: Query, Key, and Value. Key and Value are collectively referred to as the initial key-value pair. For ease of description and distinction, in subsequent embodiments, Query, Key, and Value will be abbreviated as Q-vector, K-vector, and V-vector, respectively.

[0039] Furthermore, in this embodiment, the Q vector and K vector are used to calculate the initial attention score. The initial attention score reflects the degree of association between a visual tag at a certain position in the input tag sequence and visual tags at all positions. In some implementations, for each visual tag's corresponding Q vector, a similarity score matrix S is calculated by performing a dot product between the Q vector and the K vectors at all positions. Further, this similarity score matrix S is divided by the square root of the dimension of the K vectors to obtain S', and then the SoftMax function (an activation function) is applied to calculate the weights to obtain the attention weight matrix A, where each element of the attention weight matrix A corresponds to the initial attention score of a visual tag.

[0040] Below is an example of an attention matrix A, such as... in, n represents the position of the image, N represents the total number of images (N is an integer greater than or equal to 2), and X represents the vector formed by the visual markers generated from the nth image. It is the index of each visual marker generated from the nth image. In this matrix... In this context, each element corresponds to the initial attention score of a visual tag, where i represents the position of the visual tag in the input tag sequence. Here, 0 ≤ i ≤ L-1, and i is a natural number, while L is the length of the input tag sequence.

[0041] The position of an image affects the initial attention score of a visual tag. The further back in the image is in the input tag sequence, the further back the corresponding visual tag will be in the sequence, and the smaller the assigned initial attention score will be. When using KV Cache pruning, pruning is performed based on the initial attention score, retaining those with higher scores and pruning those with lower scores. Therefore, visual tags positioned later in the sequence, due to their lower initial attention scores, will have their initial key values ​​significantly pruned, potentially leading to a loss of contextual information.

[0042] Therefore, in this embodiment, multiple initial attention scores are corrected based on the attribute information of at least two images. The attribute information reflects the importance of the image in the attention score correction process, and may include, but is not limited to, positional attributes and memory attributes. Correcting the initial attention scores using the image attribute information corrects the deviation in the allocation of initial attention scores caused by the different positions of multiple visual markers. After correcting the initial attention scores, a target attention score is obtained, which accurately reflects the importance of the visual regions of each image.

[0043] Furthermore, based on the corrected target attention scores, multiple initial key values ​​are pruned to obtain the target key values ​​to be cached. Since the target attention scores are corrected attention scores, the bias caused by the different positions of multiple visual tags when assigning initial attention scores is reduced, more accurately indicating the attention level of the current self-attention sublayer to different regions of the image during pruning. Therefore, during pruning, important components in each image can be preserved, allowing important contextual information (i.e., target key values) in each image to be cached, thus protecting the integrity of visual information to a greater extent.

[0044] In this embodiment of the disclosure, for the initial attention scores of multiple visual markers in the input marker sequence, image attribute information is introduced to correct the initial attention scores. The attribute information reflects the importance of the image in the attention score correction process. By correcting the initial attention scores with attribute information, the deviation caused by the different positions of multiple visual markers when assigning initial attention scores is corrected. The corrected target attention score more accurately indicates the attention level of different regions in the image. Furthermore, during the pruning process, important components in each image can be preserved, and important key values ​​in each image can be cached, thus protecting the integrity of visual information to a greater extent.

[0045] In this embodiment of the disclosure, an image generates at least two visual markers, each corresponding to an initial attention score. The initial attention score / target attention score corresponding to the visual markers generated for an image is simply referred to as the initial attention score / target attention score corresponding to the image. Here, " / " represents "or". In subsequent embodiments, an intermediate attention score corresponding to the visual markers generated for an image may also appear, such as a first intermediate attention score / second intermediate attention score, which can be simply referred to as the first intermediate attention score / second intermediate attention score corresponding to the image, respectively. This will not be repeated in subsequent embodiments.

[0046] In this embodiment, multiple initial key values ​​and multiple initial attention scores corresponding to multiple visual tags are obtained. Optionally, based on the weight matrix of the current self-attention sub-layer, multiple visual tags in the input tag sequence are subjected to contextual information perception processing to obtain multiple initial attention scores.

[0047] In an optional embodiment, the initial attention score corresponding to an image can be calculated using the following formula (1), but is not limited thereto.

[0048] In the above formula (1), S n S represents the initial attention score corresponding to the nth image. n X is a vector containing the initial attention scores for each visual label generated from the nth image; n represents the nth image, with values ​​from 1 to N, where N represents the total number of images and is an integer greater than or equal to 2; X represents the vector formed by the visual labels generated from the nth image. It represents the indices of the visual markers generated from the nth image. The matrix... An element corresponds to an initial attention score for a visual tag. In the formula (1), 0 ≤ i ≤ L-1, and i is a natural number, L is the length of the input label sequence, and L is an integer greater than or equal to 2. For each visual label generated in the nth image, the sum of the attention scores of the visual label and all visual labels generated in the N images is calculated, and the initial attention score of the visual label is calculated. Furthermore, the initial attention scores of each visual label generated in the nth image form a vector, which is used as the initial attention score corresponding to the nth image.

[0049] Furthermore, in this embodiment, multiple initial attention scores are corrected based on the attribute information of at least two images. Optionally, the image attribute information includes positional attributes and / or content attributes. The positional attribute of an image indicates its relative position among the at least two images, such as whether it is in front or behind. The purpose of introducing positional attributes to correct the initial attention scores is to correct the deviations caused by the different positions of multiple visual markers when assigning initial attention scores, so that the attention scores corresponding to the at least two images tend to be averaged after correction. Additionally, the content attributes of an image can be described by the relevance between the visual information and textual information corresponding to each of the at least two images, but are not limited to this. For example, the content attributes of an image can also be described by the main content or theme content of the image, such as portraits, landscapes, or architecture. The theme content refers to the main idea or emotion expressed by the image; the main content refers to the main content object in the image, such as portraits, landscapes, or architecture.

[0050] The textual information includes, but is not limited to: descriptive text, tags and categories, titles and summaries, questions and answers, etc. Introducing the content attributes of images aims to leverage the correlation between textual information and visual information in the images to guide the correction of attention scores, thereby preserving semantically relevant important components in each image during pruning.

[0051] Optionally, during the process of correcting multiple initial attention scores based on the attribute information of at least two images, the multiple initial attention scores are corrected at the image level based on the positional attributes and / or content attributes of at least two images to obtain multiple target attention scores. The positional attributes and content attributes can be chosen arbitrarily, or both can be selected; there is no limitation on this. The following explanation uses the example of correcting the initial attention scores simultaneously based on positional attributes and content attributes to illustrate the correction process, but it is not limited to this.

[0052] In one optional embodiment, multiple initial attention scores are corrected at the image level based on the positional and content attributes of at least two images to obtain multiple target attention scores.

[0053] The following describes the process of initially correcting the initial attention score based on the positional attributes of at least two images.

[0054] The position of an image affects its initial attention score. In at least two images, the later an image is positioned in the input label sequence, the later its generated visual label will be, resulting in a relatively smaller initial attention score. Therefore, in this embodiment, the initial attention score is used to characterize the image's positional attribute; a lower initial attention score indicates a later position in the at least two images, and vice versa.

[0055] Based on this, in an optional embodiment, the initial attention scores corresponding to at least two images are used as positional attributes, and the initial attention scores corresponding to at least two images are initially corrected to obtain first intermediate attention scores corresponding to at least two images. The initial correction is used to dynamically compensate for the initial attention scores of the images, preventing excessive attention scores from being assigned to images with extremely low initial attention scores, and correcting deviations caused by different image positions when assigning initial attention scores, so that the corrected attention scores tend to be more averaged.

[0056] Further optionally, for any one of the at least two images, the initial attention score corresponding to that image is normalized according to the initial attention scores corresponding to the at least two images respectively, so as to obtain the first intermediate attention score corresponding to that image, so that the first intermediate attention scores corresponding to the at least two images tend to be averaged.

[0057] In an optional embodiment, during the process of normalizing the initial attention score of any image based on the initial attention scores of at least two images, the attention percentage of the image in the at least two images is calculated based on the initial attention score of the image and the initial attention scores of the at least two images.

[0058] In this embodiment, the method for calculating the attention percentage of any image in at least two images is not limited.

[0059] In one optional embodiment, during the process of calculating the attention percentage of the image in at least two images, the average of the initial attention scores corresponding to each of the at least two images is calculated to obtain the average attention score corresponding to each of the at least two images; a first numerical calculation is performed on the average attention scores corresponding to each of the at least two images to obtain a first calculation result corresponding to each of the at least two images; the first numerical calculation includes, but is not limited to, natural exponential operations, addition operations, and multiplication operations, etc.; further, the ratio of the first calculation result corresponding to the image to the sum of the first calculation results corresponding to each of the at least two images is used as the attention percentage of the image in the at least two images.

[0060] Following the above example, this disclosure also provides two methods for calculating attention percentage, as shown in formulas (2) and (3) below, but is not limited thereto.

[0061] In formula (2), Fn represents the attention percentage, n represents the nth image, which takes values ​​from 1 to N, and N represents the total number of images, where N is an integer greater than or equal to 2. This represents the average of the initial attention scores corresponding to the nth image; This involves applying the natural exponential operation to the average of the initial attention scores for the nth image. The natural exponential operation can be considered as one example of the first numerical operation, but is not limited to it. Furthermore, An example corresponding to the first calculation result. The first calculation result corresponding to the nth image (e.g.) The sum of the first calculation results corresponding to each of the N images (e.g.) The ratio of the nth image to the total attention given to the nth image among the N images is used as the percentage of attention given to the nth image.

[0062] In the above formula (3), Fn′ represents the attention percentage, n represents the nth image, and its value ranges from 1 to N. N represents the total number of images, and N is an integer greater than or equal to 2. This represents the average of the initial attention scores corresponding to the nth image; for Multiplying by 1 is one example of using 1 as the first numerical operation, but it is not limited to this. In this example, An example corresponding to the first calculation result. The first calculation result corresponding to the nth image (e.g.) The sum of the first calculation results corresponding to each of the N images (e.g.) The ratio of the nth image to the total attention given to the nth image among the N images is used as the percentage of attention given to the nth image.

[0063] Furthermore, in this embodiment, the initial attention score corresponding to any image is normalized according to the attention ratio to obtain the first intermediate attention score corresponding to that image. After the initial attention scores of at least two images are normalized, their respective first intermediate attention scores will tend to be averaged.

[0064] In this embodiment, the method of normalizing the initial attention score corresponding to any image is not limited.

[0065] In one optional implementation, a second numerical calculation is performed based on a preset compensation factor and attention percentage to obtain a normalization coefficient. Then, based on this normalization coefficient, the initial attention score corresponding to any image is normalized to obtain a first intermediate attention score corresponding to that image.

[0066] The following are two normalization calculation methods, as shown in formulas (4) and (5). Formulas (4) and (5) are just examples of normalization and are not limited to them.

[0067] In an optional embodiment, a normalization process is shown in formula (4) above, where, Let represent the attention score of the first intermediate state corresponding to the image, γ represent the preset compensation factor, and Fn represent the attention percentage; e γ(1-Fn) This is one example of performing a second numerical calculation on a preset compensation factor and attention percentage, but it is not limited to this. When the second numerical calculation is a natural exponential operation, γ is greater than 1 to facilitate averaging the attention scores corresponding to each visual marker. It should be understood that the natural exponential operation is only an example and does not constitute a limitation on this embodiment. If the base of the second numerical operation is changed to a base less than 1, then γ is less than 1. In this example, e... γ(1-Fn)This can serve as an example of a second calculation result, namely, an example of a normalized coefficient. Furthermore, S n Let represent the initial attention score of the nth image. In the above formula (4), "*" represents a multiplication operation. By multiplying the initial attention score of the nth image by the normalization coefficient, the initial attention score corresponding to the nth image is normalized to obtain the first intermediate state attention score corresponding to the nth image. Of course, multiplication is just one method and does not constitute a limitation on this embodiment.

[0068] In another alternative embodiment, a normalization process is shown in the following formula (5):

[0069] In the above formula (5), Let γ represent the first intermediate state attention score corresponding to the image, γ represent the preset compensation factor, and Fn represent the attention ratio; (1-Fn) is an example of a second numerical calculation of the preset compensation factor and attention ratio, but is not limited to this. In this example, (1-Fn) can be used as an example of the second calculation result, that is, an example of the normalization coefficient. In formula (5), "*" represents the multiplication operation. By multiplying the initial attention score of the nth image with the normalization coefficient, the initial attention score corresponding to the nth image is normalized to obtain the first intermediate state attention score corresponding to the nth image.

[0070] In the above embodiments, a process for initial correction of multiple initial attention scores based on positional attributes is described. Initial correction can dynamically compensate for the initial attention scores of the images, preventing excessive attention scores from being allocated to images with extremely low initial attention scores. It corrects deviations in initial attention score allocation caused by different image positions, so that the first intermediate attention scores of at least two corrected images tend to be averaged. It should be understood that the method of initial correction for multiple initial attention scores in this embodiment is not limited to the normalization process described above. Any method that can make the first intermediate attention scores of at least two corrected images tend to be averaged is applicable to initial correction. For example, a correction weight corresponding to each image can be determined based on the image's positional attributes; the later the image is positioned, the greater the correction weight. Based on the correction weight corresponding to each image, the initial attention score corresponding to the visual marker generated by that image is initially corrected. For example, the initial attention score can be multiplied by the correction weight to obtain the first intermediate attention score. For example, assuming there are three images with corresponding initial attention scores of 80, 70, and 60, the lower the initial attention score of an image, the later its position among the three images, and the greater the assigned correction weight. For example, images with initial attention scores of 80, 70, and 60 are assigned correction weights of 0.6, 0.7, and 0.8, respectively. Then, the initial attention scores can be multiplied by their corresponding correction weights to obtain their respective first intermediate attention scores. For instance, 80 multiplied by 0.6 = 48, 70 multiplied by 0.7 = 49, and 60 multiplied by 0.8 = 48. It can be seen that the first intermediate attention scores for the three images tend to be averaged.

[0071] It should be understood that the above representation of the initial attention score as an integer and the adjustment weight as a decimal is merely an example and does not constitute a limitation on this embodiment.

[0072] So far, the above embodiments have focused on describing the process of initially correcting the initial attention scores of each image based on the positional attributes of each image. In the following embodiments, the process of secondarily correcting the first intermediate attention scores based on content attributes will be described.

[0073] In this embodiment, the relevance between the visual information corresponding to at least two images and the target text information is used as a content attribute. The visual information of the images refers to various visual content that can be extracted and understood from the images. For example, this could be color information, relative positional relationships between objects (e.g., objects or people), relative positional relationships between objects and space, etc. The target text information refers to the text information associated with at least a portion of the at least two images. For example, all image-associated text information can be used as the target text information, or only a portion of the image-associated text information can be selected. Optionally, when selecting only a portion of the image-associated text information as the target text information, the method of image selection is not limited. For example, images can be selected as needed based on their importance, priority, or content. The specific information dimensions used for image selection can be determined according to the application scenario.

[0074] In addition, content attributes can also be described by the image's theme content or main content. The theme content refers to the main idea or emotion expressed by the image; the main content refers to the main objects in the image, such as portraits, landscapes, or buildings.

[0075] Furthermore, in this optional embodiment, the first intermediate attention scores corresponding to each of the at least two images are further corrected based on the content attributes of the at least two images to obtain multiple target attention scores.

[0076] The secondary correction mainly involves adjusting the first intermediate attention score of an image based on contextual information provided by content attributes, thereby increasing the attention given to visual information that is highly relevant to the content attributes. For example, if the correlation between target text information and visual information is positively correlated, then if the visual information of an image is highly relevant to its content attributes, the first intermediate attention score of that image will be adaptively increased; conversely, if the visual information is poorly relevant to its content attributes, the first intermediate attention score will be adaptively decreased.

[0077] When the content attribute is implemented as topic content, if at least two images have a high repetition rate of a certain topic content, the image belonging to that topic content can be assigned relatively more attention; and the image not belonging to that topic content can be assigned relatively less attention. In some embodiments, the topic content can be preset or obtained through semantic recognition, without limitation. Based on this, when the content attribute is implemented as topic content, if at least two images have a high repetition rate of a certain topic content, the first intermediate attention score of the image belonging to that topic content is adaptively increased; and for the image not belonging to that topic content, the first intermediate attention score of that image is adaptively decreased.

[0078] When the content attribute is implemented as the main content, if a certain main content appears frequently in at least two images, relatively more attention can be allocated to the image containing that main content; and relatively less attention can be allocated to the image that does not contain that main content. In some embodiments, the main content can be preset or obtained through image recognition, without limitation. Based on this, when the content attribute is implemented as the main content, if a certain main content appears frequently in at least two images, the first intermediate attention score of the image containing that main content is adaptively increased; and the first intermediate attention score of the image that does not contain that main content is adaptively decreased.

[0079] When the relevance between the visual information of at least two images and the target text information is used as a content attribute, and taking a positive correlation between the target text information and the visual information as an example, this positive correlation is merely an example and does not constitute a limitation of this embodiment. If the visual information is highly relevant to the target text information, then the visual information should receive relatively more attention; conversely, if the visual information is less relevant to the target text information, then the visual information should receive relatively less attention. Specifically, when the content attribute is implemented as the relevance between visual information and the target text information, the secondary correction mainly involves adjusting the first intermediate attention score corresponding to the image based on the contextual information provided by the target text information, thereby increasing the attention of visual information that is highly relevant to the target text information. Taking a positive correlation between the target text information and the visual information as an example, if the visual information of an image is highly relevant to the target text information, then the first intermediate attention score of that image will be adaptively increased; if the visual information of an image is less relevant to the target text information, then the first intermediate attention score of that image will be adaptively decreased.

[0080] In an optional embodiment, the input tag sequence further includes: multiple text tags, where the aforementioned Token can be used as an example of a text tag. These text tags are obtained by preprocessing the text information associated with at least two images, such as linear embedding and positional encoding. The multiple text tags are distributed among multiple visual tags, with the distribution location related to the position of the image associated with the text information, and each image is associated with at least one text tag. Based on the multiple text tags, in the process of using the relevance between the visual information corresponding to each of the at least two images and the target text information as a content attribute to perform secondary correction on the first intermediate state attention scores corresponding to each of the at least two images to obtain multiple target attention scores, at least one target text tag is selected from the multiple text tags in the input tag sequence, and the at least one target text tag corresponds to the target text information. This embodiment does not limit the method of selecting at least one target text tag from the multiple text tags. For example, at least one text tag can be randomly selected as the target text tag; or, it can be selected based on attributes such as the importance, priority, and position of the text tag; or, it can be selected according to preset rules. Preset rules include, but are not limited to, keyword matching, character limit, etc.

[0081] Furthermore, for any one of the at least two images, attention is calculated based on at least one target text marker and a visual marker generated from that image to obtain a second intermediate attention score corresponding to that image. Optionally, for any visual marker generated from any image, the sum of the attention scores of the visual marker and at least one target text marker is calculated to obtain a second intermediate attention score corresponding to that visual marker; the second intermediate attention scores corresponding to each visual marker generated from that image are used as the second intermediate attention score corresponding to that image.

[0082] The following is one way to calculate the second intermediate state attention score, as shown in formula (6). Formula (6) is only an example of calculating the second intermediate state attention score and is not limited thereto.

[0083] As shown in formula (6) above, This represents the self-attention score of the second intermediate state corresponding to the nth image. It is a vector that includes the second intermediate self-attention scores corresponding to each visual tag generated from the nth image. In this formula (6), T represents the text tag, n represents the nth image, its value ranges from 1 to N, N represents the total number of images, and N is an integer greater than or equal to 2; This represents the vector formed by the visual markers in the nth image. It is the index of the visual marker in the nth image. Let i represent the Mth text tag in at least one target text tag. In formula (6), at least one target text tag is implemented as M text tags, where M is an integer greater than or equal to 1; t Represents an index sequence of M text tags; L represents the attention score between each visual marker and the i-th text marker in the n-th image. t M represents the length of the text tag. In the above formula (6), for each visual tag generated in the nth image, the sum of the attention scores of the visual tag and the M document tags is calculated as the second intermediate attention score of the visual tag; then, the second intermediate attention scores of each visual tag generated in the nth image form a vector as the second intermediate attention score corresponding to the nth image.

[0084] Furthermore, a third numerical calculation is performed on the first intermediate state attention score and the second intermediate state attention score corresponding to any image to obtain the target attention score corresponding to that image.

[0085] Following the example above, in The formula for calculating the attention score of the first intermediate state for any image, and In the case of the formula for calculating the second intermediate state attention score of the image, an example of a third numerical calculation is given, as shown in the following formula (7):

[0086] As shown in formula (7), Let α represent the target attention score. α is used to adjust the guidance of the text markers in the secondary correction of the first intermediate attention score corresponding to the nth image. Specifically, it involves a weighted sum of the first and second intermediate attention scores to obtain the target attention score for the nth image. The target attention score for the nth image is also a vector, including the target attention scores of each visual marker generated from the nth image. Furthermore, summing the target attention scores of each visual marker generated from at least two images yields multiple target attention scores.

[0087] In the above embodiments, the correction process was described in detail using the example of simultaneously correcting the initial attention score based on both location and content attributes. The following is a brief description of the process of correcting the initial attention score based solely on location or content attributes. Optionally, the process of correcting the initial attention score based solely on location attributes includes: using the initial attention scores corresponding to at least two images as location attributes, and correcting the initial attention scores corresponding to at least two images to obtain target attention scores corresponding to at least two images. The process of using the initial attention scores corresponding to at least two images as location attributes and correcting the initial attention scores corresponding to at least two images can be found in the detailed description of the initial correction process in the aforementioned embodiments. The only difference from the aforementioned embodiments is that the attention score obtained after correction is directly used as the target attention score, and is no longer the first intermediate state attention score. Alternatively, the process of correcting the initial attention score based solely on content attributes includes: using the relevance between the visual information and the target text information corresponding to at least two images as content attributes, and correcting the initial attention scores corresponding to at least two images to obtain target attention scores corresponding to at least two images. The process of using the relevance between the visual information of at least two images and the target text information as content attributes to correct the initial attention scores of at least two images can be found in the detailed description of the secondary correction process in the foregoing embodiments. The only difference from the foregoing embodiments is that the attention score to be corrected is no longer the first intermediate state attention score, but the initial attention score.

[0088] Furthermore, based on multiple target attention scores, multiple initial key-value pairs are pruned to obtain target key-value pairs that need to be cached. After obtaining the target key-value pairs, they can be cached, while non-target key-value pairs are not cached, thus saving storage resources consumed by caching target key-value pairs, such as video memory. In this embodiment, the method of pruning multiple initial key-value pairs based on multiple target attention scores is not limited.

[0089] In an optional embodiment, during the pruning process of multiple initial key values ​​based on multiple target attention scores, key values ​​and non-key values ​​can be determined from the multiple initial key values ​​based on the multiple target attention scores. This embodiment does not limit the method for determining key values ​​and non-key values. Furthermore, the number of key values ​​and non-key values ​​is not limited; there can be one or more key values ​​or non-key values. Optionally, the multiple target attention scores corresponding to multiple visual tags are sorted from high to low, and the initial key values ​​of the first number of target attention scores in the top column can be selected as key values, with the other key values ​​designated as non-key values. The first number can be an empirical value or an arbitrary value. Alternatively, a first attention score threshold can be set, and initial key values ​​with target attention scores greater than the first attention score threshold can be selected as key values; the other key values ​​are designated as non-key values. In an optional embodiment, non-key values ​​can be pruned (or trimmed), and the key values ​​can be used as target key values.

[0090] Furthermore, considering that the cropping operation discards a large number of non-critical key values, there is a risk of losing important visual information. To mitigate this situation, alternatively, after cropping the non-critical key values, a recycling process can be performed on these non-critical key values ​​to improve the utilization rate of visual information.

[0091] In this optional embodiment, non-critical key values ​​can be recycled to obtain secondary key values, which are then used as target key values. These secondary key values ​​are both the key values ​​to be recycled and the key values ​​to be cached. The number of secondary key values ​​can be one or more, preferably relatively limited to reduce the final number of target key values ​​to be cached. In this embodiment, the cropped non-critical key values ​​are not directly merged with the critical key values. Instead, these non-critical key values ​​are fused separately and compressed into a relatively limited number of secondary key values ​​for independent caching. This preserves important visual information while decoupling important components (such as critical key values) and secondary components (such as non-critical key values) in the image, facilitating the differentiation between important and secondary components in the image during subsequent processing based on the cached target key values ​​(e.g., decoding).

[0092] In this embodiment, the method of fusing non-critical key values ​​is not limited. Any method that can fuse these non-critical key values ​​into a relatively limited number of secondary key values ​​is applicable to this embodiment. In an optional embodiment, non-critical key values ​​can be fused in groups, and the detailed implementation process is as follows.

[0093] Optionally, non-key values ​​are divided into at least two groups, each group including at least two non-key values.

[0094] In this embodiment, the method of dividing the at least two groups is not limited. In one optional embodiment, the division can be arbitrary, as long as it divides the data into at least two groups, and each group includes at least two non-keyword key values. In another optional embodiment, based on the target attention scores corresponding to the non-keyword key values, at least two reference key values ​​are selected from the non-keyword key values. For example, the non-keyword key values ​​of the second number of target attention scores in the top column can be used as reference key values; alternatively, a second attention score threshold can be set, and key values ​​with target attention scores greater than the second attention score threshold can be selected as reference key values. Each reference key value corresponds to one group, and at least two reference key values ​​correspond to at least two groups.

[0095] Furthermore, based on the similarity between other non-key keys and at least two reference keys, other non-key keys can be assigned to corresponding groups to complete the division into at least two groups. The method of assigning other non-key keys to corresponding groups is not limited. For example, other non-key keys can be assigned to groups whose similarity to the reference keys is greater than or equal to a first similarity threshold. Alternatively, a clustering algorithm can be used to automatically determine at least two groups based on the similarity between other non-key keys and the reference keys.

[0096] Furthermore, after dividing the image into at least two groups, the non-critical key values ​​in each of the at least two groups are fused separately to obtain at least two secondary key values. These non-critical key values ​​are then fused separately and compressed into a relatively limited number of secondary key values ​​for independent caching. This process preserves important visual information while decoupling important components (such as critical key values) and secondary components (such as non-critical key values) in the image, facilitating the differentiation between important and secondary components during subsequent processing (e.g., decoding) based on the cached target key values. Optionally, for any group, the average of the non-critical key values ​​in that group is calculated as the secondary key value, or a weighted sum of the non-critical key values ​​in that group is performed to obtain the secondary key value.

[0097] Furthermore, key and secondary key values ​​are selected as target key values ​​to be cached.

[0098] In the above embodiments, pruning and recycling of non-critical key values ​​are described to compress key values ​​that need to be cached. Since the compressed target key values ​​occupy less GPU memory, memory access overhead can be reduced, thereby improving the inference speed of the model. In some cases, the visual pruning method provided in this disclosure can achieve higher accuracy at the same compression rate; and achieve a higher compression ratio at the same accuracy.

[0099] The following example illustrates how to prune multiple initial key values ​​based on multiple target attention scores to obtain the target key values ​​to be cached, as described in the above embodiments.

[0100] As shown in Figure 2, each rectangle represents a visual marker, and each visual marker corresponds to an initial key value and a target attention score. In Figure 2, rectangles marked with an "×" indicate that the initial key value of the visual marker to which that rectangle belongs has been clipped. Conversely, rectangles without an "×" indicate that the initial key value of the visual marker to which that rectangle belongs has been preserved.

[0101] In the following figures, there will be instances where rectangular frames represent visual markers, which will not be explained again.

[0102] As shown in Figure 2, taking an input tag sequence consisting of 8 visual tags as an example, each visual tag corresponds to an initial key value.

[0103] Each visual marker has a corresponding target attention score. In this embodiment, the clipping of the initial key value corresponding to each visual marker can be simply referred to as clipping of the initial key value. Further, based on the target attention score corresponding to each visual marker, key values ​​and non-key values ​​are determined from the multiple initial key values ​​corresponding to multiple visual markers. Non-key values ​​and key values ​​each have their own associated visual markers. In the example shown in Figure 2, from these 8 visual markers, sorted from left to right, the first and last are selected as the visual markers corresponding to the key values, and the other visual markers are selected as the visual markers corresponding to the non-key values. Referring to the definition of clipping of the initial key value, the clipping of the non-key values ​​of visual markers whose initial key values ​​belong to non-key values ​​can be simply referred to as clipping of the non-key values.

[0104] Furthermore, the six non-critical key values ​​in Figure 2 are pruned. The rectangles in Figure 2 marked with an "×" indicate pruning of non-critical key values. These six pruned non-critical key values ​​can then be recycled.

[0105] In this example, two empty visual markers are set to store non-critical key values, as shown in Figure 2. The initial key values ​​corresponding to the two visual markers with the highest attention scores of the cropped target are initialized, and these initialized key values ​​are used as reference key values. These reference key values ​​are used to group other non-critical key values. One reference key value corresponds to one group.

[0106] Furthermore, for each of the other non-key keys (i.e., each of the six non-key keys), the similarity to the two reference keys is calculated, and the non-key key is assigned to a group with higher similarity. This assignment method is merely an example and does not constitute a limitation of this embodiment. As shown in Figure 2, the six non-key keys are sorted from left to right, with the first and third assigned to group K. b [0] belongs to the group, the second and fourth are assigned to K. b [1] Group to which it belongs.

[0107] Furthermore, after dividing the data into two groups, the non-critical key values ​​in the two groups are merged to obtain two secondary key values. As shown in Figure 2, for any group, the average value of the non-critical key values ​​(including the reference key value) in that group is calculated as the secondary key value, or the non-critical key values ​​in that group are weighted and summed to obtain the secondary key value.

[0108] You can refer to the following formula for calculation:

[0109] In formula (8), K c [i] is the secondary key value obtained after fusion processing; K b [i] is the reference key value, where i represents the index of the reference key value. For example, in Figure 2, when i is 1, it represents K. b [1] is the reference key value. Accordingly, K c [1] represents K b [1] The secondary key value generated by the group to which it belongs. Further, K e Represents other non-key values ​​besides the reference key, i max [j] indicates a reference key value such as K b [i] K with the highest similarity e The index, as shown in Figure 2, K b [1] In the group to which i belongs, i max Pointing to and indicating reference key values ​​such as K b [1] K with the highest similarity e Then these two K e The index can be represented as i max [2] and i max [4]. ||K e || for K e The length of the key. Formula (8) above means that, for any group, the average value of the non-key values ​​in the group is calculated as the secondary key value.

[0110] Furthermore, as shown in Figure 2, key values ​​such as key values ​​KV1 and KV2, as well as two secondary key values, are used as target key values ​​to be cached.

[0111] In the above embodiments, a visual pruning method is introduced, which is not limited to the model architecture and type applied. The model type applied can be, but is not limited to, multimodal models and unimodal models. Multimodal models include, but are not limited to, visual language models and image generation models; unimodal models include, but are not limited to, visual models, image classification models, object detection models, etc. Any model type capable of processing visual modal data and based on a self-attention mechanism falls within the protection scope of this disclosure. Visual model data includes, but is not limited to, videos and multiple images. The following description uses the application of the visual pruning method to a visual language model as an example, but is not limited thereto.

[0112] This disclosure provides a visual language model comprising multiple decoder layers connected in sequence. As shown in Figure 3a, the visual language model includes multiple decoder layers connected in sequence, such as decoder layer 1 to decoder layer F. Decoder layer 1 is the first decoder layer, and decoder layer F is the last decoder layer. F is greater than 1 and is a natural number. The output label sequence of the preceding decoder layer is the input label sequence of the following decoder layer; the input label sequence of the first decoder is obtained by converting at least two images and corresponding text information; the input label sequence includes multiple visual labels, which are semantic descriptions of multiple image regions generated by the structured decomposition of at least two images, with each image generating at least one visual label.

[0113] As shown in Figure 3a, this embodiment includes four images arranged from left to right. That is, the further to the right an image is, the further back it is in the sequence. Each image is associated with corresponding text information, which is attached to the right side of each image. As shown in Figure 3a, arranged from left to right, the text information associated with the first image is "Step 1: Proceed to counter in front of tissue boxes"; the text information associated with the second image is "Step 2: Pick up middle tissue box from counter"; the text information associated with the third image is "Turn around, go to counter toilet"; and "Current step" refers to the current step.

[0114] In this embodiment, the visual language model generates a complete output through two stages: a pre-filling stage and a decoding stage. The pre-filling stage refers to the process of generating output tag information and a cache of target key values ​​based on the input tag sequence. Specifically, in the pre-filling stage, at each decoder layer, the target key values ​​obtained by pruning the input tag sequence are cached so that these target key values ​​can be reused in the decoding stage without recalculation, thereby improving computational efficiency.

[0115] In the pre-filling stage, if the current decoder layer is the first decoder layer in the visual language model, the input label sequence is generated by preprocessing the image regions in each image in the order of at least two images. If the current decoder layer is not the first decoder layer, the input label sequence of the non-first decoder layer is the output label sequence of the previous decoder layer. If the current decoder layer is the last decoder layer, an output label information is generated, such as the "historically generated label information" in Figure 3a. The output label information can be simply referred to as label information, and the label information can be visual labels or text labels.

[0116] Furthermore, during the decoding phase, for each autoregressive step of the visual language model, the output label information of the previous autoregressive step and the cached target key value of the current decoder layer are used as input to generate the next output label information. The output label sequence of the previous decoder layer serves as the input label sequence of the next decoder layer. If the current decoder layer is the first decoder layer, the input label sequence consists of the output label information generated in the pre-filling phase and the cached target key value of the first decoder layer; if the current decoder layer is not the first decoder layer, the input label sequence consists of the output label sequence of the previous decoder layer and the cached target key value corresponding to the current decoder layer; if the current decoder layer is the last decoder layer, the next output label information is generated, as shown in "Newly Generated Label Information" in Figure 3a. This process continues in a loop until the visual language model generates a complete output.

[0117] In this embodiment, decoder layer 1 is used as an example of the current self-attention sublayer to illustrate the visual pruning processing method provided in this embodiment. It should be understood that this does not constitute a limitation on the embodiments of this disclosure. The visual pruning processing method provided in the embodiments of this disclosure can be applied to any decoder layer.

[0118] In the pre-filling stage, the input tag sequence of decoder layer 1 is first obtained. This input tag sequence includes multiple visual tags and multiple text tags. The multiple visual tags are generated by decomposing the four images (as shown in Figure 3a) in the order they appear. Each image generates five visual tags, and each visual tag corresponds to a local region of a given image. Each image is divided into multiple patches, and each patch is then used for linear embedding and positional encoding to form a visual tag. In this embodiment, the visual tag of each image represents a local region of that image; a local region refers to the individual patches formed after image segmentation, i.e., a specific part of the image, and each local region has a corresponding position in the image, such as a specific image area.

[0119] Multiple text tags are obtained by transforming the text information corresponding to each image, and these text tags are distributed among multiple visual tags in the input tag sequence. Specifically, the text information corresponding to each image generates 3 text tags.

[0120] Taking the first image from left to right in Figure 3a as an example, this image can generate 5 visual tags, and the corresponding text information can be converted into 3 text tags. Similarly, in this example, from left to right, there are 4 images, each corresponding to 5 visual tags and 3 text tags. The visual tags and text tags corresponding to the 4 images can form the input tag sequence of decoder layer 1. The input tag sequence includes 20 visual tags and 12 text tags, with the 12 text tags scattered among the 20 visual tags. These tags are arranged sequentially from left to right, with the rightmost tags indicating their earlier position in the input tag sequence.

[0121] Furthermore, in this embodiment, based on the weight matrix of the current decoder layer 1, contextual information perception processing is performed on multiple visual tags in the input tag sequence to obtain multiple initial key values ​​and multiple initial attention scores corresponding to the multiple visual tags. Each visual tag corresponds to one initial key value and one initial attention score. For details on this step, please refer to the foregoing embodiments. This embodiment focuses on the process of correcting the multiple initial attention scores.

[0122] In this embodiment, a bar chart example is given, showing the average initial attention scores for the four images respectively. As shown in Figure 3b, these represent the average initial attention scores for the four images in Figure 3a. In this embodiment, [the bar chart is used]. This refers to the average initial attention score for each image, where n is the position of the image among the four images. When n=1, This represents the average initial attention score corresponding to the first image. When n takes the values ​​of 2, 3, or 4, the definition for n=1 can be referenced. As shown in Figure 3b, the horizontal axis of the bar chart before correction represents... The vertical axis represents The value of the initial attention score varies; for example, when n is 1, the vertical axis represents the average value of the initial attention score of the first image. As shown in Figure 3b, the later an image is in the sequence of four images, the lower its average initial attention score. This results in a significant cropping of the initial key values ​​of the visual tags generated by later images, leading to a loss of contextual information.

[0123] Therefore, in this embodiment, multiple initial attention scores are corrected at the image level based on the positional and content attributes of the four images.

[0124] First, the initial attention scores corresponding to each of the four images are used as positional attributes to initially adjust the initial attention scores of the four images; the earlier the image is in the frame, the higher its initial attention score. The process of using positional attributes to initially adjust the initial attention scores of the four images is described below.

[0125] For any one of the four images, calculate the average of the initial attention scores for each of the four images to obtain the average attention score for each of the four images; calculate the natural index for the average attention score for each of the four images to obtain the first calculation result for each of the four images; the ratio of the first calculation result for the image to the sum of the first calculation results for each of the four images is taken as the attention percentage of the image in the four images.

[0126] In one example, a formula for calculating the attention percentage of any image can be found in the aforementioned formula (2) or (3). Details of formula (2) or (3) will not be repeated here, but can be found in the relevant description above.

[0127] Furthermore, in this embodiment, a second numerical calculation is performed based on a preset compensation factor and attention ratio to obtain a normalization coefficient.

[0128] Furthermore, based on the normalization coefficient, the initial attention score corresponding to any image is normalized to obtain the first intermediate attention score corresponding to that image. In one example, the normalization method for the initial attention score corresponding to any image can be found in the aforementioned formula (4) or (5). Details regarding formula (4) or (5) will not be repeated here; please refer to the relevant description above.

[0129] Furthermore, the relevance between the visual information of each of the four images and the target text information is used as a content attribute to perform a secondary correction on the first intermediate state attention scores of each of the four images. This process is described below.

[0130] In this embodiment, all text tags in the input tag sequence are used as target text information corresponding to the target text tag, but it is not limited to this.

[0131] For any one of the four images, and for any visual marker generated from that image, calculate the sum of the attention scores of that visual marker and each target text marker in the input marker sequence to obtain the second intermediate attention score corresponding to that visual marker information. Use the second intermediate attention scores corresponding to each visual marker generated from that image as the second intermediate attention score for that image.

[0132] In one example, one way to calculate the second intermediate state attention score for any image is to refer to the aforementioned formula (6). Details of formula (6) will not be repeated here; please refer to the relevant description above.

[0133] In this embodiment, the calculation method for the second intermediate state attention score is combined with that in formula (6). This represents the Mth text tag in 4 images, and a total of 12 text tags; i t This represents an index sequence of 12 text tags; L represents the second intermediate state attention score for the visual markers and 12 text markers in the nth image. t Indicates the length of the text tag. In the above... In the formula, the attention scores of each visual marker generated from the nth image and the 12 document markers are summed to obtain the second intermediate state attention score corresponding to the nth image.

[0134] Furthermore, a third numerical calculation is performed on the first intermediate state attention score and the second intermediate state attention score corresponding to any image to obtain the target attention score corresponding to that image. In one example, one method of calculating the third numerical value can be found in the aforementioned formula (7). Details of formula (7) will not be repeated here, but can be found in the relevant description above.

[0135] Furthermore, combining the calculation method of the target attention score as in formula (7), the first intermediate state attention score and the second intermediate state attention score are weighted and summed to obtain the target attention score corresponding to the nth image.

[0136] In this embodiment, after correcting the initial attention scores corresponding to the four images, the target attention score for each image is obtained. Below is a bar chart example of the average target attention scores for the four images.

[0137] As shown in Figure 3b, in this embodiment, using This refers to the average target attention score for each image, where n is the position of the image among the four images. When n=1, This represents the average target attention score corresponding to the first image. When n takes the values ​​of 2, 3, or 4, the definition for n=1 can be referenced. As shown in Figure 3b, the horizontal axis of the corrected histogram represents... The vertical axis represents The vertical axis represents the average level of the target attention score in the first image. For example, when n is 1, the vertical axis represents the average level of the target attention score in the first image. As shown in Figure 3b, among the four images... The result after correction The averaged approach reduces the bias caused by assigning initial attention scores to images at different locations, and more accurately indicates the attention paid to different regions by decoder layer 1 when processing images.

[0138] Furthermore, based on multiple target attention scores, multiple initial key values ​​are pruned to obtain the target key values ​​to be cached. First, based on the target attention score corresponding to each initial key value, key values ​​and non-key values ​​are determined from multiple initial key values. The determination method can be referred to the above embodiment and will not be repeated here. After determining the non-key values, the non-key values ​​are pruned, as shown in Figure 3a. "Pruned marker" represents the visual marker to which the pruned non-key value belongs. Here, since the target attention score is a corrected attention score, the bias caused by assigning initial attention scores to images at different locations is reduced, and the attention of decoder layer 1 to different regions is more accurately indicated when processing images. Furthermore, during the pruning process, important components in each image can be preserved, so that important contextual information (i.e., target key values) in each image can be cached, thereby protecting the integrity of visual information to a greater extent.

[0139] After pruning non-critical key values, these non-critical key values ​​can be recycled to obtain secondary key values, which are then used as target key values. Preferably, the number of secondary key values ​​is relatively limited to reduce the final number of target key values ​​that need to be cached.

[0140] During the recycling process for non-critical key values, the non-critical key values ​​are divided into at least two groups, with each group containing at least two non-critical key values. During this division, at least two reference key values ​​can be selected from the non-critical key values ​​based on their corresponding target attention scores. For example, one or more non-critical key values ​​with the highest target attention scores can be selected as reference key values. Other non-critical key values ​​are then assigned to their corresponding groups based on their similarity to the at least two reference key values, thus completing the division into at least two groups. For example, other non-critical key values ​​can be assigned to groups with higher similarity to the at least two reference key values.

[0141] Furthermore, after dividing the data into at least two groups, the non-critical key values ​​in each of the at least two groups are fused to obtain at least two secondary key values. In this embodiment, for any group, the average of the non-critical key values ​​in that group is calculated as the secondary key value, or a weighted sum of the non-critical key values ​​in that group is performed to obtain the secondary key value. For example, the "recycled marker" shown in Figure 3a represents the visual marker to which the secondary key value belongs.

[0142] Furthermore, key and secondary key values ​​are used as target key values ​​to be cached. Then, the output label sequence of the current self-attention sublayer can be used as the input label sequence of the next decoder layer for pruning. The pruning process can refer to the pruning process of the current decoder layer to obtain the target key value corresponding to the next decoder layer. After pruning through F sequentially connected decoder layers, the target key values ​​are retained, realizing the pruning and recycling of the initial key values, thereby compressing the initial key values. During the decoding stage, the visual language model can read the compressed target key values ​​from the cache for attention calculation. Since the compressed target key values ​​occupy less GPU memory, memory access overhead can be reduced, thereby improving the inference speed of the decoding stage. In some cases, such as during the decoding stage, the multimodal model can require only 20% of the target key value, which reduces the initial key value memory usage by 80% (from 1.5GB to 0.3GB), and achieves a 1.5x increase in inference speed, for example, from 28ms / token to 19ms / token. At the same precision, the inference speed is improved due to the higher compression ratio.

[0143] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.

[0144] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 101 to 103 can be device A; or the execution subject of steps 101 and 102 can be device A, and the execution subject of step 103 can be device B; and so on.

[0145] Furthermore, some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0146] Figure 4 is a schematic diagram of the structure of an electronic device provided in another exemplary embodiment of this disclosure. As shown in Figure 4, the device includes a memory 44 and a processor 45.

[0147] Memory 44 is used to store computer programs and can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0148] Processor 45, coupled to memory 44, is configured to execute a computer program in memory 44 for: acquiring an input tag sequence for the current self-attention sublayer, the input tag sequence including multiple visual tags, the multiple visual tags being semantic descriptions of multiple image regions generated by structured decomposition of at least two images, with each image generating at least two visual tags; acquiring multiple initial key values ​​and multiple initial attention scores corresponding to the multiple visual tags; correcting the multiple initial attention scores based on attribute information of the at least two images to obtain multiple target attention scores, the attribute information reflecting the importance of the at least two images in the attention score correction process; and pruning the multiple initial key values ​​based on the multiple target attention scores to obtain target key values ​​to be cached.

[0149] In an optional embodiment, when the processor 45 corrects the plurality of initial attention scores based on the attribute information of the at least two images to obtain a plurality of target attention scores, it is specifically configured to: correct the plurality of initial attention scores at the image level based on the position attributes and / or content attributes of the at least two images to obtain the plurality of target attention scores.

[0150] In an optional embodiment, the processor 45 corrects the plurality of initial attention scores at the image level based on the positional and content attributes of the at least two images to obtain the plurality of target attention scores. Specifically, this involves: using the initial attention scores corresponding to each of the at least two images as positional attributes to perform initial correction on the initial attention scores corresponding to each of the at least two images to obtain a first intermediate attention score corresponding to each of the at least two images; wherein, in the at least two images, the earlier the image is in position, the larger the initial attention score corresponding to that image; using the relevance between the visual information corresponding to each of the at least two images and the target text information as content attributes to perform secondary correction on the first intermediate attention scores corresponding to each of the at least two images to obtain the plurality of target attention scores; the target text information refers to the text information associated with at least a portion of the images in the at least two images.

[0151] In an optional embodiment, when the processor 45 uses the initial attention scores corresponding to each of the at least two images as position attributes to perform initial correction on the initial attention scores corresponding to each of the at least two images to obtain the first intermediate attention scores corresponding to each of the at least two images, it specifically performs the following: for any one of the at least two images, it performs normalization processing on the initial attention score corresponding to any one of the at least two images based on the initial attention scores corresponding to each of the at least two images to obtain the first intermediate attention score corresponding to any one of the images.

[0152] In an optional embodiment, when the processor 45 normalizes the initial attention score corresponding to any image based on the initial attention scores corresponding to the at least two images respectively to obtain a first intermediate attention score corresponding to the any image, it specifically performs the following steps: calculating the attention percentage of the any image in the at least two images based on the initial attention score corresponding to the any image and the initial attention scores corresponding to the at least two images respectively; and normalizing the initial attention score corresponding to the any image based on the attention percentage to obtain a first intermediate attention score corresponding to the any image.

[0153] In an optional embodiment, when the processor 45 calculates the attention percentage of any image in the at least two images based on the initial attention score corresponding to any image and the initial attention scores corresponding to each of the at least two images, it specifically performs the following steps: calculating the average of the initial attention scores corresponding to each of the at least two images to obtain the average attention score corresponding to each of the at least two images; performing a first numerical calculation on the average attention score corresponding to each of the at least two images to obtain a first calculation result corresponding to each of the at least two images; and using the ratio of the first calculation result corresponding to any image to the sum of the first calculation results corresponding to each of the at least two images as the attention percentage of any image in the at least two images.

[0154] In an optional embodiment, when the processor 45 normalizes the initial attention score corresponding to any image based on the attention ratio to obtain the first intermediate attention score corresponding to any image, it specifically performs the following: performs a second numerical calculation based on a preset compensation factor and the attention ratio to obtain a normalization coefficient; and performs normalization processing on the initial attention score corresponding to any image based on the normalization coefficient to obtain the first intermediate attention score corresponding to any image.

[0155] In an optional embodiment, the input tag sequence further includes: multiple text tags, which are scattered among multiple visual tags, and each image is associated with at least one text tag; when the processor 45 uses the relevance between the visual information corresponding to each of the at least two images and the target text information as a content attribute to perform a secondary correction on the first intermediate attention score corresponding to each of the at least two images to obtain the multiple target attention scores, it specifically performs the following: selecting at least one target text tag from the multiple text tags, where the at least one target text tag corresponds to the target text information; for any image among the at least two images, performing attention calculation based on the at least one target text tag and the visual tag generated by the image to obtain a second intermediate attention score corresponding to the image; and performing a third numerical calculation on the first intermediate attention score and the second intermediate attention score corresponding to the image to obtain a target attention score corresponding to the image.

[0156] In an optional embodiment, when the processor 45 performs attention calculation based on the at least one target text tag and the visual tag generated from any image to obtain a second intermediate attention score corresponding to any image, it is specifically configured to: calculate the sum of the attention scores of any visual tag and the at least one target text tag for any visual tag generated from any image to obtain a second intermediate attention score corresponding to the visual tag information; and use the second intermediate attention scores corresponding to each visual tag generated from any image as the second intermediate attention score corresponding to any image.

[0157] In an optional embodiment, the processor 45 performs pruning on the plurality of initial key values ​​based on the plurality of target attention scores to obtain target key values ​​that need to be cached. Specifically, this is done by: determining key values ​​and non-key values ​​from the plurality of initial key values ​​based on the plurality of target attention scores; recycling the non-key values ​​to obtain secondary key values; and using the key values ​​and the secondary key values ​​as the target key values.

[0158] In an optional embodiment, when the processor 45 performs recycling processing on the non-critical key values ​​to obtain secondary key values, it specifically performs the following: dividing the non-critical key values ​​into at least two groups, each group including at least two non-critical key values; and performing fusion processing on the non-critical key values ​​in the at least two groups respectively to obtain at least two secondary key values.

[0159] In an optional embodiment, when the processor 45 divides the non-key key values ​​into at least two groups, it specifically performs the following: selects at least two reference key values ​​from the non-key key values ​​based on the target attention scores corresponding to the non-key key values, with each reference key value corresponding to one group; and assigns the other unselected non-key key values ​​to the corresponding groups based on the similarity between the other unselected non-key key values ​​and the at least two reference key values.

[0160] In an optional embodiment, when the processor 45 performs fusion processing on the non-critical key values ​​in the at least two groups to obtain at least two secondary key values, it is specifically configured to: calculate the average value of the non-critical key values ​​in the group as the secondary key value for any group, or perform a weighted summation on the non-critical key values ​​in the group to obtain the secondary key value.

[0161] In an optional embodiment, the above method is applied to a visual language model, which includes multiple self-attention sub-layers; in the pre-filling stage, the processor 45 uses the multiple self-attention sub-layers to execute the steps in the various methods provided in the embodiments of this disclosure to obtain a target key value; and in the decoding stage, the processor 45 uses the multiple self-attention sub-layers to perform attention calculation based on the cached target key value to obtain an output result.

[0162] Furthermore, as shown in Figure 4, the electronic device also includes other components such as a communication component 46, a display 47, a power supply component 48, and an audio component 49. Figure 4 only schematically shows some components and does not imply that the electronic device only includes the components shown in Figure 4. Additionally, the components within the dashed boxes in Figure 4 are optional, not mandatory, and their specific inclusion depends on the product form of the electronic device. The electronic device of this embodiment can be implemented as a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device, or as a server-side device such as a conventional server, cloud server, or server array. If the electronic device of this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include the components within the dashed boxes in Figure 4; if the electronic device of this embodiment is implemented as a server-side device such as a conventional server, cloud server, or server array, it may not include the components within the dashed boxes in Figure 4.

[0163] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0164] The aforementioned communication components are configured to facilitate wired or wireless communication between the device containing the communication components and other devices. The device containing the communication components can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication components receive broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication components also include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), and other technologies.

[0165] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0166] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.

[0167] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0168] Accordingly, this disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium.

[0169] Accordingly, this disclosure also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above-described method embodiments. It should be understood that each step or combination of steps in the above-described method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above-described method embodiments.

[0170] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0171] The above are merely embodiments of this disclosure and are not intended to limit the scope of this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of the claims of this disclosure.

Claims

1. A visual marker pruning method, characterized in that, include: Obtain the input label sequence of the current self-attention sublayer. The input label sequence includes multiple visual labels. The multiple visual labels are semantic descriptions of multiple image regions generated by the structured decomposition of at least two images. Each image generates at least two visual labels. Obtain multiple initial key values ​​and multiple initial attention scores corresponding to the multiple visual tags; Based on the attribute information of the at least two images, the multiple initial attention scores are corrected to obtain multiple target attention scores, wherein the attribute information reflects the importance of the at least two images in the attention score correction process; Based on the multiple target attention scores, the multiple initial key values ​​are pruned to obtain the target key values ​​that need to be cached.

2. The method according to claim 1, characterized in that, Based on the attribute information of the at least two images, the multiple initial attention scores are corrected to obtain multiple target attention scores, including: Based on the positional and / or content attributes of the at least two images, the plurality of initial attention scores are corrected at the image level to obtain the plurality of target attention scores.

3. The method according to claim 2, characterized in that, Based on the positional and content attributes of the at least two images, the initial attention scores are corrected at the image level to obtain the target attention scores, including: The initial attention scores corresponding to each of the at least two images are used as positional attributes to perform initial corrections on the initial attention scores corresponding to each of the at least two images, so as to obtain the first intermediate state attention scores corresponding to each of the at least two images; wherein, in the at least two images, the earlier the image is, the larger the initial attention score corresponding to the image. The correlation between the visual information corresponding to each of the at least two images and the target text information is used as a content attribute. The first intermediate state attention scores corresponding to each of the at least two images are then modified a second time to obtain the multiple target attention scores. The target text information refers to the text information associated with at least some of the images in the at least two images.

4. The method according to claim 3, characterized in that, Using the initial attention scores corresponding to each of the at least two images as positional attributes, the initial attention scores corresponding to each of the at least two images are initially corrected to obtain the first intermediate state attention scores corresponding to each of the at least two images, including: For any one of the at least two images, the initial attention score corresponding to any one image is normalized according to the initial attention scores corresponding to each of the at least two images to obtain the first intermediate attention score corresponding to any one image.

5. The method according to claim 4, characterized in that, Based on the initial attention scores corresponding to each of the at least two images, the initial attention score corresponding to any one image is normalized to obtain the first intermediate attention score corresponding to any one image, including: Based on the initial attention score corresponding to any one image and the initial attention scores corresponding to each of the at least two images, calculate the attention percentage of any one image in the at least two images; Based on the attention percentage, the initial attention score corresponding to any image is normalized to obtain the first intermediate attention score corresponding to any image.

6. The method according to claim 5, characterized in that, Based on the initial attention score corresponding to any one image and the initial attention scores corresponding to each of the at least two images, calculate the attention percentage of any one image in the at least two images, including: Calculate the average of the initial attention scores corresponding to each of the at least two images to obtain the average attention score corresponding to each of the at least two images; The average attention score corresponding to each of the at least two images is calculated first to obtain the first calculation result corresponding to each of the at least two images; The ratio of the first calculation result corresponding to any image to the sum of the first calculation results corresponding to each of the at least two images is taken as the attention percentage of any image in the at least two images.

7. The method according to claim 5, characterized in that, Based on the attention percentage, the initial attention score corresponding to any image is normalized to obtain the first intermediate attention score corresponding to any image, including: A second numerical calculation is performed based on the preset compensation factor and the attention ratio to obtain the normalization coefficient; Based on the normalization coefficient, the initial attention score corresponding to any image is normalized to obtain the first intermediate attention score corresponding to any image.

8. The method according to claim 3, characterized in that, The input tag sequence further includes: multiple text tags, which are scattered among the multiple visual tags, and each image is associated with at least one text tag; Using the relevance between the visual information corresponding to each of the at least two images and the target text information as a content attribute, a second correction is made to the first intermediate state attention score corresponding to each of the at least two images to obtain the plurality of target attention scores, including: Select at least one target text tag from the plurality of text tags, wherein the at least one target text tag corresponds to the target text information; For any one of the at least two images, attention is calculated based on the at least one target text tag and the visual tag generated by the image to obtain the second intermediate attention score corresponding to the image. A third numerical calculation is performed on the first intermediate attention score and the second intermediate attention score corresponding to any image to obtain the target attention score corresponding to any image.

9. The method according to claim 8, characterized in that, Attention is calculated based on the at least one target text tag and the visual tag generated from any one of the images to obtain a second intermediate attention score corresponding to any one image, including: For any visual marker generated from any image, the sum of the attention scores of the visual marker and the at least one target text marker is calculated to obtain the second intermediate state attention score corresponding to the visual marker. The second intermediate attention score corresponding to each visual marker generated for any image is taken as the second intermediate attention score for any image.

10. The method according to any one of claims 1-9, characterized in that, Based on the multiple target attention scores, the multiple initial key values ​​are pruned to obtain the target key values ​​to be cached, including: Based on the plurality of target attention scores, key and non-key key values ​​are determined from the plurality of initial key values; The non-critical key values ​​are recycled to obtain secondary key values; The key value and the secondary key value are used as the target key value.

11. The method according to claim 10, characterized in that, The non-critical key values ​​are recycled to obtain secondary key values, including: The non-key values ​​are divided into at least two groups, and each group includes at least two non-key values. The non-critical key values ​​in the at least two groups are fused to obtain at least two secondary key values.

12. The method according to claim 11, characterized in that, The non-critical key values ​​are divided into at least two groups, including: Based on the target attention score corresponding to the non-critical key value, at least two reference key values ​​are selected from the non-critical key values, and one reference key value corresponds to one group. Based on the similarity between the other unselected non-key key values ​​and the at least two reference key values, the other unselected non-key key values ​​are divided into corresponding groups.

13. The method according to claim 11, characterized in that, The non-critical key values ​​in the at least two groups are fused to obtain at least two secondary key values, including: For any group, the average of the non-critical key values ​​in the group is calculated as the secondary key value, or the non-critical key values ​​in the group are weighted and summed to obtain the secondary key value.

14. The method according to claim 10, characterized in that, Based on the plurality of target attention scores, key and non-key values ​​are determined from the plurality of initial key values, including: Sort the multiple target attention scores from high to low; Select the initial key values ​​of the first number of target attention scores in the top column as key values, and the other key values ​​as non-key values.

15. The method according to claim 10, characterized in that, The method further includes: After identifying non-critical key values ​​and before recycling them, non-critical key values ​​are pruned. After pruning non-critical key values, the non-critical key values ​​are then recycled.

16. The method according to any one of claims 1-9, characterized in that, The method is applied to a visual language model, which includes multiple self-attention sublayers; During the pre-filling stage, the method steps of weight 1 are executed using the multiple self-attention sub-layers to obtain the target key value; During the decoding phase, attention calculations are performed using the multiple self-attention sub-layers based on the cached target key values ​​to obtain the output results.

17. A visual language model, characterized in that, include: Multiple decoder layers are connected sequentially, and the output tag sequence of the previous decoder layer is the input tag sequence of the next decoder layer; The input tag sequence includes multiple visual tags, which are semantic descriptions of multiple image regions generated by the structured decomposition of at least two images, with each image generating at least one visual tag. The plurality of decoder layers are configured to perform the steps of the method according to any one of claims 1-14 during the pre-filling stage to obtain the target key value; In the decoding stage, attention is calculated based on the cached target key values ​​to obtain new labeling information.

18. An electronic device, characterized in that, The method includes a memory and a processor, the memory being used to store a computer program, and the processor being coupled to the memory for executing the computer program to implement the steps of the method according to any one of claims 1-16.

19. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the processor is enabled to perform the steps of the method according to any one of claims 1-16.

20. A computer program product, characterized in that, include: A computer program / instruction that, when executed by a processor, causes the processor to perform the steps of the method according to any one of claims 1-16.