Key value caching method and device, equipment, storage medium and product
By fusion of features and scene graphs of images and texts of multimodal large language model, the problems of wasted key-value cache memory and low inference efficiency in multimodal models are solved, and more efficient memory usage and inference performance are achieved.
Patent Information
- Application Number
- CN202510032637.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
The existing key-value caching technology in multimodal large language models leads to memory waste and inferred inference efficiency due to visual modal information redundancy.
The images and text in the hints of the multimodal large language model are encoded through pre-trained image encoder and visual encoder, a scene diagram of the image hint is constructed, the complete feature representation of the visual object is determined, and the image and text features are fused to obtain multimodal hint vector encoding, and finally it is key-value cached.
It effectively deletes a large amount of redundant visual information in the image mode, reduces the length of image encoding, avoids memory waste in key-value cache, and improves the inference efficiency of multimodal large language models.
Smart Images

Figure CN119941879A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data storage technology, and in particular to a key-value caching method, device, equipment, storage medium and product. Background Art
[0002] Key-value Cache (KV Cache) is a cache technology that optimizes the inference performance of large models. By storing key-value pairs, the calculation results of attention are reused to improve inference performance and reduce memory consumption. In the process of large-scale autoregressive inference, the output of the previous step is used as the current input, and the self-attention operation is performed. The key-value KV vector needs to be extracted for each word token in the current input. KV Cache caches the KV pairs of historical input sequences. When calculating the attention of new tokens, the attention results of repeated sequences can be directly obtained from the cache, and only the new tokens are focused on, which significantly reduces the amount of repeated calculations and improves the inference speed of the model. However, as the length of the sequence processed by the large model increases, the KV Cache occupancy of video memory also increases rapidly. Especially in multimodal large language models, multimodal KVCache includes multiple related image and text representations, and more computing resources are required to respond to the increasing input length.
[0003] Existing key-value caches still face some challenges. As the length of model processing sequences increases, the memory usage of KV Cache will also increase rapidly, especially in large multimodal models, where data processing requirements for different modalities such as images and texts are higher and occupy more cache space. Although existing solutions have proposed some optimization methods in the training, deployment, and reasoning stages of the model, such as reducing the number of KV vectors by adjusting the attention mechanism in the training stage, expelling and merging unimportant tokens in the reasoning stage, or achieving quantitative compression by reducing storage precision, they cannot effectively solve the high redundancy problem of KV Cache in multimodal models. Summary of the invention
[0004] The main purpose of this application is to provide a key-value cache method, device, equipment, storage medium and product, aiming to solve the technical problems of key-value cache memory waste and low model reasoning efficiency caused by redundancy of visual modal information.
[0005] To achieve the above purpose, the present application proposes a key-value caching method, the method comprising:
[0006] Encode the image and text in the prompt of the preset multimodal large language model according to the pre-trained image encoder and the pre-trained visual encoder respectively to obtain the image prompt encoding and the text prompt encoding;
[0007] constructing a scene graph of the image cue based on the multiple visual objects in each of the image cue codes, and determining a complete feature representation of each of the visual objects based on the scene graph;
[0008] fusing each of the image cue codes with the complete feature representation to obtain a final visual feature representation;
[0009] fusing the final visual feature representation and the textual prompt encoding to obtain a multimodal prompt vector encoding;
[0010] The key-value pairs encoded by the multimodal prompt vector are determined according to the weight matrix of the key layer and the weight matrix of the value layer, and the key-value pairs encoded by the multimodal prompt vector are key-value cached.
[0011] In one embodiment, the step of constructing a scene graph of the image prompt according to the multiple visual objects in each of the image prompt codes comprises:
[0012] Extracting multiple visual objects in each of the image prompt codes, and generating an object node set according to the visual objects;
[0013] Extracting attributes of each of the visual objects, and generating an attribute node set according to the attributes;
[0014] Extracting the relationship between each of the visual objects, and generating an edge set between object nodes according to the relationship;
[0015] Determining an adjacency matrix of each of the image prompt codes according to the edge set;
[0016] A scene graph of an image prompt is constructed according to the object node set, the attribute node set, the edge set and the adjacency matrix.
[0017] In one embodiment, the step of determining a complete feature representation of each of the visual objects according to the scene graph comprises:
[0018] Determine attribute features, edge features, and neighbor node features of each object node according to the scene graph prompted by the image;
[0019] The attribute features, the edge features and the neighbor node features of each object node are fused to obtain a complete feature representation of each visual object.
[0020] In one embodiment, the step of fusing the attribute features, the edge features, and the neighbor node features of each object node to obtain a complete feature representation of each visual object includes:
[0021] Determine a similarity value between each object node and each neighbor node feature corresponding to the object node, and perform weighted averaging on the neighbor node features according to the similarity value to obtain an object-level fusion rule;
[0022] Determine the average value of the edge feature and the average value of the attribute feature of each object node, and obtain edge-level fusion rules and attribute-level fusion rules respectively;
[0023] The object-level fusion rule, the edge-level fusion rule and the attribute-level fusion rule are subjected to feature fusion in combination with preset parameters to obtain a complete feature representation of each visual object.
[0024] In one embodiment, the step of fusing each of the image hint codes with the complete feature representation to obtain a final visual feature representation comprises:
[0025] encoding each of the image prompts as an original image feature, determining a cosine similarity between the original image feature and the complete feature representation of each visual object, and obtaining a similarity matrix;
[0026] Determining the most relevant visual object corresponding to each of the original image features according to the similarity matrix;
[0027] Determine, according to the similarity matrix and the most relevant visual object, a maximum similarity set corresponding to a plurality of most relevant target original image features corresponding to each of the visual objects;
[0028] An average value of the maximum similarity set of each of the visual objects is determined to obtain a final visual feature representation of each of the visual objects.
[0029] In one embodiment, after the step of caching the key-value pairs encoded by the multimodal hint vector, the method further includes:
[0030] The cache time in the key-value cache determines the cumulative attention score in the original text key-value pair before the most recent preset time period;
[0031] Eliminate the remaining text key-value pairs other than the preset number of text key-value pairs with the highest cumulative attention scores in the original text key-value pairs to obtain the eliminated key-value pairs;
[0032] Determine a similarity matrix between the removed key-value pairs and the remaining text key-value pairs;
[0033] Determine, from the eliminated key-value pairs, the word element with the greatest similarity to each word element in the remaining text key-value pairs according to the remaining text key-value pairs and the similarity matrix;
[0034] Determine the maximum similarity set of each word in the removed key-value pairs according to the similarity matrix and the word with the greatest similarity;
[0035] The mean of each word and other words in the removed key-value pair is determined respectively, and the final mean of the mean and each word is determined, and each word in the key-value cache is updated to the final mean.
[0036] In addition, to achieve the above purpose, the present application also proposes a key-value cache device, which includes:
[0037] A prompt encoding module, used to encode the image and text in the prompt of the preset multimodal large language model according to the pre-trained image encoder and the pre-trained visual encoder, respectively, to obtain the image prompt encoding and the text prompt encoding;
[0038] a feature determination module, configured to construct a scene graph of the image prompt according to the plurality of visual objects in each of the image prompt codes, and determine a complete feature representation of each of the visual objects according to the scene graph;
[0039] A feature fusion module, used for fusing each of the image prompt codes with the complete feature representation to obtain a final visual feature representation;
[0040] A coding fusion module, used for fusing the final visual feature representation and the text prompt coding to obtain a multimodal prompt vector coding;
[0041] The key-value caching module is used to determine the key-value pairs encoded by the multimodal prompt vector according to the weight matrix of the key layer and the weight matrix of the value layer, and to perform key-value caching on the key-value pairs encoded by the multimodal prompt vector.
[0042] In addition, to achieve the above-mentioned purpose, the present application also proposes a key-value caching device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the key-value caching method described above.
[0043] In addition, to achieve the above objectives, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the key-value caching method described above are implemented.
[0044] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the key-value caching method described above are implemented.
[0045] The present application provides a key-value caching method, which encodes the image and text in the prompt of a preset multimodal large language model respectively, constructs a scene graph of image prompts according to multiple visual objects in each image prompt code, and determines the complete feature representation of each visual object according to the scene graph; fuses each image prompt code and the complete feature representation, fuses the final visual feature representation obtained with the text prompt code, and obtains a multimodal prompt vector code; determines the key-value pairs of the multimodal prompt vector code according to the weight matrix of the key layer and the weight matrix of the value layer, and key-value caches the key-value pairs of the multimodal prompt vector code. Since the present application constructs a scene graph of image prompts, and then realizes the fusion of the original features of the image and the features of the scene graph objects according to the scene graph, it effectively deletes a large amount of redundant visual information in the image modality and reduces the length of the image code, thereby avoiding memory waste in the key-value cache and improving the reasoning efficiency of the multimodal large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0048] Figure 1 A flowchart of the key-value cache method embodiment 1 of the present application is provided;
[0049] Figure 2 A schematic diagram of the process of multimodal prompt encoding in the key-value caching method of this application;
[0050] Figure 3 This is a schematic diagram of an example process of feature fusion in the key-value cache method of this application;
[0051] Figure 4 A flowchart of the second embodiment of the key-value cache method of the present application is provided;
[0052] Figure 5 A schematic diagram of the process of constructing a scene graph in the key-value cache method of this application;
[0053] Figure 6 A flowchart of the third embodiment of the key-value cache method of the present application is provided;
[0054] Figure 7 This is a flow chart of KV pair elimination in the key-value cache method of this application;
[0055] Figure 8 This is a flow chart of KV pair merging in the key-value caching method of this application;
[0056] Fig. 9 This is a schematic diagram of the module structure of the key-value cache device according to an embodiment of the present application;
[0057] Fig.10 Schematic diagram of the device structure of the hardware operating environment involved in the key-value caching method in the embodiment of the present application.
[0058] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0059] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0060] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0061] The main solution of the embodiment of the present application is: according to the pre-trained image encoder and the pre-trained visual encoder, the image and text in the prompt of the preset multimodal large language model are respectively encoded to obtain the image prompt code and the text prompt code; according to the multiple visual objects in each of the image prompt codes, a scene graph of the image prompt is constructed, and the complete feature representation of each of the visual objects is determined according to the scene graph; each of the image prompt codes and the complete feature representation are merged to obtain the final visual feature representation; the final visual feature representation is merged with the text prompt code to obtain the multimodal prompt vector code; according to the weight matrix of the key layer and the weight matrix of the value layer, the key-value pairs of the multimodal prompt vector code are respectively determined, and the key-value pairs of the multimodal prompt vector code are key-value cached.
[0062] The existing key-value cache still faces some challenges. As the length of the model processing sequence increases, the memory usage of KV Cache will also increase rapidly, especially in large multimodal models, where the data processing requirements of different modalities such as images and texts are higher and the cache space occupied is also larger. Although the existing solutions have proposed some optimization methods in the training, deployment and reasoning stages of the model, such as reducing the number of KV vectors by adjusting the attention mechanism in the training stage, expelling and merging unimportant tokens in the reasoning stage, or reducing the storage precision to achieve quantitative compression, they cannot effectively solve the high redundancy problem of KV Cache in multimodal models.
[0063] The present application provides a solution, by respectively encoding the image and text in the prompt of the preset multimodal large language model, constructing a scene graph of the image prompt according to the multiple visual objects in each image prompt code, and determining the complete feature representation of each visual object according to the scene graph; fusing each image prompt code and the complete feature representation, fusing the obtained final visual feature representation and the text prompt code, and obtaining a multimodal prompt vector code; determining the key-value pairs of the multimodal prompt vector code according to the weight matrix of the key layer and the weight matrix of the value layer, and key-value caching the key-value pairs of the multimodal prompt vector code. Since the present application constructs a scene graph of image prompts, and then realizes the fusion of the original features of the image and the features of the scene graph objects according to the scene graph, it effectively deletes a large amount of redundant visual information in the image modality and reduces the length of the image code, thereby avoiding memory waste in the key-value cache and improving the reasoning efficiency of the multimodal large language model.
[0064] It should be noted that the execution subject of the method of this embodiment can be a computing service device with key-value caching, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc.; it can also be a key-value caching device with the same or similar functions. This embodiment and the following embodiments will be described by taking the key-value caching device as an example.
[0065] Based on this, the embodiment of the present application provides a key-value caching method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the key-value caching method of the present application.
[0066] In this embodiment, the key-value caching method includes steps S10 to S50:
[0067] Step S10, encoding the image and text in the prompt of the preset multimodal large language model according to the pre-trained image encoder and the pre-trained visual encoder respectively, to obtain the image prompt code and the text prompt code.
[0068] It should be noted that the prompt of the multimodal large language model contains a set of related images and texts. The trained image encoder and visual encoder are used to encode the images and texts in the prompts respectively to obtain the vector encoding of the multimodal prompt (MPrompt) Where X T and X I Represent the Embedding Representation (ER) of the prompt text and prompt image respectively, M and N represent the number of image tokens and text tokens respectively. The process of multimodal prompt encoding can be referred to Figure 2 shown.
[0069] Step S20, constructing a scene graph of the image prompt according to the multiple visual objects in each of the image prompt codes, and determining a complete feature representation of each of the visual objects according to the scene graph.
[0070] It can be understood that for image prompts, the target detection module is used to extract multiple visual objects (VO) in each image separately, and generate corresponding n region-level visual features to obtain Corresponding to the n visual objects in the kth image. The scene graph generation tool is used to extract the properties of the visual objects and the relationship between the visual objects, and a scene graph (SG) of the image prompt is constructed. Then, the complete feature representation of each visual object is determined according to the scene graph, denoted as F = {f1, f2, ..., f n},f i =h′ i .
[0071] Step S30, fusing each of the image hint codes with the complete feature representation to obtain a final visual feature representation.
[0072] It can be understood that for each image hint, the original feature representation X obtained by the image decoder is I-k , the representation of the visual object obtained from the scene graph F k , and the two are combined to obtain the final visual feature representation. The Visual Scene Graph (VSG) can intelligently extract the core objects in the image and their corresponding relationships, without having to store every pixel or irrelevant visual features in the image. This compression method significantly reduces the space occupied by the cache while maintaining the key information of the image, especially when processing high-resolution images or multiple pictures, which can fully relieve the pressure on the video memory.
[0073] In a feasible implementation, step S30 may include steps S301 to S304:
[0074] Step S301: encode each of the image prompts as original image features, determine the cosine similarity between the original image features and the complete feature representation of each visual object, and obtain a similarity matrix.
[0075] It can be understood that for each image hint, the original feature representation X obtained by the image decoder is I-k , and the representation of the visual object obtained from the scene graph F k , calculate the cosine similarity between the two and get the similarity matrix S.
[0076] Step S302: determining the most relevant visual object corresponding to each of the original image features according to the similarity matrix.
[0077] Step S303: determining a maximum similarity set corresponding to a plurality of most relevant target original image features corresponding to each of the visual objects according to the similarity matrix and the most relevant visual objects.
[0078] Step S304: determine the average value of the maximum similarity set of each of the visual objects to obtain a final visual feature representation of each of the visual objects.
[0079] It is worth noting that, according to the similarity matrix S, each original feature of the image can be matched to the most relevant visual object in the scene graph. i There may be multiple matching image original features, and these original features and f are calculated i The final visual feature representation of the image prompt is obtained by averaging the values of
[0080] You can refer to Figure 3 The feature fusion process is illustrated with an example. There are 4 visual objects in the scene graph. According to the similarity matrix, the most similar visual object can be found for each original feature of the image. Among them, object a is the object closest to feature 1 and feature 3, object b is the object closest to feature 2, feature 4 and feature 5, and object d is the object closest to feature 6. Taking object a as an example, the maximum similarity set of a, that is, features 1 and 3, is averaged and fused to obtain the final visual feature representation of object a. Similarly, the final visual feature representations of objects b, c, and d in the scene graph can be obtained.
[0081] Step S40: fusing the final visual feature representation and the text prompt code to obtain a multimodal prompt vector code.
[0082] It can be understood that the final multimodal prompt vector encoding is obtained by combining the embedded representation of the text prompt in is the embedding representation of the text prompt, It is a representation that combines the original features of the image and the features of the scene graph objects.
[0083] Step S50 , determining the key-value pairs encoded by the multimodal prompt vector according to the weight matrix of the key layer and the weight matrix of the value layer respectively, and performing key-value caching on the key-value pairs encoded by the multimodal prompt vector.
[0084] It can be understood that the KV vector (i.e., key-value pair) of the multimodal prompt H is finally calculated, and the key-value pair encoded by the multimodal prompt vector is cached. According to the formula K=HW K ,V=HWV , where K represents the key vector and V represents the value vector. The KV tensor of the multimodal cue encoding H can be calculated, where W K ,W V Represent the weight matrices of the key layer and value layer respectively.
[0085] The present embodiment provides a key-value caching method, which encodes the image and text in the prompt of a preset multimodal large language model respectively, constructs a scene graph of the image prompt according to multiple visual objects in each image prompt code, and determines the complete feature representation of each visual object according to the scene graph; fuses each image prompt code and the complete feature representation, fuses the final visual feature representation obtained with the text prompt code, and obtains a multimodal prompt vector code; determines the key-value pairs of the multimodal prompt vector code according to the weight matrix of the key layer and the weight matrix of the value layer, and key-value caches the key-value pairs of the multimodal prompt vector code. Since the present embodiment constructs a scene graph of the image prompt, and then realizes the fusion of the original image features and the scene graph object features according to the scene graph, a large amount of redundant visual information in the image modality is effectively deleted, and the length of the image code is reduced, thereby avoiding the waste of memory and computing resource consumption in the key-value cache during the reasoning process, and improving the reasoning efficiency of the multimodal large language model.
[0086] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can refer to the above introduction, and will not be repeated later. Figure 4 , step S20, the key-value caching method further includes steps S201 to S205:
[0087] Step S201, extracting multiple visual objects in each of the image prompt codes, and generating an object node set based on the visual objects.
[0088] It can be understood that the target detection module is used to extract multiple visual objects in each image separately, and generate corresponding n region-level visual features to obtain Corresponding to the n visual objects in the k-th image, a set of object nodes is generated according to the visual objects.
[0089] Step S202: extracting the attributes of each of the visual objects, and generating an attribute node set according to the attributes.
[0090] Step S203, extracting the relationship between each of the visual objects, and generating an edge set between object nodes according to the relationship.
[0091] It should be understood that a scene graph generation tool may be used to extract the attributes of visual objects and the relationships between visual objects, and generate a set of attribute nodes and a set of edges between object nodes.
[0092] Step S204: determining an adjacency matrix of each of the image prompt codes according to the edge set.
[0093] It can be understood that the adjacency matrix M of each image prompt encoding is determined according to the relationship between each visual object in the edge set. k .
[0094] Step S205 , constructing a scene graph of image prompts according to the object node set, the attribute node set, the edge set and the adjacency matrix.
[0095] Here is an example to illustrate the process of building a scene graph. You can refer to Figure 5 The scene graph in the figure has four visual objects, namely object a, object b, object c and object d. Object a has attributes ① and ②, and its neighbor nodes are objects b, c and d, connected by edges 1, 2 and 3 respectively. Finally, a scene graph G is generated for each image. k = {O k ,E k ,A k ,M k},in represents the set of visual objects in the k-th image, Represents the set of edges between object nodes, Represents a collection of object attributes, M k It is the adjacency matrix, which represents the connection relationship between object nodes in the image.
[0096] In a feasible implementation manner, after step S205, steps S206 to S207 may also be included:
[0097] Step S206, determining the attribute features, edge features and neighbor node features of each object node according to the scene graph prompted by the image.
[0098] Step S207, fusing the attribute features, the edge features and the neighbor node features of each object node to obtain a complete feature representation of each visual object.
[0099] It should be understood that the scene graph is a heterogeneous graph containing object nodes, attribute nodes and edges. Any object node in the scene graph has a set of neighbor nodes, edge sets and attribute sets. In order to obtain a complete feature representation of each visual object in the scene graph, it is necessary to extract and fuse the attribute features, edge features and neighbor node features of each object.
[0100] In this embodiment, the attribute features, edge features and neighbor node features of each object node are determined according to the scene graph prompted by the image, and the attribute features, edge features and neighbor node features of each object node are fused to obtain a complete feature representation of each visual object. By fusing the attribute features, edge features and neighbor node features of each object, a complete feature representation of each visual object in the scene graph is obtained.
[0101] In a feasible implementation, step S207 may include steps S2071 to S2073:
[0102] Step S2071, determining the similarity value between each object node and each neighbor node feature corresponding to the object node, and performing weighted averaging on the neighbor node features according to the similarity value to obtain an object-level fusion rule.
[0103] Step S2072, determining the average value of the edge feature and the average value of the attribute feature of each object node, and obtaining edge-level fusion rules and attribute-level fusion rules respectively.
[0104] Step S2073, performing feature fusion on the object-level fusion rule, the edge-level fusion rule, and the attribute-level fusion rule in combination with preset parameters to obtain a complete feature representation of each visual object.
[0105] It can be understood that the initial feature of object node i is h i , use the attention mechanism to combine the information that the node needs to pay attention to in the three layers, and calculate the fusion rules of the neighbor object level, edge level, and attribute level of node i respectively to obtain h′ iobj , h′ iedg , h′ iatt .
[0106] Here we will explain it with a specific example. Take object node a as an example. The neighbor nodes of a are b, c, and d. The connected edges are 1, 2, and 3. The attributes are ① and ②. First, calculate the fusion rule of object node a at the object level, calculate the similarity between node a and its neighbor nodes b, c, and d, and perform weighted average of the neighbor nodes according to the similarity to obtain the object-level fusion rule h′ iobj ; Secondly, calculate the edge-level and attribute-level fusion rules, average all edge features and attribute features of node a, and obtain the edge-level and attribute-level fusion rules h′ iedg and h′ iatt Finally, the following formula is used to perform three-level feature fusion:
[0107]
[0108] Among them, α and β are trainable parameters, and finally the complete feature representation of each object in the scene graph is obtained, which is recorded as F = {f1, f2, ..., f n}, f i =h′ i .
[0109] In this embodiment, the attention mechanism is used to combine the information that the node needs to pay attention to at the three layers, and the fusion rules of the node's neighbor object level, edge level, and attribute level are calculated respectively, and finally a complete feature representation of each object in the scene graph is obtained.
[0110] In this embodiment, multiple visual objects in the image cue encoding are generated into object nodes, attribute nodes and edges, and a scene graph is constructed according to the relationship between the object nodes. The visual scene graph can intelligently extract the core objects in the image and their corresponding relationships, without the need to store every pixel or irrelevant visual feature in the image.
[0111] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above description, and will not be described in detail later. Figure 6 After step S50, the key-value caching method further includes steps A10 to A60:
[0112] Step A10, determining a cumulative attention score in the original text key-value pairs whose cache time in the key-value cache is before the most recent preset time period.
[0113] It is understandable that after the key-value cache, in order to further reduce memory usage and improve the collaborative processing capability of multimodal information, the KV Cache content is updated. This process is divided into two sub-steps: image-first KV pair elimination and KV pair merging.
[0114] First, when the KV Cache exceeds the maximum capacity, some KV pairs stored in the KV Cache need to be removed. Since the visual information has been integrated before entering the KV Cache, only the text KV pairs with lower importance are removed here. First, the KV pairs corresponding to the M prompts that appeared in the recent period and all image KV pairs are retained, and the Accumulated Attention Score (AAS) is calculated for all the text KV pairs remaining in the KV Cache.
[0115] Step A20, removing the remaining text key-value pairs other than the preset number of text key-value pairs with the highest cumulative attention scores in the original text key-value pairs, to obtain the removed key-value pairs.
[0116] It should be understood that the N text KV pairs with the highest cumulative attention scores are removed and recorded as K c . Eliminate the remaining KV pairs with lower attention scores. The remaining text key-value pairs that need to be eliminated are denoted as K e =KK c The KV pair elimination process can be found in Figure 7 shown.
[0117] Step A30, determining a similarity matrix between the removed key-value pairs and the remaining text key-value pairs.
[0118] It can be understood that given the set K of keys of the removed prompts e =KK c , using the many-to-one nearest neighbor matching algorithm to derive K e With K c The similarity matrix between , and shares the similarity matrix with the set V of values.
[0119] Step A40, determining, from the eliminated key-value pairs, a word element having the greatest similarity to each word element in the remaining text key-value pairs according to the remaining text key-value pairs and the similarity matrix.
[0120] Step A50, determining the maximum similarity set of each word in the removed key-value pairs according to the similarity matrix and the word with the greatest similarity.
[0121] Understandably, K e Each word-unit token in can find K according to the similarity matrix c The token with the greatest similarity to itself, and K c Each k i There is also a maximum similarity set K sim .
[0122] Step A60, respectively determine the mean of each word and other words in the removed key-value pair, determine the final mean of the mean and each word, and update each word in the key-value cache to the final mean.
[0123] Here we will explain the process of KV pair merging with examples. Figure 8 In the figure, K c In the example, the maximum similarity set of token1 is token4 and token6, the maximum similarity set of token2 is token7, and the maximum similarity set of token3 is token5. Calculate k respectively. i With K sim The mean of each token in k iTake token 1 in the figure as an example, calculate the average of token 1 and each token in the maximum similarity set, that is, calculate the average of token 1 and token 4, and calculate the average of token 1 and token 6, and then calculate the average of these averages with token 1 itself, so as to update the retained token 1. Similarly, update the retained set K c The V vector in the KV Cache no longer calculates the similarity matrix separately, but directly shares the similarity matrix with the K vector, and updates the V vector at the same time as the K vector is updated.
[0124] In this embodiment, by accumulating attention scores and dynamically calculating the correlation between text and image KV pairs, text KV pairs with higher similarity are merged into image KV pairs, thereby achieving cross-modal KV cache compression. This method not only reduces the storage redundancy of text information, but also enhances the connection between text and images, improving the efficiency of multimodal reasoning. Compared with existing single-modality optimization schemes, the proposed merging algorithm can more efficiently utilize multimodal information, especially in large-model reasoning tasks that need to process images and texts at the same time, which can significantly improve reasoning performance while reducing memory usage.
[0125] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the key-value caching method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0126] This application also provides a key-value cache device, please refer to Fig. 9 , the key-value cache device comprises:
[0127] A prompt encoding module 10 is used to encode the image and text in the prompt of the preset multimodal large language model according to a pre-trained image encoder and a pre-trained visual encoder, respectively, to obtain an image prompt code and a text prompt code;
[0128] A feature determination module 20, configured to construct a scene graph of the image cue according to the plurality of visual objects in each of the image cue codes, and determine a complete feature representation of each of the visual objects according to the scene graph;
[0129] A feature fusion module 30, used for fusing each of the image hint codes with the complete feature representation to obtain a final visual feature representation;
[0130] A coding fusion module 40, configured to fuse the final visual feature representation with the text prompt coding to obtain a multimodal prompt vector coding;
[0131] The key-value caching module 50 is used to determine the key-value pairs encoded by the multimodal prompt vector according to the weight matrix of the key layer and the weight matrix of the value layer, and to perform key-value caching on the key-value pairs encoded by the multimodal prompt vector.
[0132] The key-value cache device provided by the present application adopts the key-value cache method in the above embodiment to solve the technical problem. Compared with the prior art, the beneficial effects of the key-value cache device provided by the present application are the same as the beneficial effects of the key-value cache method provided by the above embodiment, and other technical features in the key-value cache device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0133] The present application provides a key-value cache device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the key-value cache method in the above-mentioned embodiment one.
[0134] Reference below Fig.10 , which shows a schematic diagram of the structure of a key-value cache device suitable for implementing an embodiment of the present application. The key-value cache device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig.10 The key-value cache device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0135] like Fig.10As shown, the key-value cache device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the key-value cache device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the key-value cache device to communicate with other devices wirelessly or by wire to exchange data. Although the key-value cache device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0136] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0137] The key-value cache device provided by the present application adopts the key-value cache method in the above embodiment to solve the technical problems of key-value cache. Compared with the prior art, the beneficial effects of the key-value cache device provided by the present application are the same as the beneficial effects of the key-value cache method provided by the above embodiment, and other technical features in the key-value cache device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0138] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0139] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0140] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the key-value caching method in the above-mentioned embodiment.
[0141] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0142] The computer-readable storage medium may be included in the key-value cache device; or may exist independently without being assembled into the key-value cache device.
[0143] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0144] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0145] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0146] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned key-value caching method, and can solve technical problems. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the key-value caching method provided in the above-mentioned embodiment, and will not be repeated here.
[0147] The present application also provides a computer program product, including a computer program, which implements the steps of the key-value caching method as described above when executed by a processor.
[0148] The computer program product provided by the present application can solve the technical problem. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the key-value cache method provided by the above embodiment, which will not be described in detail here.
[0149] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A key-value caching method, characterized in that: The method includes: Encode the image and text in the prompt of the preset multimodal large language model according to the pre-trained image encoder and the pre-trained visual encoder respectively to obtain the image prompt encoding and the text prompt encoding; constructing a scene graph of the image cue based on the multiple visual objects in each of the image cue codes, and determining a complete feature representation of each of the visual objects based on the scene graph; fusing each of the image cue codes with the complete feature representation to obtain a final visual feature representation; fusing the final visual feature representation and the textual prompt encoding to obtain a multimodal prompt vector encoding; The key-value pairs encoded by the multimodal prompt vector are determined according to the weight matrix of the key layer and the weight matrix of the value layer, and the key-value pairs encoded by the multimodal prompt vector are key-value cached.
2. The method according to claim 1, characterized in that The step of constructing a scene graph of the image prompt according to the multiple visual objects in each of the image prompt codes comprises: Extracting multiple visual objects in each of the image prompt codes, and generating an object node set according to the visual objects; Extracting attributes of each of the visual objects, and generating an attribute node set according to the attributes; Extracting the relationship between each of the visual objects, and generating an edge set between object nodes according to the relationship; Determining an adjacency matrix of each of the image prompt codes according to the edge set; A scene graph of an image prompt is constructed according to the object node set, the attribute node set, the edge set and the adjacency matrix.
3. The method according to claim 2, characterized in that The step of determining a complete feature representation of each visual object according to the scene graph comprises: Determine attribute features, edge features, and neighbor node features of each object node according to the scene graph prompted by the image; The attribute features, the edge features and the neighbor node features of each object node are fused to obtain a complete feature representation of each visual object.
4. The method according to claim 3, characterized in that The step of fusing the attribute features, the edge features and the neighbor node features of each object node to obtain a complete feature representation of each visual object includes: Determine a similarity value between each object node and each neighbor node feature corresponding to the object node, and perform weighted averaging on the neighbor node features according to the similarity value to obtain an object-level fusion rule; Determine the average value of the edge feature and the average value of the attribute feature of each object node, and obtain edge-level fusion rules and attribute-level fusion rules respectively; The object-level fusion rule, the edge-level fusion rule and the attribute-level fusion rule are subjected to feature fusion in combination with preset parameters to obtain a complete feature representation of each visual object.
5. The method according to claim 1, characterized in that The step of fusing each of the image hint codes with the complete feature representation to obtain a final visual feature representation comprises: encoding each of the image prompts as an original image feature, determining a cosine similarity between the original image feature and the complete feature representation of each visual object, and obtaining a similarity matrix; Determining the most relevant visual object corresponding to each of the original image features according to the similarity matrix; Determine, according to the similarity matrix and the most relevant visual object, a maximum similarity set corresponding to a plurality of most relevant target original image features corresponding to each of the visual objects; An average value of the maximum similarity set of each of the visual objects is determined to obtain a final visual feature representation of each of the visual objects.
6. The method according to claim 1, characterized in that After the step of caching the key-value pairs encoded by the multimodal prompt vector, the method further includes: The cache time in the key-value cache determines the cumulative attention score in the original text key-value pair before the most recent preset time period; Eliminate the remaining text key-value pairs other than the preset number of text key-value pairs with the highest cumulative attention scores in the original text key-value pairs to obtain the eliminated key-value pairs; Determine a similarity matrix between the removed key-value pairs and the remaining text key-value pairs; Determine, from the eliminated key-value pairs, the word element with the greatest similarity to each word element in the remaining text key-value pairs according to the remaining text key-value pairs and the similarity matrix; Determine the maximum similarity set of each word in the removed key-value pairs according to the similarity matrix and the word with the greatest similarity; The mean of each word and other words in the removed key-value pair is determined respectively, and the final mean of the mean and each word is determined, and each word in the key-value cache is updated to the final mean.
7. A key-value cache device, characterized in that: The key-value cache device comprises: A prompt encoding module, used to encode the image and text in the prompt of the preset multimodal large language model according to the pre-trained image encoder and the pre-trained visual encoder, respectively, to obtain the image prompt encoding and the text prompt encoding; a feature determination module, configured to construct a scene graph of the image prompt according to the plurality of visual objects in each of the image prompt codes, and determine a complete feature representation of each of the visual objects according to the scene graph; A feature fusion module, used for fusing each of the image prompt codes with the complete feature representation to obtain a final visual feature representation; A coding fusion module, used for fusing the final visual feature representation and the text prompt coding to obtain a multimodal prompt vector coding; The key-value caching module is used to determine the key-value pairs encoded by the multimodal prompt vector according to the weight matrix of the key layer and the weight matrix of the value layer, and to perform key-value caching on the key-value pairs encoded by the multimodal prompt vector.
8. A key-value cache device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the key-value caching method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the key-value caching method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the key-value cache method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
KV cache compression and eviction lexical element recovery method and system for large-scale language model reasoning
CN120975245A
Multi-modal parallel reasoning method, electronic equipment and program product
CN121071824A
Multi-modal data processing method and electronic equipment
CN121505378A
An adaptive kv cache compression method, system and device for an audio-text multimodal large model
CN122511270A