Data processing method and system, electronic equipment, storage medium and computer program product
By evaluating the importance of visual marking and text marking of visual language models and cache compression, the problem of large and low-efficiency key-value cache overhead when dealing with large amounts of visual markings is solved, and more efficient resource management and inference speed is achieved.
Patent Information
- Application Number
- CN202510602367.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When existing visual language models deal with a large number of visual markers and generate growth outputs, there are problems such as high storage cost and low update efficiency of key-value caches, resulting in slow model inference speed and unreasonable resource management.
By obtaining multiple visual markers, text markers and initial key value data of the visual language model, using these data to determine the key text markers, and evaluating the importance of the visual markers based on the key text markers, cross-modal attention weight distribution is obtained. Then, the initial key-value data is cached and compressed to generate the target key-value data.
While saving key-value cache overhead, it ensures the efficiency of the model to perform subsequent inference based on target key-value data, and improves the inference speed and resource management efficiency of the visual language model.
Smart Images

Figure CN120123992A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of large model technology and key-value caching technology. Specifically, it relates to a data processing method, system, electronic device, storage medium, and computer program product. Background Art
[0002] In the application field of large visual language models (LVLMs), with the improvement of model complexity and the enhancement of visual information processing capabilities, the demand for computing resources during the model's reasoning process has increased significantly, especially in the processing of visual tokens.
[0003] Currently, application scenarios of image and video understanding require LVLMs to be able to quickly generate long-context outputs. For example, in tasks such as visual question answering and image text generation, LVLMs need to process high-resolution images and video sequences, which poses higher requirements for the decoding efficiency and memory usage efficiency of LVLMs.
[0004] However, when existing LVLMs process a large number of visual tokens and generate long outputs, there may be problems such as high storage costs and low update efficiency of key-value caches. LVLMs usually involve key-value cache data in multiple modalities (such as text modality and visual modality), which leads to high cache consumption of key-value data and may result in contention for memory bandwidth. This not only affects the reasoning speed of LVLMs but also limits the application scope of LVLMs in long-content understanding and generation tasks.
[0005] As can be seen from the above, existing visual language models lack effective key-value data cache management strategies when facing high visual information density and long text generation requirements, resulting in unreasonable resource allocation and low model efficiency. Especially when processing high-frame-rate videos and high-resolution images, the performance and efficiency of LVLMs are severely challenged.
[0006] For the above problems, no effective solutions have been proposed yet. Summary of the Invention
[0007] Embodiments of this application provide a data processing method, system, electronic device, storage medium, and computer program product to at least solve the technical problem in related technologies that the key-value data cache overhead of visual language models is large and affects the model's reasoning efficiency.
[0008] According to one aspect of the embodiments of the present application, a data processing method is provided, including: obtaining a plurality of visual tokens, a plurality of text tokens, and initial key-value data corresponding to a vision-language model; using the initial key-value data to determine key text tokens among the plurality of text tokens; using the initial key-value data and the distribution positions of the key text tokens among the plurality of text tokens to evaluate the importance of the plurality of visual tokens, obtaining an evaluation result, where the evaluation result is used to characterize the cross-modal attention weight distribution between the plurality of visual tokens and the key text tokens; and performing cache compression processing on the initial key-value data according to the evaluation result to obtain target key-value data.
[0009] According to one aspect of the embodiments of the present application, a data processing method is provided, including: obtaining multi-modal input data; extracting, from the target key-value data, the key-value data to be reused corresponding to the multi-modal input data; and performing inference calculation based on the multi-modal input data and the key-value data to be reused to generate a target answer; where the target key-value data is obtained by using the data processing method of any one of the above in the historical inference rounds of the vision-language model.
[0010] According to one aspect of the embodiments of the present application, a data processing method is provided, including: obtaining a data processing request through a first application programming interface, where the request data carried in the data processing request includes: multi-modal input data; and returning a data processing response through a second application programming interface, where the response data carried in the data processing response includes: a target answer, the target answer is generated by performing inference calculation based on the multi-modal input data and the key-value data to be reused, the key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by using the data processing method of any one of the above in the historical inference rounds of the vision-language model.
[0011] According to one aspect of the embodiments of the present application, a data processing method is provided, including: obtaining a current input data processing dialogue request, where the request data carried in the data processing dialogue request includes: multi-modal input data; in response to the data processing dialogue request, returning a data processing dialogue reply, where the information carried in the data processing dialogue reply includes: a target answer, the target answer is generated by performing inference calculation based on the multi-modal input data and the key-value data to be reused, the key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by using the data processing method of any one of the above in the historical inference rounds of the vision-language model; and displaying the target answer in a graphical user interface.
[0012] According to one aspect of the embodiments of the present application, a data processing method is provided, including: responding to an input instruction acting on an operation interface, and displaying multi-modal input data on the operation interface; responding to a processing instruction acting on the operation interface, and displaying a target answer on the operation interface; wherein, the target answer is generated by performing inference calculation based on the multi-modal input data and the key-value data to be reused, and the key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by using the data processing method of any one of the above in the historical inference rounds of the vision-language model.
[0013] According to one aspect of the embodiments of the present application, a data processing system is provided, including: a client for sending multi-modal input data; a server connected to the client for extracting the key-value data to be reused corresponding to the multi-modal input data from the target key-value data, and performing inference calculation based on the multi-modal input data and the key-value data to be reused to generate a target answer; the client is further configured to output the target answer; wherein, the target key-value data is obtained by using the data processing method of any one of the above in the historical inference rounds of the vision-language model.
[0014] According to one aspect of the embodiments of the present application, an electronic device is provided, including: a memory storing an executable program; a processor for running the program, wherein, when the program runs, it executes the data processing method of any one of the above.
[0015] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, the computer-readable storage medium includes a stored executable program, wherein, when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the data processing method of any one of the above.
[0016] According to one aspect of the embodiments of the present application, a computer program product is provided, including a computer program, and the computer program implements the data processing method of any one of the above when executed by a processor.
[0017] In the embodiments of the present application, a plurality of visual tokens, a plurality of text tokens, and initial key-value data corresponding to a vision-language model are obtained; using the initial key-value data, key text tokens among the plurality of text tokens are determined; using the initial key-value data and the distribution positions of the key text tokens among the plurality of text tokens, importance evaluation is performed on the plurality of visual tokens to obtain an evaluation result, where the evaluation result is used to characterize the cross-modal attention weight distribution between the plurality of visual tokens and the key text tokens; according to the evaluation result, cache compression processing is performed on the initial key-value data to obtain target key-value data. Thus, in the embodiments of the present application, an elite observation window corresponding to the key text tokens is constructed, and further, the importance of the visual tokens is evaluated based on the elite observation window, and accordingly, the initial key-value data is cached and compressed. This compression process is related to the importance of the visual tokens, which can save the key-value cache overhead while ensuring the efficiency of the subsequent inference of the model based on the target key-value data. That is to say, the present application achieves the purpose of compressing the key-value data of the vision-language model based on the importance evaluation result of the plurality of visual tokens, thereby realizing the technical effects of saving the key-value data cache overhead of the vision-language model and increasing the inference efficiency of the model based on the target key-value data, and further solving the technical problems of large key-value data cache overhead and affecting the model inference efficiency in the related art.
[0018] It is easy to note that the above general description and the following detailed description are only for exemplifying and explaining the present application, and do not constitute a limitation to the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:
[0020] Figure 1 is a schematic diagram of an application scenario of a data processing method according to an embodiment of the present application;
[0021] Figure 2 is a flowchart of a data processing method according to an embodiment of the present application;
[0022] Figure 3 is a schematic diagram of an optional key-value data processing process according to an embodiment of the present application;
[0023] Figure 4 is a flowchart of another data processing method according to an embodiment of the present application;
[0024] Figure 5 is a flowchart of yet another data processing method according to an embodiment of the present application;
[0025] Figure 6It is a flowchart of another data processing method according to an embodiment of the present application;
[0026] Figure 7 It is a flowchart of another data processing method according to an embodiment of the present application;
[0027] Figure 8 It is a structural block diagram of a data processing device according to an embodiment of the present application;
[0028] Figure 9 It is a structural block diagram of another data processing device according to an embodiment of the present application;
[0029] Figure 10 It is a structural block diagram of another data processing device according to an embodiment of the present application;
[0030] Figure 11 It is a structural block diagram of another data processing device according to an embodiment of the present application;
[0031] Figure 12 It is a structural block diagram of another data processing device according to an embodiment of the present application;
[0032] Figure 13 It is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] The technical solution provided by this application is mainly implemented using large model technology. Here, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one hundred trillion model parameters. A large model can also be called a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one hundred million parameters is produced. Such a model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0036] It should be noted that in actual applications, the pre-trained model can be fine-tuned with a small number of samples so that the large model can be applied to different tasks. For example, large models can be widely applied in the fields of natural language processing (NLP), computer vision, speech processing, etc. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., and can also be widely applied to tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. Therefore, the main application scenarios of large models include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0037] First, some nouns or terms that appear during the description of the embodiments of this application are applicable to the following explanations.
[0038] Visual Language Model (VLM): It refers to an artificial intelligence model that has the ability to process both images and text and is used to understand cross-modal information to achieve functions such as image captioning and visual question answering.
[0039] Key-Value Cache (KV cache): It refers to caching the intermediate calculation results (including the key matrix and value matrix) during the operation of the VLM for reuse in the subsequent model inference process to reduce repeated calculations and accelerate the decoding process of the model.
[0040] According to an embodiment of the present application, a data processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0041] Considering that the number of model parameters of the large model is huge and the computing resources of the mobile terminal are limited, the above method provided by the embodiments of the present application can be applied to Figure 1 the application scenarios shown, but not limited thereto. In the operating environment corresponding to the application scenarios shown Figure 1 a large model is deployed in the server 10. The server 10 can be connected to one or more client devices 20 through a local area network connection, a wide area network connection, an Internet connection, or other types of data networks. Here, the client devices 20 can include but are not limited to: smart phones, tablet computers, laptop computers, palmtop computers, personal computers, smart home devices, in-vehicle devices, etc. The client device 20 can interact with the user through a graphical user interface to realize the invocation of the large model, and further realize the method provided by the embodiments of the present application.
[0042] In the embodiments of the present application, the system composed of the client device and the server can execute the following steps: The client device uploads a plurality of visual tokens, a plurality of text tokens, and initial key-value data corresponding to the vision-language model to the server; the server uses the initial key-value data to determine the key text tokens among the plurality of text tokens, evaluates the importance of the plurality of visual tokens using the initial key-value data and the distribution positions of the key text tokens among the plurality of text tokens to obtain an evaluation result, and performs cache compression processing on the initial key-value data according to the evaluation result to obtain target key-value data, where the evaluation result is used to characterize the cross-modal attention weight distribution between the plurality of visual tokens and the key text tokens. Further, the server returns the target key-value data to the client device.
[0043] It should be noted that with the rapid development of high-performance computing units, in the operating environments of other application scenarios, the above method provided by the embodiments of the present application can also be applied to a model all-in-one machine. In an optional embodiment, a variety of models are built into the model all-in-one machine, and the user can select and adjust one model according to needs to obtain his own model. Thus, the high-performance computing unit built into the model all-in-one machine can directly call the adjusted model to execute the above method provided by the embodiments of the present application. In another optional embodiment, a trained model is built into the large model all-in-one machine. Thus, the high-performance computing unit built into the model all-in-one machine can directly call the model to execute the above method provided by the embodiments of the present application.
[0044] Further, when the user needs to train their own model, they can also upload their own dataset through the client. This dataset is sent from the client to the server, enabling the server to adjust the pre-trained model with this dataset to obtain the user's own model, which is then deployed to the production environment. To facilitate the user's model adjustment requirements, the server can provide complete adjustment tools, development frameworks, and processes, supporting multiple adjustment strategies, so that the adjusted model can better adapt to different field applications and achieve high customization.
[0045] Under the above operating environment, the present application provides a data processing method as Figure 2 shown. Figure 2 It is a flowchart of the data processing method according to an embodiment of the present application. As Figure 2 shown, the method may include the following steps:
[0046] Step S202, obtaining a plurality of visual tokens, a plurality of text tokens, and initial key-value data corresponding to the vision-language model;
[0047] Step S204, using the initial key-value data to determine the key text tokens among the plurality of text tokens;
[0048] Step S206, using the initial key-value data and the distribution positions of the key text tokens among the plurality of text tokens to evaluate the importance of the plurality of visual tokens, obtaining an evaluation result, where the evaluation result is used to represent the cross-modal attention weight distribution between the plurality of visual tokens and the key text tokens;
[0049] Step S208, according to the evaluation result, performing cache compression processing on the initial key-value data to obtain target key-value data.
[0050] In the application scenario, visual tokens and text tokens can be the basic units for the vision-language model to process information. Among them, visual tokens are used to encode the input visual data (such as images or video frames), and text tokens are used to encode the input text prompts. The above-mentioned plurality of visual tokens can be obtained by the vision-language model extracting features from the input visual data. The above-mentioned plurality of text tokens can be obtained by the vision-language model extracting features from the input text prompts.
[0051] The above-mentioned initial key-value data (Original KV Cache) can be the cache data corresponding to the key vector and value vector generated by the vision-language model in the pre-fill stage. The initial key-value data can be obtained by converting the hidden state representation of the vision-language model through a weight matrix, and the initial key-value data can be used for attention calculation in the subsequent decoding stage.
[0052] In some exemplary application scenarios, during the process of determining key text tokens, an elite observation window can also be constructed by using initial key-value data and multiple text tokens, where the elite observation window is used to characterize the distribution positions of key text tokens among the multiple text tokens.
[0053] Based on the running results of the Self-Attention mechanism among multiple text tokens, an Elite Observation Window is constructed. Specifically, based on the initial key-value data, by evaluating the degree of mutual attention among multiple text tokens, which can be characterized by an attention value, key text tokens are filtered out from the multiple text tokens using this attention value. These key text tokens can be some non-consecutive text tokens among the multiple text tokens. Based on the distribution positions of these key text tokens among the multiple text tokens, the above-mentioned elite observation window can be created.
[0054] It is easy to notice that the establishment of the above-mentioned elite observation window does not necessarily depend on all text tokens or a continuous segment of text tokens. Instead, by analyzing the relationships among multiple text tokens, key text tokens that have a key impact on the evaluation of visual tokens are selected, thereby constructing a more appropriate elite observation window for subsequent evaluation of the importance of visual tokens. This strategy for constructing the elite observation window can improve the stability and effectiveness of the evaluation of the importance of visual tokens, providing an optimization basis for the compression strategy of key-value data. That is to say, the above-mentioned key text tokens can be at least one text token among the multiple text tokens that do not have a strictly continuous relationship.
[0055] Furthermore, cross-modal attention calculation is performed between the key text tokens selected from the elite observation window and the visual tokens corresponding to the initial key-value data, so as to achieve the evaluation of the importance of multiple visual tokens. Thus, the evaluation results can reflect the distribution of attention weights when each visual token interacts cross-modally with the key text tokens. The attention weight corresponding to each visual token can characterize the relative importance (or contribution degree) of the visual token in the model inference decision. During the above-mentioned process of importance evaluation, the information interaction between the visual modality and the text modality is considered, ensuring that the visual language model can retain visual information that has a high impact on the final output during cache compression, thereby reducing the loss of model performance caused by cache compression.
[0056] Furthermore, based on the evaluation results, it is possible to determine a part of the visual tokens among the multiple visual tokens that can be safely removed from the initial key-value data. Based on this, when performing cache compression processing (KV CacheCompression) on the initial key-value data, the key-value data removal can be performed for this part of the visual tokens, which can reduce the memory resources occupied by the key-value data cache while ensuring that the model performance will not be significantly affected. In addition, in the application scenario, all text tokens and text key-value data in the text modality can be fully retained; or other text tokens except for the key text tokens among the multiple text tokens can be removed, and the key-value data corresponding to these other text tokens can be synchronously removed.
[0057] After cache compression processing, the target key-value data can include the visual key-value data and text key-value data retained according to the evaluation results, thereby supporting the reuse of the target key-value data by the vision-language model in subsequent inference calculations and improving the model decoding efficiency.
[0058] It is easy to notice that the embodiments of the present application can perform refined management on the KV cache of the vision-language model (especially LVLMs) when processing cross-modal information. The introduction of the elite observation window improves the accuracy and consistency of the visual token importance evaluation, enabling cache compression to more effectively retain key visual information while removing redundant visual information, not only significantly reducing the model's decoding latency and memory occupancy, but also ensuring that the model maintains performance comparable to that under the full cache condition on the compressed KV cache. In particular, when processing high-resolution visual inputs and long text outputs, the above-mentioned solution provided by the embodiments of the present application has more significant advantages, achieving efficiency improvement and resource management optimization in vision-language model inference.
[0059] Through the above steps S202 to S208, the embodiments of the present application obtain multiple visual tokens, multiple text tokens, and initial key-value data corresponding to the vision-language model; use the initial key-value data to determine the key text tokens among the multiple text tokens; use the initial key-value data and the distribution positions of the key text tokens among the multiple text tokens to evaluate the importance of the multiple visual tokens, and obtain an evaluation result, where the evaluation result is used to characterize the cross-modal attention weight distribution between the multiple visual tokens and the key text tokens; according to the evaluation result, perform cache compression processing on the initial key-value data to obtain target key-value data. Thus, the embodiments of the present application construct an elite observation window corresponding to the key text tokens, and further evaluate the importance of the visual tokens based on the elite observation window, and accordingly perform cache compression on the initial key-value data. This compression process is related to the importance of the visual tokens, which can save the key-value cache overhead while ensuring the efficiency of the subsequent inference of the model based on the target key-value data. That is to say, the present application achieves the purpose of compressing the key-value data of the vision-language model based on the importance evaluation result of the multiple visual tokens, thereby realizing the technical effects of saving the key-value data cache overhead of the vision-language model and increasing the inference efficiency of the model based on the target key-value data, and further solving the technical problems of large key-value data cache overhead and affecting the model inference efficiency in the related art.
[0060] In an alternative embodiment, in step S202, obtaining multiple visual tokens, multiple text tokens, and initial key-value data includes the following method steps:
[0061] Step S221, obtaining text prompt data and visual prompt data of the vision-language model;
[0062] Step S222, performing word segmentation conversion on the text prompt data to obtain multiple text tokens;
[0063] Step S223, performing segmentation processing on the visual prompt data to obtain multiple visual tokens;
[0064] Step S224, performing pre-filling calculation based on the multiple visual tokens and the multiple text tokens to obtain initial key-value data.
[0065] In the above alternative embodiment, the text prompt data can be user-readable text instructions or text questions, which are used to guide the processing manner of the vision-language model for the visual prompt data. The visual prompt data can be images, video frames, or other forms of visual inputs that the vision-language model needs to analyze. The above text prompt data and visual prompt data can be included in the multi-modal input data of the vision-language model, enabling the vision-language model to perform inference and generate responses for specific tasks or scenarios.
[0066] Further, the text prompt data undergoes a word segmentation operation in natural language processing and is converted into multiple text tokens. Each text token represents a word or phrase, and text tokens are the basic units (Tokens) for the vision-language model to process text inputs. Through the above analysis and conversion, the text prompt data can be captured and used by the vision-language model in a structured manner, facilitating subsequent processing and the invocation of the attention mechanism.
[0067] Further, the images or video frames in the visual prompt data are segmented into multiple visual tokens. Each visual token can be a part of the image, such as an object. In an application scenario, a vision Transformer network (such as Vision Transformer, abbreviated as ViT) can also be used to convert the pixel data in the above visual prompt data into a vector representation to obtain the above multiple visual tokens. The above segmentation process can help the vision-language model identify and understand the key components in the visual prompt data.
[0068] In the pre-fill stage, the vision-language model performs forward propagation calculations with multiple visual tokens and multiple text tokens as inputs. The pre-fill stage generates key-value pair data, which can serve as the basic data for the key-value data. The key-value pair data can be used to construct the key-value cache of the vision-language model or to update the key-value cache of the vision-language model. The key-value data is an important part of the attention mechanism calculation, helping the vision-language model quickly retrieve context information when processing sequential data, avoiding repeated calculations, and thus accelerating the inference speed. After the pre-fill calculation is completed, the obtained key-value data is saved in the key-value cache to provide support for the subsequent decoding and generation processes of the vision-language model.
[0069] Through the above steps S221 to S224, in the embodiment of the present application, before the vision-language model performs a cross-modal task, it first comprehensively and structurally preprocesses the input data, generating corresponding tokens and initial key-value data for the data of each modality. The above solution not only enables the vision-language model to better integrate visual information and language information, but also provides effective context memory for the subsequent decoding stage. Especially in the inference process of a large-scale vision-language model, the key-value cache constructed in the pre-fill stage is crucial for avoiding redundant calculations and improving the model operation efficiency. Overall, the embodiment of the present application significantly improves the efficiency and performance of the vision-language model in processing complex visual and language tasks through fine input preprocessing and the generation of preliminary key-value data.
[0070] In an optional embodiment, in step S224, pre-fill calculations are performed based on multiple visual tokens and multiple text tokens to obtain initial key-value data, including the following method steps:
[0071] Step S2241: Perform embedding encoding calculation on multiple visual tokens to obtain visual representation data in the hidden state;
[0072] Step S2242: Perform embedding encoding calculation on multiple text tokens to obtain text representation data in the hidden state;
[0073] Step S2243: Perform weight conversion calculation based on the visual representation data and the text representation data to obtain initial key-value data.
[0074] In the above optional embodiment, by performing embedding encoding calculation on multiple visual tokens, each visual token is converted into a vector representation that can be understood by the visual language model, and visual representation data in the hidden state is obtained. By performing embedding encoding calculation on multiple text tokens, each text token is converted into a corresponding vector representation, and text representation data in the hidden state is obtained. The above visual representation data and text representation data are expressed in the hidden space of the visual language model.
[0075] The above embedding encoding calculation on multiple visual tokens can be completed by ViT or a similar structure, converting the pixel information of the input image into hidden states, and these hidden states carry the feature information of the input image, providing a basic representation of the visual modality for subsequent modality interaction. The above embedding encoding calculation on multiple text tokens can be completed by the embedding layer of the visual language model to ensure that the text information can exist in a form compatible with the visual information, facilitating the fusion of cross-modal information.
[0076] Furthermore, through weight conversion calculation, the visual representation data and text representation data in the hidden state are converted into corresponding key-value data. Specifically, multiply the visual representation data by the weight parameter matrix of the pre-trained model to obtain visual key-value data, for example, including a visual key vector and a visual value vector. Multiply the text representation data by the weight parameter matrix to obtain text key-value data, for example, including a text key vector and a text value vector. In addition, in the pre-filling stage, query data corresponding to the visual representation data and the text representation data can also be calculated, for example, a query vector. In addition to storing the key vector and the value vector, the initial key-value data can also store the query vector. The above key vector, value vector, and query vector can be stored in the cache to support the subsequent attention mechanism calculation and decoding process of the model.
[0077] Through the above steps S2241 to S2243, the embodiments of the present application can achieve efficient preprocessing of visual input and text input, uniformly convert input data of multiple modalities into the hidden space of the vision-language model, and facilitate the fusion and processing of cross-modal information. The above solution not only improves the understanding ability of the vision-language model for input data, but also optimizes the calculation process, provides accurate context representation for subsequent attention calculation and decoding, thereby reducing redundant calculations, saving computing resources and accelerating the inference process without sacrificing the model performance. In addition, through the generation of initial key-value data, key information can be quickly located and retrieved in the subsequent decoding stage, avoiding repeated processing of all input data, and further improving the running efficiency and response speed of the vision-language model.
[0078] In an alternative embodiment, the initial key-value data includes: text attention vectors corresponding to multiple text tokens. In step S204, using the initial key-value data to determine the key text tokens among the multiple text tokens includes the following method steps:
[0079] Step S241, perform self-attention calculation on the multiple text tokens using the text attention vectors to obtain a text uni-modal attention matrix, where the text uni-modal attention matrix is used to record the attention weights between the multiple text tokens;
[0080] Step S242, select key text tokens from the multiple text tokens using the text uni-modal attention matrix.
[0081] In addition, in the implementation manner of constructing an elite observation window using the initial key-value data and multiple text tokens, after selecting the key text tokens according to the above steps S241 and S242, an elite observation window can be further constructed based on the key text tokens.
[0082] The above text attention vectors may include at least one of query vectors, key vectors, and value vectors corresponding to multiple text tokens. For example, in an application scenario, when constructing an elite observation window, the text attention vectors corresponding to multiple text tokens used may include: text query vectors (denoted as ) and text key vectors (denoted as ).
[0083] The text uni-modal attention matrix (Text Uni-modal Attention Matrix) is generated through self-attention calculation and records the attention weights between multiple text tokens. "Uni-modal" means that the attention calculation is only limited to the text modality and does not involve cross-modal interactions. The text uni-modal attention matrix helps to determine the correlation and information flow between multiple text tokens and provides a basis for the selection of key text tokens in subsequent steps.
[0084] Furthermore, by analyzing the text unimodal attention matrix, the vision-language model can identify key text tokens among multiple text tokens, which are considered to play a key role in text information transmission and understanding. Specifically, based on the attention weights of the text tokens, key text tokens are selected, and text tokens with higher attention weights are considered more critical text tokens.
[0085] Furthermore, after determining the key text tokens, the vision-language model will construct an elite observation window corresponding to the key text tokens. Through the elite observation window, the vision-language model can more focusedly evaluate the importance of visual tokens, thereby making more biased decisions when compressing the visual key-value cache, which not only reduces the resource consumption of the key-value cache but also ensures the maintenance of the model performance.
[0086] Through the above steps S241 to S243, the embodiments of the present application calculate the self-attention of text tokens by using the text attention vector, and then select key text tokens (and an elite observation window can also be further constructed). The above solution can help achieve accurate evaluation of the importance of visual tokens. Selecting key text tokens not only improves the efficiency of the vision-language model in processing visual information but also ensures that the vision-language model can still maintain high performance even in a highly compressed visual key-value cache environment. Specifically, by selecting key text tokens, the vision-language model avoids relying on all text tokens or consecutive text tokens, reduces redundant calculations in the evaluation process, and at the same time ensures the multi-perspective consistency and stability of visual token evaluation, thereby realizing resource optimization and inference acceleration in the operation of the vision-language model.
[0087] In an alternative embodiment, the text attention vector includes: a text query vector and a text key vector. In step S241, using the text attention vector, self-attention calculation is performed on multiple text tokens to obtain a text unimodal attention matrix, including the following method steps:
[0088] Step S2411, determining a reference text token from multiple text tokens;
[0089] Step S2412, calculating the association tightness between multiple text tokens and the reference text token by using the text query vector and the text key vector to obtain a calculation result;
[0090] Step S2413, generating a text unimodal attention matrix by using the calculation result.
[0091] In the above optional embodiments, the reference text marker can be the last text marker among multiple text markers. By using the text query vector and the text key vector, the tightness of the association between multiple text markers and the reference text marker is calculated. Specifically, through the dot product operation between the text query vector and the text key vector, and then through normalization transformation, the association tightness score (which can be used as the calculation result) is obtained. This association tightness score can reflect the attention degree of each text marker to the reference text marker, providing a basis for generating a more appropriate elite observation window and ensuring that important visual information can be accurately captured.
[0092] It should be noted that in the attention mechanism, the dot product operation between the text query vector and the text key vector essentially measures the similarity between the text query vector and the text key vector, that is, the degree of proximity in the semantic space. The higher the dot product score, the closer the corresponding text marker is associated with the reference text marker in terms of content. After normalization (such as the Softmax function), the above dot product score is converted into a probability distribution, that is, the association tightness score. The association tightness score intuitively reflects the importance and relevance of each text marker relative to the reference text marker. This scoring mechanism ensures that when generating the elite observation window, the key text markers can be selected based on the true association degree between text markers, so that in the subsequent visual marker evaluation, the visual markers corresponding to the key text markers are considered as the focus, achieving the accurate capture and efficient utilization of important visual information.
[0093] In the application scenario, using the text attention vector, self-attention calculation is performed on multiple text markers to obtain the text unimodal attention matrix. The specific calculation method can be as shown in the following formula (1).
[0094] Formula (1)
[0095] In the above formula (1), represents the text unimodal attention matrix, represents the transpose of the text key vector of, represents the hidden state dimension of the vision-language model, represents the real number field, represents the number of text markers, The function is the probability distribution conversion function.
[0096] Through the above formula (1), in the process of calculating the text unimodal attention matrix, through and the dot product operation between them, the correlation score between each text marker and other text markers can be calculated, and then it is scaled by dividing by and then through the application of The function converts these scaled correlation scores into a probability distribution, obtaining a text unimodal attention matrix . The text unimodal attention matrix The weight elements in it will be used in the subsequent weighted summation process to determine the degree of emphasis the visual language model places on different parts when understanding text information.
[0097] Specifically, after selecting the last text token among multiple text tokens as the reference text token, the tightness of association between the multiple text tokens and the reference text token is calculated, and the obtained association tightness score can be the result obtained by using function for to perform normalization calculation.
[0098] It should be noted that the calculated association tightness scores constitute the elements of the text unimodal attention matrix, and the size of this text unimodal attention matrix is , where represents the number of text tokens. Each association tightness score represents the attention weight of the corresponding text token to the last text token (reference text token), and these attention weights illustrate the correlation between each text token and the reference text token at the semantic level.
[0099] Through the above steps S2411 to S2413, the embodiments of the present application can accurately identify the key parts in text information, select key text tokens, and further construct an elite observation window, thereby supporting subsequent effective importance evaluation of visual tokens. The generation of the text unimodal attention matrix not only strengthens the semantic understanding and correlation analysis within the text modality but also provides accurate guidance for cross-modal importance evaluation of visual tokens. The above scheme can significantly improve the efficiency and effect of the key-value data compression strategy, reduce redundant information in the visual key-value cache, and ensure the stability of the model performance. Especially when dealing with complex or long-sequence visual language tasks, the above scheme has more obvious advantages. Therefore, the embodiments of the present application can provide an effective key-value cache management strategy for the efficient inference of visual language models.
[0100] In an alternative embodiment, in step S242, using the text unimodal attention matrix to select key text tokens from multiple text tokens includes the following method steps:
[0101] Step S2421, according to the text unimodal attention matrix and a preset correlation threshold, select key text tokens from multiple text tokens, where the correlation threshold is used to control the sparsity of the selection of key text tokens.
[0102] In the above optional embodiments, during the process of selecting key text markers from multiple text markers, a preset correlation threshold is used to control the sparsity of selecting key text markers according to the text unimodal attention matrix. By applying the above correlation threshold, a text marker is recognized as a key text marker only when the attention weight of the text marker reaches or exceeds the correlation threshold.
[0103] In the application scenario, key text markers are selected according to the following formula (2).
[0104] Formula (2)
[0105] In the above formula (2), represents the index of the key text marker, that is, among the text markers, the th text marker is determined as the key text marker. identifies the integer index from 0 to corresponding to each specific text marker. represents the preset correlation threshold. When takes 0, all text markers will be determined as key text markers. When takes 1, the text marker with the highest attention weight among the
[0106] text markers will be determined as the key text marker. By introducing the preset correlation threshold, the above solution can flexibly select the part of text markers that truly contribute to the importance evaluation of visual markers, rather than simply using all text markers or a fixed number of text markers. This effectively reduces the noise interference in the evaluation process, enhances the stability and accuracy of the evaluation, and provides a more accurate basis for formulating the compression strategy.
[0107] Through the above step S2421, the embodiment of the present application realizes in-depth mining of information within the text modality by selecting key text tokens from multiple text tokens. Combining the text unimodal attention matrix and a preset correlation threshold can specifically highlight the text information segments that play important roles when the vision-language model understands the input data. The above text token screening mechanism based on the attention mechanism ensures that subsequent cross-modal evaluations can focus on more important text regions. Furthermore, when performing key-value cache compression, the vision-language model can more accurately determine which key-value data corresponding to visual tokens are crucial during the decoding process and which key-value data corresponding to visual tokens are redundant information that can be safely eliminated. This helps with the effective management and utilization of the cache space, while maintaining or approaching the performance level of the vision-language model with a full key-value cache, significantly improving the running efficiency and resource savings of large vision-language models when processing long-sequence input and output tasks.
[0108] In an alternative embodiment, the initial key-value data includes: visual key vectors corresponding to multiple visual tokens, text query vectors and text key vectors corresponding to multiple text tokens. In step S206, using the initial key-value data and the distribution positions of the key text tokens among the multiple text tokens, an importance evaluation is performed on the multiple visual tokens to obtain an evaluation result, including the following method steps:
[0109] Step S261, constructing a target query vector and a target key vector using the distribution positions, visual key vectors, text query vectors, and text key vectors;
[0110] Step S262, performing cross-modal attention evaluation calculation based on the target query vector and the target key vector to obtain a target attention matrix, where the target attention matrix is used to evaluate the importance of multiple visual tokens during the inference process of the vision-language model;
[0111] Step S263, performing pooling calculation on the target attention matrix in the text token dimension to obtain an evaluation result, where the evaluation result includes importance scores corresponding to multiple visual tokens respectively.
[0112] The distribution positions of the above key text tokens among the multiple text tokens can be characterized by an elite observation window.
[0113] In the above alternative embodiment, the above visual key vectors are key (Key) vectors corresponding to multiple visual tokens during the pre-padding stage. Combining the text query vectors and text key vectors in the elite observation window with the visual key vectors corresponding to multiple visual tokens to construct a target query vector and a target key vector that can more accurately evaluate the importance of visual tokens.
[0114] Further, during the process of cross-modal attention evaluation calculation, the target query vector and the target key vector are used to evaluate the importance of visual tokens through the attention mechanism, and the target attention matrix is obtained. This target attention matrix reflects the correlation between each visual token and the key text token, thereby quantifying the relative importance of visual tokens in the model inference process.
[0115] The above pooling calculation refers to aggregating the target attention matrix to reduce the data dimension while retaining key information. The above pooling calculation is performed on the dimension of text tokens, so as to comprehensively evaluate the importance of each visual token from the perspective of text. The evaluation results include the importance scores of each visual token after pooling. These importance scores are used to guide the subsequent cache compression strategy to ensure that the model can still maintain good performance while reducing storage requirements.
[0116] Through the above steps S261 to S263, the embodiments of the present application can perform all-round and refined importance evaluation on visual tokens in the vision-language model. The introduction of the elite observation window ensures that the evaluation process focuses on text segments that have a significant impact on model inference, while the cross-modal attention evaluation calculation and pooling operation further quantify the relative importance of visual tokens, providing a strong basis for subsequent cache compression. Thus, when the vision-language model decodes, it can more effectively utilize limited cache resources, specifically retain important visual information, reduce the storage requirements for non-critical visual data, and thus significantly improve the computational efficiency and memory usage efficiency in the decoding stage. The above solution optimizes the processing method of visual tokens. Especially when processing high-resolution image or video data, it can more precisely control the resource consumption of the model, enabling the vision-language model to exhibit better performance and resource management capabilities in large-scale data processing and long-sequence output tasks.
[0117] In an optional embodiment, in step S261, the target query vector and the target key vector are constructed by using the distribution position, the visual key vector, the text query vector, and the text key vector, including the following method steps:
[0118] Step S2611, according to the distribution position, extract the target query vector corresponding to the key text token from the text query vector;
[0119] Step S2612, according to the distribution position, extract the key text key vector corresponding to the key text token from the text key vector;
[0120] Step S2613, use the visual key vector and the key text key vector to construct the target key vector.
[0121] The distribution position of the above-mentioned key text markers among multiple text markers can be characterized by an elite observation window.
[0122] In the above optional embodiment, the above-mentioned elite observation window can be used to determine the index of the key text identifier. . Based on this, a target query vector is extracted from the text query vector , and a key text key vector is extracted from the text key vector . Further, the visual key vector and the key text key vector are combined to construct a target key vector. The key text key vector . Further, the visual key vector and the key text key vector are combined to construct a target key vector.
[0123] In an application scenario, the above-mentioned visual key vector can be a visual key matrix, the text query vector can be a text query matrix, the text key vector can be a text key matrix, the above-mentioned target query vector can be a target query matrix, and the above-mentioned target key vector can be a target key matrix. Based on this, the construction method of the target query matrix can be as shown in the following formula (3).
[0124] Formula (3)
[0125] The construction method of the above-mentioned target key matrix can be as shown in the following formula (4).
[0126] Formula (4)
[0127] In the above formulas (3) and (4), represents the number of key text markers. Vision represents the number of visual identifiers.
[0128] By combining the visual key vector and the key text key vector to construct a target key vector, the key information of the visual modality and the text modality is fused. Through calculation or weighting, the key text key vector screened by the elite observation window is integrated with the visual key vector to generate a more comprehensive and representative target key vector. The construction of the target key vector enables the model to simultaneously consider visual features and key text information in the attention mechanism calculation, thereby more comprehensively understanding the input and making more accurate decisions. Thus, it not only promotes the modality fusion inside the vision-language model but also improves the depth and breadth of the model's understanding of the input data, providing a more accurate basis for subsequent cache compression and decoding processes.
[0129] Further, based on the target query vector and the target key vector, cross-modal attention evaluation calculation can be implemented in the manner shown in the following formula (5) to obtain a target attention matrix .
[0130] Formula (5)
[0131] Further, for the target attention matrix perform pooling calculation on the text token dimension to obtain importance scores corresponding to multiple visual tokens .
[0132] Formula (6)
[0133] Based on the above Formula (6), average pooling can be performed on the target attention matrix along the text dimension to obtain the importance scores corresponding to the above multiple visual tokens, and the above evaluation results can be obtained.
[0134] Through the above steps S2611 to S2613, the embodiments of the present application can significantly improve the efficiency and accuracy of the vision-language model when processing cross-modal information. The use of the elite observation window optimizes the screening and utilization of text information, ensures that the model focuses on the text parts crucial for understanding visual tokens, reduces ineffective calculations, and accelerates the inference process of the model. The construction of the target query vector and the target key vector further refines the vectors participating in the attention mechanism calculation, reduces the consumption of computing resources, and at the same time maintains the stability of the model performance. The above solutions enable the vision-language model to achieve fast, efficient, and accurate cross-modal understanding and generation when processing large-scale visual and language data, especially in the processing of long-sequence outputs and high-resolution visual inputs, showing significant performance improvement and resource optimization effects.
[0135] In an alternative embodiment, in step S208, according to the evaluation results, cache compression processing is performed on the initial key-value data to obtain target key-value data, including the following method steps:
[0136] Step S281, based on the evaluation results, determine the distribution intensity parameter and the distribution skewness parameter corresponding to the importance score;
[0137] Step S282, use the distribution intensity parameter and the distribution skewness parameter to determine the target compression ratio corresponding to the vision-language model;
[0138] Step S283, according to the evaluation results and the target compression ratio, discard some key-value data in the initial key-value data to obtain target key-value data.
[0139] In the above optional embodiments, the evaluation results include importance scores of multiple visual markers. The distribution intensity parameter and the distribution skewness parameter are metrics used to quantify the distribution characteristics of the importance scores. The distribution intensity parameter reflects the degree of concentration of the model's attention to visual information, that is, the higher the sum of the importance scores, the larger the parameter value of the distribution intensity parameter, indicating that the model attaches more importance to the visual modality; while the distribution skewness parameter measures the degree of skewness of the distribution of the importance scores, revealing the difference between important visual markers and ordinary visual markers.
[0140] Based on the distribution intensity parameter and the distribution skewness parameter, the target compression ratio that should be applied at different layers of the vision-language model can be determined. The target compression ratio refers to the amount of key-value data that the model should retain relative to the original data volume. By analyzing the degree of attention to visual markers and the distribution characteristics of scores at a specific layer, the maximum compression rate that the layer can withstand can be intelligently determined without significantly degrading the model performance. Thus, the embodiments of the present application support dynamic, hierarchical compression, which is more flexible than static or uniform compression and more adaptable to the needs of the vision-language model when processing diverse data.
[0141] After determining the target compression ratio, a part of the key-value data in the initial key-value data is discarded according to the evaluation results and the target compression ratio, thereby obtaining the target key-value data. The partial discarding process of the above key-value data can essentially optimize the utilization of storage space, reduce the cache load, and at the same time ensure that the model's access to key visual information is not affected or minimally affected.
[0142] It should be noted that in the embodiments of the present application, during the process of partially discarding key-value data, the specific key-value data to be discarded is the key-value data removed from the key-value data corresponding to visual tokens in the order of non-critical or low importance evaluated. This means that the key-value data corresponding to visual tokens with higher scores under the elite observation window is retained, while the key-value data corresponding to visual tokens with lower scores is discarded, so as to achieve the target compression ratio, optimize the storage space, reduce the cache load, and minimize the impact on the model performance. Through the above steps S281 to S283, the embodiments of the present application can implement an intelligent and dynamic cache compression strategy, especially for the visual part in the vision-language model. By accurately calculating the distribution intensity parameter and distribution skewness parameter of the importance score, the target compression ratio of each layer of the model can be intelligently determined, so as to selectively discard non-critical key-value data while ensuring the complete retention of key visual information. The above solution not only improves the resource utilization efficiency of the model, but also largely avoids the performance degradation caused by excessive compression, providing strong support for the efficient operation and real-time application of the vision-language model. Especially for scenarios dealing with high-resolution images or videos, the above compression strategy can significantly reduce the storage and computing requirements of key-value data, enabling the model to maintain or even improve performance under limited hardware resources.
[0143] In an alternative embodiment, the vision-language model includes multiple model layers. In step S281, based on the evaluation result, to determine the distribution intensity parameter corresponding to the importance score, the following method steps are included:
[0144] Step S2811, based on the evaluation result, determine the in-layer importance scores of multiple visual tokens for the model layer;
[0145] Step S2812, perform a summation calculation on the in-layer importance scores to obtain the distribution intensity parameter, where the distribution intensity parameter is used to characterize the degree of dependence of the inference process of the model layer on visual information.
[0146] The evaluation results include importance scores corresponding to multiple visual tokens. A vision-language model typically includes multiple model layers, each of which is used to process and fuse different aspects of visual information and text information. By combining the importance score of each visual token in the evaluation results with the internal calculation and decision-making process of a specific model layer, the contribution of the visual token in the inference process of that model layer is quantitatively evaluated, thereby obtaining the intra-layer importance score. Thus, the embodiments of the present application not only evaluate the importance of visual tokens, but also delve into the model layer to analyze the degree of dependence of different model layers on visual information, so as to be able to more finely manage the resource requirements of each model layer and improve resource utilization efficiency. Further, by performing a summation operation on the intra-layer importance scores of all visual tokens within the model layer, a distribution intensity parameter is obtained. This distribution intensity parameter is used to reflect the overall degree of dependence of the model layer on visual information during the inference process. The above distribution intensity parameter provides key information for subsequent cache compression strategies to support differential management of the visual key-value caches of different model layers, that is, for model layers that are more dependent on visual information, more cache resources can be allocated, while for model layers that are less dependent on visual information, the cache resources can be appropriately reduced to achieve efficient allocation of resources and avoid unnecessary redundant storage.
[0147] In the application scenario, the distribution intensity parameter can be calculated in the manner of formula (7) as follows .
[0148] Formula (7)
[0149] In the above formula (7), for a certain model layer, the intra-layer importance score can be calculated according to formula (6) for the information within that layer, that is .
[0150] Through the above steps S2811 to S2812, the embodiments of the present application can perform refined resource management on multiple model layers of the vision-language model, especially in the processing of visual information. By determining the intra-layer importance score of visual tokens for each layer, the model can identify which layers have a higher degree of dependence on visual information, so that when compressing the cache, a higher retention rate is given to the visual key-value caches of these layers. The distribution intensity parameter obtained by summation calculation further quantifies the overall demand of each model layer for visual information, providing data support for dynamically adjusting the cache allocation strategy. The above cache management based on intra-layer importance and distribution intensity enables the model to significantly reduce the total cache demand while maintaining high decoding quality and performance, achieving effective optimization of the memory and bandwidth corresponding to the key-value data cache. Especially when processing long-sequence inputs and large-scale visual data, this solution can keep the model running efficiently and avoid the negative impact that resource bottlenecks may bring to the model performance and response speed.
[0151] In an alternative embodiment, the vision - language model includes multiple model layers. In step S281, based on the evaluation results, the distribution skewness parameter corresponding to the importance score is determined, including the following method steps:
[0152] Step S2813, based on the evaluation results, determine multiple groups of importance scores corresponding to the multiple model layers respectively;
[0153] Step S2814, using the multiple groups of importance scores, calculate multiple groups of statistical parameters corresponding to the multiple model layers respectively, where the statistical parameters include at least one of the following: score mean and score standard deviation;
[0154] Step S2815, using the multiple groups of importance scores and the multiple groups of statistical parameters, calculate the distribution skewness parameter, where the distribution skewness parameter is used to characterize the difference in the dependence degree of the inference processes of the multiple model layers on visual information.
[0155] The above - mentioned evaluation results (Evaluation Results) include the importance scores corresponding to multiple visual tokens. The importance scores are grouped according to different model layers (Model Layers) in the vision - language model because the characteristics of each model layer and the way of processing visual information may be different, resulting in different dependence degrees of these model layers on visual tokens. Further, by analyzing the attention mechanism of each model layer, the contribution degree of visual tokens to the model inference process is quantified, and multiple groups of importance scores are determined, providing a basis for subsequent statistical analysis and compression strategy formulation.
[0156] It should be noted that in the process of quantifying the contribution degree of visual tokens to the model inference process to determine multiple groups of importance scores, first, identify the attention patterns of each model layer in the model, especially focusing on those model layers that can effectively connect visual and text information. Then, based on the attention distribution of these model layers, calculate the degree to which each visual token is attended to by different text tokens to form the importance score. The above - mentioned process of forming the importance score takes into account the characteristics of different model layers. For example, some model layers may focus more on local detail processing, while some model layers may pay more attention to global features. Finally, group these importance scores according to the model layer, and each group of scores represents the key degree of visual tokens in the corresponding model layer. Through the above - mentioned method, it is possible to understand how visual information is utilized in different layers, so as to formulate a compression strategy to ensure that while reducing the storage burden of key - value data, the high - performance operation of the model can still be maintained.
[0157] The above-mentioned multiple sets of statistical parameters include: the mean score and the standard deviation of the scores calculated for multiple sets of importance scores (i.e., the visual marker scores for each layer). The above-mentioned multiple sets of statistical parameters reflect the average importance of different model layers to visual information and the volatility of the score distribution. The mean score provides an average perspective on the degree of dependence on visual information, while the standard deviation of the scores reflects the concentration and dispersion of the score distribution. Through the above mean score and standard deviation of the scores, a comprehensive understanding of the visual processing characteristics of the model layer can be achieved.
[0158] Furthermore, by analyzing the mean importance score and the standard deviation of the scores of the model layer, the differences in the processing of visual information by different model layers are quantified to obtain the distribution skewness parameter. The distribution skewness parameter can reveal the imbalance in the dependence of multiple model layers on visual information. That is to say, some model layers may be highly dependent on visual information, while other model layers may be less dependent on visual information. The above distribution skewness parameter helps to understand the architecture characteristics of the model and formulate personalized key-value cache compression strategies for different model layers. It can also guide the visual language model to reasonably allocate the compression budget during the key-value data compression process, ensuring that the key visual information processing in the model layer is not or less negatively affected, while optimizing the overall resource utilization efficiency.
[0159] In the application scenario, the distribution skewness parameter can be calculated in the manner of the following formula (8) .
[0160] Formula (8)
[0161] In the above formula (8), represents the mean score of the importance score, represents the standard deviation of the importance score.
[0162] Through the above steps S2813 to S2815, the embodiments of the present application can achieve a fine analysis of the visual information processing characteristics of each layer of the visual language model. By determining the importance score and calculating the statistical parameters, the differences in the degree of dependence of different model layers on visual information are quantified. The above solution not only deepens the understanding of the internal working principle of the model, but also provides strong data support for dynamically adjusting the visual KV cache compression strategy. The calculation of the distribution skewness parameter enables the model to allocate the compression budget differently according to the characteristics and requirements of each layer for processing visual information, realizing the efficient utilization of resources. The above differential management strategy, especially when processing visual information of high resolution or video sequences, can effectively avoid the loss of key visual features, ensure the high performance of the model, and at the same time significantly reduce the memory occupancy and accelerate the decoding process, providing a basis for the large-scale application of the visual language model.
[0163] In an alternative embodiment, in step S282, using the distribution intensity parameter and the distribution skewness parameter to determine the target compression ratio corresponding to the vision-language model, including the following method steps:
[0164] Step S2821, obtain the preset compression ratio corresponding to the vision-language model;
[0165] Step S2822, use the distribution intensity parameter and the distribution skewness parameter to perform normalization correction on the preset compression ratio to obtain the target compression ratio.
[0166] In the above alternative embodiment, the preset compression ratio refers to the original compression target value set in the vision-language model. The preset compression ratio is used to reflect the initial requirements of users or system administrators for resource optimization of the vision-language model. The preset compression ratio is usually set based on prior knowledge of the performance and resource consumption of the vision-language model under different compression conditions.
[0167] Furthermore, using the distribution intensity parameter and the distribution skewness parameter to perform normalization correction on the preset compression ratio, the preset compression ratio is adjusted to a target compression ratio that better conforms to the current running state of the model. The above normalization correction process can be implemented based on a dynamic allocation strategy, which takes into account the visual information processing characteristics and visual information processing requirements of the vision-language model on multiple model layers, and determines the compression degree of the visual key-value cache of each model layer in a more refined manner. Thus, the target compression ratio can better adapt to the actual running situation of the vision-language model, ensuring that while effectively reducing the resource consumption of key-value cache data compression, minimizing the impact on the model inference performance.
[0168] In the application scenario, perform normalization processing on the distribution intensity parameter to obtain ; perform normalization processing on the distribution skewness parameter to obtain . Calculate the target compression ratio according to the following formula (9)
[0169] Formula (9)
[0170] In the above formula (9), represents the preset compression ratio.
[0171] Through the above steps S2821 to S2822, the embodiments of the present application achieve personalized and dynamic adjustment of the key-value cache compression strategy in the vision-language model. By obtaining the preset compression ratio and using the distribution intensity parameter and the distribution skewness parameter for correction, this solution can intelligently adapt to the resource requirements of the model under different tasks and data inputs. This dynamic adjustment mechanism ensures that the compression of the key-value cache is not only based on static compression targets, but also fully considers the characteristics of visual information processing inside the model, as well as the dynamic changes in the importance distribution of visual tokens, thus achieving a better balance between resource optimization and model performance. In practical applications, the above strategy based on dynamic compression ratio adjustment can significantly improve the running efficiency of the vision-language model. Especially when dealing with large-scale visual data and long-sequence output tasks, it can more precisely control resource consumption and avoid performance degradation caused by excessive compression.
[0172] In an alternative embodiment, in step S283, according to the evaluation result and the target compression ratio, partial key-value data in the initial key-value data is discarded to obtain the target key-value data, including the following method steps:
[0173] Step S2831, based on the evaluation result, sort the importance scores of multiple visual tokens to obtain a sorting result;
[0174] Step S2832, according to the sorting result and the target compression ratio, select the tokens to be compressed from multiple visual tokens;
[0175] Step S2833, discard the visual key-value data corresponding to the tokens to be compressed in the initial key-value data to obtain the target key-value data.
[0176] In the above alternative embodiment, the evaluation result includes the importance scores corresponding to multiple visual tokens. Sorting the importance scores means arranging multiple visual tokens in descending or ascending order to obtain a sorting result. This sorting result directly determines the target object of subsequent compression operations.
[0177] The above target compression ratio specifies the proportion of visual key-value data to be compressed from the initial key-value data. Based on the sorting result and the target compression ratio, the vision-language model can clearly identify which visual tokens belong to the compressible range, that is, the tokens to be compressed. The process of selecting the tokens to be compressed is dynamic, and the target compression ratio can be flexibly adjusted according to the actual situation to balance model performance and resource consumption. By intelligently selecting the tokens to be compressed, it can ensure that even after the compression operation, sufficient visual information is still maintained to support an efficient inference process, while reducing the storage burden brought by redundant visual information.
[0178] It should be noted that in the process of the visual language model identifying which visual markers are within the compressible range based on the sorting results and the target compression ratio, the visual language model first obtains the importance scores of all visual markers through the elite observation window, and then sorts them in descending order according to these importance scores. According to the preset target compression ratio, the model can start from the end of the sorted list and sequentially identify and mark those visual markers with lower scores, which are considered redundant or non-critical as markers to be compressed.
[0179] Furthermore, part of the visual key-value data corresponding to the mark to be compressed is removed from the initial key-value data, thereby obtaining the target key-value data. The above scheme reduces the memory usage during model operation and optimizes the data processing process, so that the model can decode and generate output faster, especially when processing high-resolution images or video content, which can significantly reduce the required cache space while maintaining the performance of the model. The above discarding process is implemented based on cross-modal evaluation results, ensuring that sensitivity to key visual information can be maintained even during the compression process, and the model performance will not be degraded due to unreasonable key-value data compression.
[0180] Through the above steps S2831 to S2833, the embodiment of the present application can realize the intelligent compression of the visual key data of the visual language model. With the help of the sorting of evaluation results and the selection of markers based on the target compression ratio, the model can selectively remove the visual information that contributes less to the reasoning process, while ensuring the complete retention of key visual features. The above scheme can significantly improve the operating efficiency of the model in a resource-constrained environment, reduce storage requirements, but do not sacrifice model performance. Especially when processing large-scale inputs with a large number of visual markers, the above compression strategy can effectively reduce memory pressure and speed up decoding, thereby achieving effective management and use optimization of resources while maintaining the quality of model reasoning. By intelligently compressing visual key data, the technical solution of the embodiment of the present application enables the visual language model to operate more efficiently and economically in various application scenarios (such as visual question answering, image description generation, etc.), demonstrating the advancement and practicality of the visual language model in the field of multimodal information processing.
[0181] According to the aforementioned data processing method, in an exemplary application scenario, the following is provided: Figure 3 An optional key-value data processing process is shown in FIG. Figure 3As shown, after obtaining the text prompt and the input image, feature extraction and transformation are performed on the text prompt to obtain multiple text identifiers. After feature extraction and transformation on the input image, multiple visual identifiers are obtained. Further, through the pre-fill stage, an intermediate representation of the hidden state is generated, and this intermediate representation includes text representation data and visual representation data. Then, based on this intermediate representation, weight calculation is performed to obtain the initial key-value data. Further, based on the above intermediate representation of the hidden state and the initial key-value data, key text identifiers are selected, and these key text identifiers are used to construct an elite observation window. Further, based on the key text identifiers, the importance scores of the visual identifiers are evaluated; based on these importance scores, the distribution intensity parameters are calculated, and the distribution skewness parameters are calculated through probability density distribution analysis. Further, based on the distribution skewness parameters and the distribution intensity parameters, the compression budget is reallocated, and according to the reallocation result, the initial key-value data is compressed to obtain the target key-value data.
[0182] In summary, in the embodiment of the present application, by selecting key text markers from multiple text markers (or an elite observation window corresponding to the key text markers can also be constructed), and further evaluating the importance of visual markers based on the elite observation window, the initial key-value data is cached and compressed accordingly. This compression process is related to the importance of visual markers, which can save the key-value cache overhead while ensuring the efficiency of the subsequent inference of the model based on the target key-value data. That is to say, the present application achieves the purpose of compressing the key-value data of the vision-language model based on the evaluation results of the importance of multiple visual markers, thereby realizing the technical effects of saving the key-value data cache overhead of the vision-language model and increasing the inference efficiency of the model based on the target key-value data, and further solving the technical problems of large key-value data cache overhead and affecting the model inference efficiency in the related technologies.
[0183] In the foregoing operating environment, the present application also provides another data processing method as shown in Figure 4 Figure [Figure number not provided in the original, so it remains as shown]. Figure 4 is a flowchart of another data processing method according to an embodiment of the present application. As shown in Figure 4 Figure [Figure number not provided in the original, so it remains as shown], this data processing method includes:
[0184] Step S401, obtaining multi-modal input data;
[0185] Step S402, extracting the key-value data to be reused corresponding to the multi-modal input data from the target key-value data;
[0186] Step S403, performing inference calculation based on the multi-modal input data and the key-value data to be reused to generate a target answer.
[0187] The above target key-value data is obtained by using any one of the above data processing methods in the historical inference rounds of the vision-language model.
[0188] The above multi-modal input data is the data input into the vision-language model. The multi-modal input data includes the fusion information of visual information (such as images, videos) and text information (such as descriptions, instructions). The multi-modal input data is the basis for the vision-language model to perform tasks.
[0189] The target key-value data is obtained by caching and compressing according to the aforementioned data processing method. The target key-value data contains the key information required by the model during the inference process. The key-value data to be reused extracted from the target key-value data is historical information related to the current multi-modal input data. This extraction mechanism allows the model to avoid repeated calculations, especially in processing continuous inference tasks, which can significantly improve the inference speed and reduce the consumption of computing resources.
[0190] Furthermore, the vision-language model uses the current multi-modal input data and the extracted key-value data to be reused for inference calculation to generate a target answer. Based on this, the vision-language model accelerates the processing of new tasks using the key-value data of historical inference rounds. The generation of the target answer reflects the accuracy of the model's understanding of the input and the effective utilization of historical information.
[0191] Through the above steps S401 to S403, the embodiments of the present application obtain multi-modal input data; extract the key-value data to be reused corresponding to the multi-modal input data from the target key-value data; perform inference calculation based on the multi-modal input data and the key-value data to be reused to generate a target answer; the above target key-value data is obtained by adopting any one of the above data processing methods in the historical inference rounds of the vision-language model. Thus, the embodiments of the present application achieve key-value reuse in the inference process based on the target key-value data, accelerating model inference. And the target key-value data is obtained by the aforementioned data processing method. The aforementioned data processing method constructs an elite observation window corresponding to the key text markers, and further evaluates the importance of visual markers based on the elite observation window, and accordingly caches and compresses the initial key-value data. This compression process is related to the importance of visual markers, which can save the key-value cache overhead while ensuring the efficiency of the model's subsequent inference based on the target key-value data. That is to say, the present application achieves the purpose of accelerating model inference by reusing key values based on the compressed target key-value data, thereby realizing the technical effects of saving the key-value data cache overhead of the vision-language model and increasing the inference efficiency of the model based on the target key-value data, and further solving the technical problems of large key-value data cache overhead and affecting model inference efficiency in the related art.
[0192] It should be noted that the preferred embodiments of the above steps S401 to S403 can be referred to the relevant descriptions above, and will not be elaborated here.
[0193] Under the aforementioned operating environment, the present application also provides asFigure 5 Another data processing method shown as follows. Figure 5 It is a flowchart of another data processing method according to an embodiment of the present application. As Figure 5 shown, the data processing method includes:
[0194] Step S501, obtaining a data processing request through a first application programming interface. Among them, the request data carried in the data processing request includes: multimodal input data;
[0195] Step S502, returning a data processing response through a second application programming interface. Among them, the response data carried in the data processing response includes: a target answer, which is generated by performing inference calculation based on the multimodal input data and the key-value data to be reused. The key-value data to be reused is extracted from the target key-value data.
[0196] The above-mentioned target key-value data is obtained by using the data processing method of any one of the above in the historical inference rounds of the vision-language model.
[0197] The above data processing method of the embodiment of the present application can run on a cloud server to provide a data processing cloud service for a client. The client calls the first application programming interface to send a data processing request. After the cloud server obtains the data processing request through the first application programming interface, it generates a target answer according to the data processing method, and further returns the target answer to the client through the second application programming interface.
[0198] It should be noted that the target key-value data used in the process of generating the above target answer can be obtained according to the data processing method of any one of the foregoing.
[0199] The above first application programming interface and the second application programming interface can be either the same application programming interface or different application programming interfaces. In an optional embodiment, the interface parameters in the above first application programming interface and the second application programming interface may include but are not limited to: interface global identifier, interface signature key, interface timestamp, interface request identifier, system call credential identifier, etc. The above first application programming interface can use a get request (GET) or a post request (POST) as the interface request method to obtain a file processing request. The above second application programming interface can use a lightweight data interchange format (such as JavaScript Object Notation format, abbreviated as JSON format) to feedback a file processing response.
[0200] The embodiments of the present application can implement an efficient and flexible visual language model inference framework. When receiving and processing multi-modal input data, the visual language model can quickly generate a target answer through pre-optimized target key-value data, significantly improving the processing speed and resource utilization efficiency of visual language tasks. Especially in application scenarios that require real-time interaction or large-scale data processing, such as image description generation in an online customer service system, visual question answering on a social media platform, etc., the above solution can provide a faster response time, lower computational cost, and more stable model performance. By reusing the target key-value data obtained in previous inference rounds, when processing a new multi-modal input request, the visual language model can immediately utilize the key information stored previously, avoiding redundant calculations, and enabling the visual language model to still operate efficiently in a high-load and high-real-time environment.
[0201] Through the above steps S501 to S502, the embodiments of the present application obtain multi-modal input data; extract the key-value data to be reused corresponding to the multi-modal input data from the target key-value data; perform inference calculations based on the multi-modal input data and the key-value data to be reused to generate a target answer; the above target key-value data is obtained by using any one of the above data processing methods in the historical inference rounds of the visual language model. Thus, the embodiments of the present application achieve key-value reuse in the inference process based on the target key-value data, accelerating model inference. And the target key-value data is obtained by the foregoing data processing method. The foregoing data processing method constructs an elite observation window corresponding to the key text markers, and further evaluates the importance of the visual markers based on the elite observation window, and accordingly caches and compresses the initial key-value data. This compression process is related to the importance of the visual markers, which can save the key-value cache overhead while ensuring the efficiency of the model's subsequent inference based on the target key-value data. That is to say, the present application achieves the purpose of accelerating model inference by performing key-value reuse based on the compressed target key-value data, thereby realizing the technical effects of saving the key-value data cache overhead of the visual language model and increasing the inference efficiency of the model based on the target key-value data, and further solving the technical problems in the related art that the key-value data cache overhead of the visual language model is large and affects the model inference efficiency.
[0202] It should be noted that the preferred embodiments of the above steps S501 to S502 can be referred to the foregoing related descriptions and will not be elaborated here.
[0203] Under the foregoing operating environment, the present application also provides another data processing method as Figure 6 shown. Figure 6 is a flowchart of another data processing method according to the embodiments of the present application, as Figure 6 shown, and this data processing method includes:
[0204] Step S601, obtain the current input data processing dialogue request. Among them, the request data carried in the data processing dialogue request includes: multimodal input data;
[0205] Step S602, in response to the data processing dialogue request, return a data processing dialogue reply. Among them, the information carried in the data processing dialogue reply includes: a target answer, which is generated by performing inference calculations based on the multimodal input data and the key-value data to be reused. The key-value data to be reused is extracted from the target key-value data;
[0206] Step S603, display the target answer in the graphical user interface.
[0207] The above-mentioned target key-value data is obtained by adopting any one of the above data processing methods in the historical inference rounds of the vision-language model.
[0208] In the embodiment of the present application, the vision-language model receives a data processing dialogue request from the user. This data processing dialogue request is put forward in the form of a data processing dialogue, which means that the model will communicate with the user. The questions or instructions put forward by the user will include multimodal input data. When the user sends a data processing dialogue request through the graphical user interface or other interfaces, the model parses these request data to understand the user's needs and prepare corresponding responses.
[0209] The above data processing dialogue reply not only includes the target answer, that is, the answer or feedback generated by the model according to the multimodal input data provided by the user and the key-value data to be reused saved previously, but may also include other auxiliary information. The key-value data to be reused is extracted from the target key-value data. These key-value data to be reused are the intermediate representations generated by the model in the historical inference rounds and can be used to accelerate the processing of subsequent identical or similar requests. By reusing the key-value data, the model can skip redundant calculation processes and directly call the saved information related to the current request, thereby significantly improving the response speed and calculation efficiency. Especially when dealing with a large amount of visual information and long text sequences, the effect is more obvious.
[0210] Furthermore, present the target answer to the user. As the human-computer interaction interface, the graphical user interface converts the target answer output by the model into an intuitive and easy-to-understand form, which is convenient for the user to understand and use. The above display process ensures that the user can receive the model's answer in a timely manner, enhancing the user experience. Especially in multimodal tasks, the combination of visual and language information makes the answer content more rich and vivid.
[0211] Through the above steps S601 to S603, the embodiment of the present application obtains the current input data processing dialogue request. Among them, the request data carried in the data processing dialogue request includes: multimodal input data; in response to the data processing dialogue request, a data processing dialogue reply is returned. The information carried in the data processing dialogue reply includes: a target answer, which is generated by performing inference calculation based on the multimodal input data and the key-value data to be reused. The key-value data to be reused is extracted from the target key-value data; the target answer is displayed in the graphical user interface; the above target key-value data is obtained by using any one of the above data processing methods in the historical inference rounds of the vision-language model. Thus, the embodiment of the present application realizes key-value reuse in the inference process based on the target key-value data, accelerating model inference. And the target key-value data is obtained by the foregoing data processing method. The foregoing data processing method constructs an elite observation window corresponding to the key text markers, and further evaluates the importance of the visual markers based on the elite observation window, and accordingly caches and compresses the initial key-value data. This compression process is related to the importance of the visual markers, which can save the key-value cache overhead while ensuring the efficiency of the subsequent inference of the model based on the target key-value data. That is to say, the present application achieves the purpose of accelerating model inference by performing key-value reuse based on the compressed target key-value data, thereby realizing the technical effects of saving the key-value data cache overhead of the vision-language model and increasing the inference efficiency of the model based on the target key-value data, and further solving the technical problems of large key-value data cache overhead and affecting model inference efficiency in the related art.
[0212] It should be noted that the preferred implementation manners of the above steps S601 to S603 can be referred to the foregoing related descriptions, and will not be elaborated here.
[0213] In the foregoing operating environment, the present application also provides another data processing method as Figure 7 shown. Figure 7 It is a flowchart of another data processing method according to the embodiment of the present application. As Figure 7 shown, this data processing method includes:
[0214] Step S701, in response to an input instruction acting on the operation interface, display multimodal input data on the operation interface;
[0215] Step S702, in response to a processing instruction acting on the operation interface, display a target answer on the operation interface.
[0216] The above target answer is generated by performing inference calculation based on the multimodal input data and the key-value data to be reused. The key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by using any one of the above data processing methods in the historical inference rounds of the vision-language model.
[0217] The above data processing method provided by the embodiments of the present application can utilize the aforementioned target key-value data to implement a visual data processing solution, which is convenient for human-computer interaction. The user inputs multi-modal input data by triggering an input instruction on the operation interface. Further, the user triggers the system to generate a target answer according to the aforementioned data processing method by triggering a processing instruction on the operation interface. Further, the system displays the target answer in the operation interface for the user.
[0218] It should be noted that the target key-value data utilized in the process of generating the above target answer can be obtained according to any one of the aforementioned data processing methods.
[0219] According to the above method steps, a visual solution for data processing functions is provided. The terminal device provides a graphical user interface, and at least a data processing scenario is displayed in the graphical user interface. The display content of the graphical user interface further includes an input component (such as a text input box, a voice input control, etc.) and a display component (such as a text display window, an image display window, etc.). The user inputs multi-modal input data through the input component. After detecting the user's input behavior, the target answer is obtained by using the target key-value data, and further, the target answer is displayed through the display component in the graphical user interface.
[0220] The embodiments of the present application achieve efficient and intuitive interaction between the vision-language model and the user. The user can freely input multi-modal data including images, texts, etc. through the operation interface. After the model receives the data, it uses the target key-value data obtained by the key-value data caching and compression solution implemented by the aforementioned data processing method for inference calculation, and quickly generates a high-quality target answer. The above solution not only improves the response speed of the model, but also optimizes the resource utilization efficiency, ensuring that even when processing large-scale multi-modal data, the vision-language model can maintain high performance and stability. In addition, the visual display on the operation interface enhances the user's understanding of the input and output, and improves the user experience. That is to say, this solution significantly improves the interactivity and processing efficiency of the vision-language model in practical applications, and provides technical support for the intelligent understanding and generation of multi-modal information.
[0221] Through the above steps S601 to S603, the embodiments of the present application respond to the input instructions acting on the operation interface and display multi-modal input data on the operation interface; respond to the processing instructions acting on the operation interface and display the target answer on the operation interface; the above target answer is generated by performing inference calculations based on the multi-modal input data and the key-value data to be reused, and the key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by using any one of the above data processing methods in the historical inference rounds of the vision-language model. Thus, the embodiments of the present application achieve key-value reuse in the inference process based on the target key-value data, accelerating model inference. And the target key-value data is obtained by the foregoing data processing method, and the foregoing data processing method constructs an elite observation window corresponding to the key text markers, and further evaluates the importance of the visual markers based on the elite observation window, and accordingly caches and compresses the initial key-value data, and this compression process is related to the importance of the visual markers, which can save the key-value cache overhead while ensuring the efficiency of the subsequent inference of the model based on the target key-value data. That is to say, the present application achieves the purpose of accelerating model inference by performing key-value reuse based on the target key-value data after compression processing, thereby realizing the technical effects of saving the key-value data cache overhead of the vision-language model and increasing the inference efficiency of the model based on the target key-value data, and further solving the technical problems in the related art that the key-value data cache overhead of the vision-language model is large and affects the model inference efficiency.
[0222] It should be noted that the preferred embodiments of the above steps S701 to S702 can be referred to the relevant descriptions above, and will not be repeated here.
[0223] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0224] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0225] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc), and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0226] According to an embodiment of the present application, there is also provided an apparatus embodiment for implementing the above data processing method. Figure 8 is a schematic structural diagram of a data processing apparatus according to an embodiment of the present application, as Figure 8 shown, the apparatus includes: an acquisition module 801, configured to acquire a plurality of visual tokens, a plurality of text tokens, and initial key-value data corresponding to a vision-language model; a construction module 802, configured to determine key text tokens among the plurality of text tokens by using the initial key-value data; an evaluation module 803, configured to evaluate the importance of the plurality of visual tokens by using the initial key-value data and the distribution positions of the key text tokens among the plurality of text tokens, and obtain an evaluation result, where the evaluation result is used to characterize the cross-modal attention weight distribution between the plurality of visual tokens and the key text tokens; and a processing module 804, configured to perform cache compression processing on the initial key-value data according to the evaluation result to obtain target key-value data.
[0227] Optionally, the above acquisition module 801 is further configured to: acquire text prompt data and visual prompt data of the vision-language model; perform word segmentation conversion on the text prompt data to obtain a plurality of text tokens; perform segmentation processing on the visual prompt data to obtain a plurality of visual tokens; and perform pre-filling calculation based on the plurality of visual tokens and the plurality of text tokens to obtain initial key-value data.
[0228] Optionally, the above acquisition module 801 is further configured to: perform embedding encoding calculation on the plurality of visual tokens to obtain visual representation data in a hidden state; perform embedding encoding calculation on the plurality of text tokens to obtain text representation data in a hidden state; and perform weight conversion calculation based on the visual representation data and the text representation data to obtain initial key-value data.
[0229] Optionally, the initial key-value data includes: text attention vectors corresponding to multiple text tokens, and the above-mentioned construction module 802 is further configured to: perform self-attention calculation on the multiple text tokens by using the text attention vectors to obtain a text unimodal attention matrix, where the text unimodal attention matrix is used to record the attention weights between the multiple text tokens; select key text tokens from the multiple text tokens by using the text unimodal attention matrix.
[0230] Optionally, the above-mentioned construction module 802 is further configured to: construct an elite observation window based on the key text tokens.
[0231] Optionally, the text attention vector includes: a text query vector and a text key vector, and the above-mentioned construction module 802 is further configured to: determine a reference text token from the multiple text tokens; calculate the association tightness between the multiple text tokens and the reference text token by using the text query vector and the text key vector to obtain a calculation result; generate a text unimodal attention matrix by using the calculation result.
[0232] Optionally, the above-mentioned construction module 802 is further configured to: select key text tokens from the multiple text tokens according to the text unimodal attention matrix and a preset correlation threshold, where the correlation threshold is used to control the sparsity of the selection of key text tokens.
[0233] Optionally, the initial key-value data includes: visual key vectors corresponding to multiple visual tokens, text query vectors and text key vectors corresponding to multiple text tokens, and the above-mentioned evaluation module 803 is further configured to: construct a target query vector and a target key vector by using the distribution position, the visual key vector, the text query vector and the text key vector; perform cross-modal attention evaluation calculation based on the target query vector and the target key vector to obtain a target attention matrix, where the target attention matrix is used to evaluate the importance of the multiple visual tokens in the inference process of the vision-language model; perform pooling calculation on the target attention matrix in the text token dimension to obtain an evaluation result, where the evaluation result includes importance scores corresponding to the multiple visual tokens respectively.
[0234] Optionally, the above-mentioned evaluation module 803 is further configured to: extract a target query vector corresponding to the key text token from the text query vector according to the distribution position; extract a key text key vector corresponding to the key text token from the text key vector according to the distribution position; construct a target key vector by using the visual key vector and the key text key vector.
[0235] Optionally, the above processing module 804 is further configured to: based on the evaluation result, determine a distribution intensity parameter and a distribution skewness parameter corresponding to the importance score; use the distribution intensity parameter and the distribution skewness parameter to determine a target compression ratio corresponding to the vision-language model; and discard some of the key-value data in the initial key-value data according to the evaluation result and the target compression ratio to obtain target key-value data.
[0236] Optionally, the above processing module 804 is further configured to: based on the evaluation result, determine the intra-layer importance scores of multiple visual markers for the model layer; perform a summation calculation on the intra-layer importance scores to obtain a distribution intensity parameter, where the distribution intensity parameter is used to characterize the degree of dependence of the inference process of the model layer on visual information.
[0237] Optionally, the above processing module 804 is further configured to: based on the evaluation result, determine multiple groups of importance scores corresponding to multiple model layers; use the multiple groups of importance scores to calculate multiple groups of statistical parameters corresponding to the multiple model layers, where the statistical parameters include at least one of the following: score mean and score standard deviation; use the multiple groups of importance scores and the multiple groups of statistical parameters to calculate a distribution skewness parameter, where the distribution skewness parameter is used to characterize the difference in the degree of dependence of the inference processes of the multiple model layers on visual information.
[0238] Optionally, the above processing module 804 is further configured to: obtain a preset compression ratio corresponding to the vision-language model; use the distribution intensity parameter and the distribution skewness parameter to perform a normalization correction on the preset compression ratio to obtain a target compression ratio.
[0239] Optionally, the above processing module 804 is further configured to: based on the evaluation result, sort the importance scores of multiple visual markers to obtain a sorting result; select the markers to be compressed from the multiple visual markers according to the sorting result and the target compression ratio; and discard the visual key-value data corresponding to the markers to be compressed in the initial key-value data to obtain target key-value data.
[0240] It should be noted here that the above acquisition module 801, construction module 802, evaluation module 803, and processing module 804 correspond to steps S202 to S208 in the embodiment. The instances and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the content disclosed in the foregoing embodiments.
[0241] It should be noted that the above modules or units may be hardware components or software components stored in a memory and processed by one or more processors, and the above modules may also be part of a device and may run in a computer terminal.
[0242] According to an embodiment of the present application, another device embodiment for implementing the above data processing method is further provided. Figure 9It is a schematic structural diagram of another data processing device according to an embodiment of the present application. As Figure 9 shown, the device includes: an acquisition module 901, configured to acquire multimodal input data; an extraction module 902, configured to extract key-value data to be reused corresponding to the multimodal input data from target key-value data; a calculation module 903, configured to perform inference calculation based on the multimodal input data and the key-value data to be reused, and generate a target answer; wherein, the target key-value data is obtained by adopting the data processing method of any one of the above in the historical inference rounds of the vision-language model.
[0243] It should be noted here that the above acquisition module 901, extraction module 902, and calculation module 903 correspond to steps S401 to S403 in the embodiment. The instances and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the foregoing embodiments.
[0244] According to an embodiment of the present application, there is also provided another device embodiment for implementing the above data processing method. Figure 10 It is a schematic structural diagram of another data processing device according to an embodiment of the present application. As Figure 10 shown, the device includes: a request module 1001, configured to obtain a data processing request through a first application programming interface, wherein the request data carried in the data processing request includes: multimodal input data; a response module 1002, configured to return a data processing response through a second application programming interface, wherein the response data carried in the data processing response includes: a target answer, the target answer is generated by performing inference calculation based on the multimodal input data and the key-value data to be reused, the key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by adopting the data processing method of any one of the above in the historical inference rounds of the vision-language model.
[0245] It should be noted here that the above request module 1001 and response module 1002 correspond to steps S501 to S502 in the embodiment. The instances and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the foregoing embodiments.
[0246] According to an embodiment of the present application, there is also provided another device embodiment for implementing the above data processing method. Figure 11 It is a schematic structural diagram of another data processing device according to an embodiment of the present application. As Figure 11As shown in the figure, the device includes: an acquisition module 1101, configured to acquire a current input data processing dialogue request, where the request data carried in the data processing dialogue request includes: multimodal input data; a return module 1102, configured to return a data processing dialogue reply in response to the data processing dialogue request, where the information carried in the data processing dialogue reply includes: a target answer, which is generated by performing inference calculation based on the multimodal input data and the key-value data to be reused, and the key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by using any one of the above data processing methods in the historical inference rounds of the vision-language model; a display module 1103, configured to display the target answer in a graphical user interface.
[0247] It should be noted here that the above acquisition module 1101, return module 1102, and display module 1103 correspond to steps S601 to S603 in the embodiment. The instances and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the content disclosed in the foregoing embodiments.
[0248] According to an embodiment of the present application, another device embodiment for implementing the above data processing method is further provided. Figure 12 It is a structural schematic diagram of another data processing device according to an embodiment of the present application. As Figure 12 shown in the figure, the device includes: an input module 1201, configured to display multimodal input data on an operation interface in response to an input instruction acting on the operation interface; a processing module 1202, configured to display a target answer on the operation interface in response to a processing instruction acting on the operation interface; where the target answer is generated by performing inference calculation based on the multimodal input data and the key-value data to be reused, and the key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by using any one of the above data processing methods in the historical inference rounds of the vision-language model.
[0249] It should be noted here that the above input module 1201 and processing module 1202 correspond to steps S701 to S702 in the embodiment. The instances and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the content disclosed in the foregoing embodiments.
[0250] It should be noted that the preferred implementation manners of this embodiment can refer to the relevant descriptions in the foregoing embodiments, and will not be repeated here.
[0251] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in the above embodiments, but are not limited to the schemes provided in the above embodiments.
[0252] Embodiments of the present application may provide a data processing system, including: a client for sending multimodal input data; a server connected to the client for extracting reusable key-value data corresponding to the multimodal input data from target key-value data, and performing inference calculations based on the multimodal input data and the reusable key-value data to generate a target answer; the client is further configured to output the target answer; wherein, the target key-value data is obtained by using the data processing method described in any one of the above in the historical inference rounds of a vision-language model.
[0253] Embodiments of the present application may provide an electronic device, including: a memory storing an executable program; a processor for running the program, wherein when the program runs, it executes the data processing method described in any one of the above.
[0254] Figure 13 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 13 shown, the electronic device 130 may include: one or more (only one is shown in the figure) processors 132, a memory 134, a storage controller, and a peripheral interface.
[0255] The above-mentioned electronic device can be understood as an integrated intelligent terminal, including but not limited to servers, desktop computers, personal computers (abbreviated as PC), model all-in-ones, etc., and the above-mentioned model in the above embodiments of the present application can be pre-installed in the electronic device.
[0256] Specifically, the electronic device may pre-install various types of models, including but not limited to models in the fields of natural language processing, visual processing, speech processing, code processing, multimodal task processing, etc., so as to provide a variety of model selections. In different product forms, the electronic device may support one or more model usage methods, including but not limited to model training, model invocation, model fine-tuning, model deployment, model inference and application, etc. In some product forms, the electronic device also supports model management, including but not limited to multi-type model management (supporting the management of various types of models such as discriminative and generative), model version control (supporting the control of different model versions), model evaluation (evaluating the performance and effect of the model based on a model evaluation tool), etc. In other product forms, the electronic device may also create an application based on the model, provide the ability to call an application programming interface (abbreviated as API), and can call the model into the created application through the API interface, and at the same time provide an application management tool to realize the management and monitoring of the application.
[0257] Furthermore, the electronic device may further include data management (supporting the creation and management of model tuning data sets), a training center (providing rich training resources to help users learn and master artificial intelligence (AI) technology), and basic control capabilities (providing enterprise-level basic control capabilities to ensure the security and efficient operation of the system). Through the above functions, a comprehensive and integrated AI development, training, deployment, and application device is provided.
[0258] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the data processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the data processing method in the above embodiments. The memory may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely set relative to the processor, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.
[0259] The processor can call the executable program stored in the memory through the transmission device to execute the data processing method in any one of the above embodiments.
[0260] Those of ordinary skill in the art can understand that the structure shown Figure 13 is only schematic, and the electronic device can also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a mobile Internet device (MID). Figure 13 It does not limit the structure of the above electronic device. For example, the electronic device 130 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 13 , or have a different configuration from that shown Figure 13 .
[0261] Those of ordinary skill in the art can understand that all or part of the steps in the various data processing methods in the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0262] Embodiments of the present application also provide a computer-readable storage medium. Optionally, in this embodiment, the above computer-readable storage medium includes a stored executable program, wherein when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the data processing method of any one of the foregoing items.
[0263] Optionally, in this embodiment, the above storage medium may be located in an electronic device.
[0264] Optionally, in this embodiment, the computer-readable storage medium is set to store an executable program, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the data processing method in any one of the above embodiments.
[0265] Embodiments of the present application also provide a computer program product. Optionally, in this embodiment, the above computer program product may include a computer program, and the above computer program implements the data processing method provided in the above embodiment when executed by a processor.
[0266] Embodiments of the present application also provide a computer program product. Optionally, the above computer program product may include a non-volatile computer-readable storage medium, and the above non-volatile computer-readable storage medium may be used to store a computer program, and the above computer program implements the data processing method provided in the above embodiment when executed by a processor.
[0267] Embodiments of the present application also provide a computer program. Optionally, in this embodiment, the above computer program implements the data processing method provided in the above embodiment when executed by a processor.
[0268] In the above embodiments of the present application, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0269] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the above unit division is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of units or modules can be in an electrical or other form.
[0270] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0271] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0272] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the above-mentioned data processing method in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, ROM, RAM, mobile hard disks, magnetic disks or optical discs that can store program codes.
[0273] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A data processing method, characterized in that: include: Obtain multiple visual tags, multiple text tags, and initial key-value data corresponding to the visual language model; Determine a key text tag among the plurality of text tags using the initial key value data; Using the initial key value data and the distribution position of the key text marker in the multiple text markers, the multiple visual markers are evaluated for importance to obtain an evaluation result, wherein the evaluation result is used to characterize the cross-modal attention weight distribution between the multiple visual markers and the key text marker; According to the evaluation result, cache compression processing is performed on the initial key-value data to obtain target key-value data.
2. The data processing method according to claim 1, characterized in that: Acquiring the plurality of visual marks, the plurality of text marks, and the initial key value data includes: Acquire text prompt data and visual prompt data of the visual language model; Performing word segmentation conversion on the text prompt data to obtain the multiple text tags; Segmenting the visual prompt data to obtain the multiple visual markers; Pre-filling calculation is performed based on the multiple visual tags and the multiple text tags to obtain the initial key value data.
3. The data processing method according to claim 2, characterized in that: Performing pre-fill calculation based on the multiple visual tags and the multiple text tags to obtain the initial key value data includes: Performing embedded coding calculation on the multiple visual markers to obtain visual representation data in a hidden state; Performing embedded coding calculation on the multiple text tags to obtain text representation data in a hidden state; A weight conversion calculation is performed based on the visual representation data and the text representation data to obtain the initial key value data.
4. The data processing method according to claim 1, characterized in that: The initial key value data includes: text attention vectors corresponding to the multiple text tags, and using the initial key value data, determining the key text tag among the multiple text tags includes: Using the text attention vector, performing self-attention calculation on the multiple text tags to obtain a text unimodal attention matrix, wherein the text unimodal attention matrix is used to record the attention weights between the multiple text tags; The key text tag is selected from the multiple text tags using the text unimodal attention matrix.
5. The data processing method according to claim 4, characterized in that: The text attention vector includes: a text query vector and a text key vector. The text attention vector is used to perform self-attention calculation on the multiple text tags to obtain the text unimodal attention matrix, which includes: determining a reference text token from the plurality of text tokens; Using the text query vector and the text key vector, calculating the association closeness between the multiple text tags and the reference text tag to obtain a calculation result; Using the calculation results, the text unimodal attention matrix is generated.
6. The data processing method according to claim 4, characterized in that: Using the text unimodal attention matrix, selecting the key text tag from the multiple text tags includes: The key text marker is selected from the multiple text markers according to the text unimodal attention matrix and a preset relevance threshold, wherein the relevance threshold is used to control the sparseness of the selection of the key text marker.
7. The data processing method according to claim 1, characterized in that: The initial key value data includes: visual key vectors corresponding to the multiple visual tags, text query vectors and text key vectors corresponding to the multiple text tags; using the initial key value data and the distribution positions of the key text tags in the multiple text tags, the multiple visual tags are evaluated for importance, and the evaluation results include: constructing a target query vector and a target key vector using the distribution position, the visual key vector, the text query vector and the text key vector; Performing a cross-modal attention evaluation calculation based on the target query vector and the target key vector to obtain a target attention matrix, wherein the target attention matrix is used to evaluate the importance of the multiple visual markers in the reasoning process of the visual language model; A pooling calculation of the text tag dimension is performed on the target attention matrix to obtain the evaluation result, wherein the evaluation result includes importance scores corresponding to the multiple visual tags respectively.
8. The data processing method according to claim 7, characterized in that: Constructing the target query vector and the target key vector by using the distribution position, the visual key vector, the text query vector and the text key vector comprises: Extracting a target query vector corresponding to the key text tag from the text query vector according to the distribution position; Extracting a key text key vector corresponding to the key text tag from the text key vector according to the distribution position; The target key vector is constructed using the visual key vector and the key text key vector.
9. The data processing method according to claim 1, characterized in that: According to the evaluation result, cache compression processing is performed on the initial key value data to obtain the target key value data including: Based on the evaluation results, determining a distribution intensity parameter and a distribution skewness parameter corresponding to the importance score; Determining a target compression ratio corresponding to the visual language model using the distribution intensity parameter and the distribution skewness parameter; According to the evaluation result and the target compression ratio, part of the key-value data in the initial key-value data is discarded to obtain the target key-value data.
10. The data processing method according to claim 9, characterized in that: The visual language model includes multiple model layers, and based on the evaluation result, determining the distribution intensity parameter corresponding to the importance score includes: Based on the evaluation results, determining intra-layer importance scores of the multiple visual markers for the model layer; The importance scores within the layer are summed up to obtain the distribution intensity parameter, wherein the distribution intensity parameter is used to characterize the degree of dependence of the reasoning process of the model layer on the visual information.
11. The data processing method according to claim 9, characterized in that: The visual language model includes multiple model layers, and based on the evaluation result, determining the distribution skewness parameter corresponding to the importance score includes: Based on the evaluation results, determining a plurality of groups of importance scores corresponding to the plurality of model layers respectively; Using the multiple groups of importance scores, calculating multiple groups of statistical parameters corresponding to the multiple model layers respectively, wherein the statistical parameters include at least one of the following: a score mean and a score standard deviation; The distribution skewness parameter is calculated using the multiple groups of importance scores and the multiple groups of statistical parameters, wherein the distribution skewness parameter is used to characterize the difference in the degree of dependence of the reasoning process of the multiple model layers on the visual information.
12. The data processing method according to claim 9, characterized in that: Determining the target compression ratio corresponding to the visual language model by using the distribution intensity parameter and the distribution skewness parameter includes: Obtaining a preset compression ratio corresponding to the visual language model; The preset compression ratio is normalized and corrected by using the distribution intensity parameter and the distribution skewness parameter to obtain the target compression ratio.
13. The data processing method according to claim 9, characterized in that: According to the evaluation result and the target compression ratio, part of the key-value data in the initial key-value data is discarded to obtain the target key-value data including: Based on the evaluation result, ranking the importance scores of the multiple visual markers to obtain a ranking result; Selecting a mark to be compressed from the multiple visual marks according to the sorting result and the target compression ratio; The visual key value data corresponding to the mark to be compressed in the initial key value data is discarded to obtain the target key value data.
14. A data processing method, characterized in that: include: Obtain multimodal input data; Extracting the key-value data to be reused corresponding to the multimodal input data from the target key-value data; Performing reasoning calculation based on the multimodal input data and the key-value data to be reused to generate a target answer; The target key-value data is obtained by using the data processing method described in any one of claims 1 to 13 in the historical reasoning round of the visual language model.
15. A data processing method, characterized in that: include: Obtaining a data processing request through a first application programming interface, wherein the request data carried in the data processing request includes: multimodal input data; A data processing response is returned through a second application programming interface, wherein the response data carried in the data processing response includes: a target answer, wherein the target answer is generated by performing reasoning calculation based on the multimodal input data and the key-value data to be reused, wherein the key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by using the data processing method described in any one of claims 1 to 13 in the historical reasoning rounds of the visual language model.
16. A data processing method, characterized in that: include: Acquire a currently input data processing dialogue request, wherein the request data carried in the data processing dialogue request includes: multimodal input data; In response to the data processing dialogue request, a data processing dialogue reply is returned, wherein the information carried in the data processing dialogue reply includes: a target answer, the target answer is generated by performing reasoning calculation based on the multimodal input data and the key-value data to be reused, the key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by using the data processing method described in any one of claims 1 to 13 in the historical reasoning round of the visual language model; The target response is presented within a graphical user interface.
17. A data processing method, characterized in that: include: In response to an input instruction acting on the operation interface, displaying multimodal input data on the operation interface; In response to a processing instruction acting on the operation interface, a target answer is displayed on the operation interface; Wherein, the target answer is generated by performing reasoning calculation based on the multimodal input data and the key-value data to be reused, the key-value data to be reused is extracted from the target key-value data, and the target key-value data is obtained by using the data processing method described in any one of claims 1 to 13 in the historical reasoning round of the visual language model.
18. A data processing system, characterized in that: include: Client, used to send multimodal input data; A server connected to the client, configured to extract the to-be-reused key-value data corresponding to the multimodal input data from the target key-value data, and to perform reasoning calculation based on the multimodal input data and the to-be-reused key-value data to generate a target answer; The client is also used to output the target answer; The target key-value data is obtained by using the data processing method described in any one of claims 1 to 13 in the historical reasoning round of the visual language model.
19. An electronic device, characterized in that: include: A memory storing an executable program; A processor, used to run the program, wherein the program executes the data processing method described in any one of claims 1 to 17 when running.
20. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the data processing method according to any one of claims 1 to 17.
21. A computer program product, characterized in that The invention comprises a computer program, which implements the data processing method according to any one of claims 1 to 17 when being executed by a processor.
Citation Information
Patent Citations
Information processing method, electronic equipment and computer readable storage medium
CN119669278A
Self-adaptive prefix key value cache compression method and device and electronic equipment
CN119761430A
Cited By
Code generation method and device based on large language model, equipment and storage medium
CN120406959A
Model task processing acceleration method and device, equipment and medium
CN121053432A
Model task processing acceleration method, apparatus, device, and medium
CN121053432B
Visual language model processing method and device, storage medium and program product
CN121212216A
Problem processing method, device and equipment
CN122240797A