Model processing method, device, computer, storage medium and program product

By introducing business hidden vectors as intermediate states in the large language model, dimensional compression and cross-layer data multiplexing, the problem of excessive KV-Cache memory usage is solved, the model deployment efficiency and performance are improved, and the cost is reduced.

CN119005264BActive Publication Date: 2025-07-08TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411126490.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-15
Publication Date
2025-07-08
Estimated Expiration
2044-08-15

AI Technical Summary

Technical Problem

During the autoregression generation process of existing large language models (LLM), due to the large memory space occupied by key-value cache (KV-Cache), the model deployment efficiency is low and the processing of input data will lead to performance degradation.

Method used

By introducing an intermediate state between the service hidden state and the key-value state, that is, the service hidden vector, and the projection matrix is used for dimensional compression and quantization processing, the video memory usage of key-value cache is reduced, and cross-layer data multiplexing is performed to avoid directly cache the key-value state.

Benefits of technology

The video memory usage is reduced to less than half of the original, which improves the model deployment efficiency and performance, reduces deployment costs, and improves the video memory utilization rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119005264B_ABST
    Figure CN119005264B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a model processing method, apparatus, computer, storage medium, and program product. The method includes: in a first service processing network, adding a service hidden vector obtained by compressing a first service hidden state of service data through a first projection matrix to a key-value cache; performing attention processing on a first key vector and a first value vector obtained by transforming the service hidden vector through a second projection matrix to obtain a second service hidden state; transmitting cross-layer transfer parameters in the first key vector, the first value vector, and the service hidden vector to a second service processing network, and performing attention processing on a second key vector and a second value vector determined by the cross-layer transfer parameters based on the second service hidden state to obtain a third service hidden state; restoring information of the service hidden state obtained by the last service processing network to obtain a service processing result. By using the present application, the video memory occupancy of the model can be reduced, and the model deployment efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a model processing method, apparatus, computer, storage medium, and program product. Background Art

[0002] Current large language models (LLMs) are generally implemented based on the Transformer architecture, especially the Transformer in the decoder mode. During the autoregressive generation process, to avoid duplicate calculations in the model, the model will save some previous intermediate calculation results, namely the key-value cache (KV-Cache). Among them, the size of the video memory space occupied by the KV-Cache is mainly determined by several factors, which can be denoted as Sp = 2LSH. Sp is used to represent the size of the video memory space occupied by the KV-Cache, L represents the number of Transformers included in the current LLM, S represents the length of the vector elements included in the sample, H represents the size of the KV state, 2 represents the key vector (K) and the value vector (V). Moreover, the model scale of the LLM is very large, and with the advent of the long text era, S is also getting larger and larger, which leads to the fact that the size of the video memory space occupied by the KV-Cache (i.e., the proportion of memory occupation) is getting larger and larger, resulting in the limitation of the batch size during model deployment, and further leading to low model deployment efficiency. Therefore, it is extremely important to reduce the size of the video memory space occupied by the KV-Cache. Currently, generally, unimportant partial vector elements in the input data are discarded, or multiple vector elements are fused to reduce the input length of the model. However, the input data of the model is the original attribute of the sample, and the complete input data contains the complete information of the sample. Processing the vector elements in the input data will cause data loss, thereby reducing the model performance. Summary of the Invention

[0003] Embodiments of this application provide a model processing method, apparatus, computer, storage medium, and program product, which can reduce the video memory occupation of the model, improve the model deployment efficiency, reduce the model deployment cost, and improve the model performance.

[0004] On the one hand, embodiments of this application provide a model processing method, which includes:

[0005] In the first service processing network, obtain the first service hidden state corresponding to the service data, compress the dimension of the first service hidden state by using the first projection matrix to obtain a service hidden vector, and add the service hidden vector to the key-value cache;

[0006] Obtain the business hidden vector from the key-value cache, perform key-value conversion on the business hidden vector using the second projection matrix to obtain the first key vector and the first value vector, and perform attention processing on the first key vector and the first value vector to obtain the second business hidden state;

[0007] Obtain the cross-layer transfer parameter from the first key vector, the first value vector, and the business hidden vector, and pass the cross-layer transfer parameter to the second business processing network;

[0008] In the second business processing network, determine the second key vector and the second value vector based on the cross-layer transfer parameter, and perform attention processing on the second key vector and the second value vector based on the second business hidden state to obtain the third business hidden state;

[0009] When the second business processing network is the last business processing network, perform information restoration on the third business hidden state to obtain the business processing result for the business data.

[0010] One aspect of the embodiments of the present application provides a model processing method, and the method includes:

[0011] In the first initial processing network of the initial processing model, obtain the first sample hidden state corresponding to the business sample, perform dimension compression on the first sample hidden state using the first initial projection matrix to obtain the sample hidden vector, and add the sample hidden vector to the sample key-value cache;

[0012] Obtain the sample hidden vector from the sample key-value cache, perform key-value conversion on the sample hidden vector using the second initial projection matrix to obtain the first sample key vector and the first sample value vector, and perform attention processing on the first sample key vector and the first sample value vector to obtain the second sample hidden state;

[0013] Obtain the sample transfer parameter from the first sample key vector, the first sample value vector, and the sample hidden vector, and pass the sample transfer parameter to the second initial processing network in the initial processing model;

[0014] In the second initial processing network, determine the second sample key vector and the second sample value vector based on the sample transfer parameter, and perform attention processing on the second sample key vector and the second sample value vector based on the second sample hidden state to obtain the third sample hidden state;

[0015] When the second initial processing network is the last initial processing network, perform information restoration on the third sample hidden state to obtain the sample processing result for the business sample;

[0016] Adjust the parameters of the initial processing model based on the sample processing results to obtain a business processing model with converged parameters; the business processing model includes a first projection matrix corresponding to the first initial projection matrix and a second projection matrix corresponding to the second initial projection matrix.

[0017] An embodiment of the present application provides a model processing device on the one hand. The device includes:

[0018] A parameter compression module, configured to obtain a first business hidden state corresponding to business data in a first business processing network, compress the dimension of the first business hidden state using the first projection matrix to obtain a business hidden vector, and add the business hidden vector to the key-value cache;

[0019] A key-value processing module, configured to obtain a business hidden vector from the key-value cache, perform key-value conversion on the business hidden vector using the second projection matrix to obtain a first key vector and a first value vector, and perform attention processing on the first key vector and the first value vector to obtain a second business hidden state;

[0020] A cross-layer transfer module, configured to obtain cross-layer transfer parameters from the first key vector, the first value vector, and the business hidden vector, and transfer the cross-layer transfer parameters to a second business processing network;

[0021] The key-value processing module is further configured to, in the second business processing network, determine a second key vector and a second value vector based on the cross-layer transfer parameters, and perform attention processing on the second key vector and the second value vector based on the second business hidden state to obtain a third business hidden state;

[0022] A business prediction module, configured to, when the second business processing network is the last business processing network, restore the information of the third business hidden state to obtain a business processing result for the business data.

[0023] Wherein, when compressing the dimension of the first business hidden state using the first projection matrix to obtain a business hidden vector, the parameter compression module may be used to:

[0024] Compress the dimension of the first business hidden state using the first projection matrix to obtain a first hidden vector;

[0025] Perform quantization compression on the first hidden vector to obtain a business hidden vector;

[0026] When performing key-value conversion on the business hidden vector using the second projection matrix to obtain a first key vector and a first value vector, the key-value processing module may be used to:

[0027] Perform quantization restoration on the business hidden vector to obtain a second hidden vector;

[0028] Determine the product of the first key projection matrix in the second projection matrix and the second hidden vector as the first key vector, and determine the product of the first value projection matrix in the second projection matrix and the second hidden vector as the first value vector.

[0029] Wherein, when quantizing and compressing the first hidden vector to obtain the service hidden vector, the parameter compression module can be used for:

[0030] Group the first hidden vector to obtain N vector groups; N is a positive integer; each vector group includes vector elements;

[0031] Obtain the quantization levels, obtain the vector statistical values respectively corresponding to the N vector groups, and determine the quantization parameters respectively corresponding to the N vector groups based on the quantization levels and the vector statistical values respectively corresponding to the N vector groups;

[0032] Use the quantization parameter corresponding to each vector group to perform compression processing on the vector elements included in the vector group, and obtain the service hidden vectors respectively corresponding to the N vector groups.

[0033] Wherein, when quantizing and restoring the service hidden vector to obtain the second hidden vector, the key-value processing module can be used for:

[0034] Obtain the service hidden vectors and quantization parameters respectively corresponding to the N vector groups, and use the quantization parameter corresponding to each vector group to perform quantization restoration processing on the service hidden vector corresponding to the vector group to obtain the restored hidden vectors respectively corresponding to the N vector groups;

[0035] Merge the restored hidden vectors respectively corresponding to the N vector groups to obtain the second hidden vector.

[0036] Wherein, when performing attention processing on the first key vector and the first value vector to obtain the second service hidden state, the key-value processing module can be used for:

[0037] Determine the product of the first query matrix and the first service hidden state as the first query vector;

[0038] Obtain the vector dimension of the first query vector, and perform feature fusion on the first query vector and the first key vector based on the vector dimension to obtain vector attention;

[0039] Use the vector attention to perform weighted processing on the first value vector to obtain the second service hidden state.

[0040] Wherein, the first service hidden state includes S first service sub-vectors, and S is a positive integer;

[0041] The key-value processing module can be used for:

[0042] If i is 1, obtain the i-th service hidden vector from the key-value cache, perform key-value conversion on the i-th service hidden vector using the second projection matrix to obtain the i-th first key vector and the i-th first value vector, and perform attention processing on the i-th first key vector and the i-th first value vector to obtain the i-th second service sub-vector; i is used to indicate the position of the i-th first service sub-vector among the S first service sub-vectors.

[0043] If i is a positive integer greater than 1 and less than or equal to S, obtain the first query vector of the i-th first service sub-vector, obtain the service hidden vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively from the key-value cache, perform key-value conversion on the first service hidden vector to the i-th service hidden vector respectively using the second projection matrix to obtain the first key vectors and the first value vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively, determine the sub-attentions corresponding to the first first service sub-vector to the i-th first service sub-vector respectively according to the first key vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively and the first query vector corresponding to the i-th first service sub-vector, and use the sub-attentions corresponding to the first first service sub-vector to the i-th first service sub-vector respectively to perform weighted processing on the first value vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively to obtain the i-th second service sub-vector.

[0044] When i is S, combine the first second service sub-vector to the S-th second service sub-vector to form the second service hidden state.

[0045] Among them, when obtaining the cross-layer transfer parameter from the first key vector, the first value vector and the service hidden vector, this cross-layer transfer module can be used for:

[0046] Determine the service hidden vector as the cross-layer transfer parameter;

[0047] When determining the second key vector and the second value vector based on the cross-layer transfer parameter in the second service processing network, this key-value processing module can be used for:

[0048] In the second service processing network, perform key-value conversion on the cross-layer transfer parameter using the third projection matrix in the second service processing network to obtain the second key vector and the second value vector.

[0049] Among them, when obtaining the cross-layer transfer parameter from the first key vector, the first value vector and the service hidden vector, this cross-layer transfer module can be used for:

[0050] Determine the first key vector and the first value vector as the cross-layer transfer parameter;

[0051] In the second service processing network, when determining the second key vector and the second value vector based on the cross-layer transfer parameters, the key-value processing module can be used for:

[0052] In the second service processing network, obtain the second key vector and the second value vector from the cross-layer transfer parameters.

[0053] Among them, when performing attention processing on the first key vector and the first value vector to obtain the second service hidden state, the key-value processing module can be used for:

[0054] Obtain position association information based on the first service hidden state, perform feature fusion on the first key vector, the first value vector, and the position association information to obtain full-scale key-value data;

[0055] Determine the product of the first query matrix and the first service hidden state as the first query vector;

[0056] Perform attention processing on the first query vector using the full-scale key-value data to obtain the second service hidden state.

[0057] Among them, the device further includes:

[0058] A data acquisition module, configured to acquire the number of service processing networks included in the service processing model and acquire the range of parameter reuse layers;

[0059] A layer determination module, configured to obtain a divisor of the number of networks from the range of parameter reuse layers and determine the divisor of the number of networks as the reuse layer;

[0060] A network determination module, configured to determine the (a*d + 1)-th service processing network in the service processing model as the first service processing network, and determine the service processing networks other than the first service processing network included in the service processing model as the second service processing network; d is the reuse layer; a is a natural number.

[0061] Among them, the device further includes:

[0062] A storage transfer module, configured to store the key-value cache corresponding to the first service object in the hardware memory if it is detected that the first service object is in an offline state; the first service object is a service object that provides service data;

[0063] The storage transfer module is further configured to, when it is detected that the object state of the first service object changes from the offline state to the online state, obtain the key-value cache from the hardware memory and execute the process of obtaining the service hidden vector from the key-value cache.

[0064] An embodiment of the present application provides a model processing device on the one hand. The device includes:

[0065] A sample compression module, which is used to obtain a first sample hidden state corresponding to a service sample in a first initial processing network of an initial processing model, perform dimensionality compression on the first sample hidden state by using a first initial projection matrix to obtain a sample hidden vector, and add the sample hidden vector to a sample key-value cache;

[0066] A sample processing module, which is used to obtain a sample hidden vector from the sample key-value cache, perform key-value conversion on the sample hidden vector by using a second initial projection matrix to obtain a first sample key vector and a first sample value vector, and perform attention processing on the first sample key vector and the first sample value vector to obtain a second sample hidden state;

[0067] A parameter processing module, which is used to obtain a sample transfer parameter from the first sample key vector, the first sample value vector and the sample hidden vector, and transfer the sample transfer parameter to a second initial processing network in the initial processing model;

[0068] The sample processing module is further used to determine a second sample key vector and a second sample value vector based on the sample transfer parameter in the second initial processing network, and perform attention processing on the second sample key vector and the second sample value vector based on the second sample hidden state to obtain a third sample hidden state;

[0069] A sample prediction module, which is used to restore information of the third sample hidden state to obtain a sample processing result for the service sample when the second initial processing network is the last initial processing network;

[0070] A model training module, which is used to adjust parameters of the initial processing model based on the sample processing result to obtain a service processing model with converged parameters; the service processing model includes a first projection matrix corresponding to the first initial projection matrix and a second projection matrix corresponding to the second initial projection matrix.

[0071] Wherein, the device further includes:

[0072] A model acquisition module, which is used to acquire a basic model and determine a first initial processing network and a second initial processing network from the basic model;

[0073] A dependency acquisition module, which is used to determine network dependencies and topological dependencies in the basic model based on the model architectures of the first initial processing network, the second initial processing network and the basic model;

[0074] A sparsity processing module, which is used to obtain neurons to be optimized from the basic model based on the network dependencies and topological dependencies, and perform sparsity processing on the neurons to be optimized in the basic model to obtain an initial processing model.

[0075] Wherein, the device further includes:

[0076] An associated acquisition module, configured to acquire a base model and an associated model of the base model, and acquire associated model parameters in the associated model; the task similarity between the associated model task corresponding to the associated model and the base model task corresponding to the base model is greater than or equal to a model association threshold.

[0077] A parameter initialization module, configured to perform parameter initialization on the base model parameters in the base model based on the associated model parameters to obtain an initial processing model.

[0078] One aspect of the embodiments of the present application provides a computer device, including a processor, a memory, and an input / output interface;

[0079] The processor is respectively connected to the memory and the input / output interface. Among them, the input / output interface is used to receive and output data, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device including the processor executes the model processing method in one aspect of the embodiments of the present application.

[0080] One aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the model processing method in one aspect of the embodiments of the present application.

[0081] One aspect of the embodiments of the present application provides a computer program product, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various alternative manners in one aspect of the embodiments of the present application. In other words, when the computer instructions are executed by the processor, the methods provided in various alternative manners in one aspect of the embodiments of the present application are implemented.

[0082] Implementing the embodiments of the present application will have the following beneficial effects:

[0083] In an embodiment of the present application, in the first service processing network, the first service hidden state corresponding to the service data is obtained, the dimensionality of the first service hidden state is compressed using the first projection matrix to obtain a service hidden vector, and the service hidden vector is added to the key-value cache; the service hidden vector is obtained from the key-value cache, and the key-value conversion of the service hidden vector is performed using the second projection matrix to obtain a first key vector and a first value vector, and attention processing is performed on the first key vector and the first value vector to obtain a second service hidden state; from the first key vector, the first value vector, and the service hidden vector, a cross-layer transfer parameter is obtained, and the cross-layer transfer parameter is passed to the second service processing network; in the second service processing network, a second key vector and a second value vector are determined based on the cross-layer transfer parameter, and attention processing is performed on the second key vector and the second value vector based on the second service hidden state to obtain a third service hidden state; when the second service processing network is the last service processing network, information restoration is performed on the third service hidden state to obtain a service processing result for the service data. Through the above process, an intermediate state, that is, a service hidden vector, is introduced between the service hidden state and the key-value state (i.e., the key vector and the value vector). Instead of caching the key-value state, the introduced intermediate state is cached, and this intermediate state is obtained after dimensionality compression of the service hidden state, so that only one copy of data needs to be stored in the cache of the key-value state, and this data has undergone dimensionality compression, thereby reducing the size of the video memory space occupied by the key-value cache to less than half of the original, reducing the video memory occupancy of the model. The reduction in video memory occupancy can increase the batch size, thereby improving the model deployment efficiency, improving the video memory utilization rate, reducing the model deployment cost, and at the same time improving the model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0085] Figure 1 It is a network interaction architecture diagram for model processing provided by an embodiment of the present application;

[0086] Figure 2 It is a schematic diagram of a model processing scenario provided by an embodiment of the present application;

[0087] Figure 3 It is a flowchart of a model processing method provided by an embodiment of the present application;

[0088] Figure 4 It is a schematic diagram of a model compression scenario provided by an embodiment of the present application;

[0089] Figure 5 is another method flowchart for model processing provided by an embodiment of the present application;

[0090] Figure 6 is a schematic diagram of a cross-layer reuse scenario provided by an embodiment of the present application;

[0091] Figure 7 is a schematic diagram of a cross-layer reuse scenario after vector quantization provided by an embodiment of the present application;

[0092] Figure 8 is a schematic diagram of model partitioning provided by an embodiment of the present application;

[0093] Figure 9a is a schematic diagram of a cross-layer reuse scenario provided by an embodiment of the present application Figure 1 ;

[0094] Figure 9b is a schematic diagram of a cross-layer reuse scenario provided by an embodiment of the present application Figure 2 ;

[0095] Figure 9c is a schematic diagram of a cross-layer reuse scenario provided by an embodiment of the present application Figure 3 ;

[0096] Figure 9d is a schematic diagram of a cross-layer reuse scenario provided by an embodiment of the present application Figure 4 ;

[0097] Figure 10 is yet another method flowchart for model processing provided by an embodiment of the present application;

[0098] Figure 11 is still another method flowchart for model processing provided by an embodiment of the present application;

[0099] Figure 12 is a method flowchart of the training process for model processing provided by an embodiment of the present application;

[0100] Figure 13 is a schematic diagram of the operating performance of a model provided by an embodiment of the present application;

[0101] Figure 14 is a schematic diagram of a model processing device provided by an embodiment of the present application;

[0102] Figure 15 is another schematic diagram of a model processing device provided by an embodiment of the present application;

[0103] Figure 16 is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0104] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts shall fall within the protection scope of the present application.

[0105] Among them, if it is necessary to collect data of an object (such as a user, etc.) in the present application, a prompt interface or a pop-up window is displayed before and during the collection. The prompt interface or the pop-up window is used to prompt the user that some data is being collected currently. Only after obtaining the confirmation operation of the user on the prompt interface or the pop-up window, the relevant steps of data acquisition are started, otherwise it ends. Moreover, the obtained user data will be used in reasonable, legal scenarios or uses, etc. Optionally, in some scenarios where user data needs to be used but the user's authorization has not been obtained, authorization can also be requested from the user, and the user data will be used when the authorization is passed. That is to say, the use of user data in the present application complies with the relevant provisions of laws and regulations.

[0106] In the embodiments of the present application, please refer to Figure 1 , Figure 1 which is a network interaction architecture diagram for model processing provided by the embodiments of the present application, as Figure 1As shown, the computer device 101 can obtain service data from the local space, perform prediction on the service data to obtain a service processing result; alternatively, the computer device 101 can obtain service data sent by service devices (such as service device 102a, service device 102b, or service device 102c, etc.), and perform prediction on the service data to obtain a service processing result. Among them, when the computer device 101 performs prediction on the service data to obtain a service processing result, an intermediate state, that is, a service hidden vector, is introduced between the service hidden state and the key-value state. The computer device 101 can compress the dimension of the service hidden state to obtain a service hidden vector, cache the service hidden vector, and project the service hidden vector into a key vector and a value vector when needed, so that the key-value cache only needs to store one piece of data, and this data is obtained based on dimension compression, thus reducing the video memory space occupied by the key-value cache to less than half of the original, greatly reducing the video memory space occupied by the key-value cache and improving the model deployment efficiency. At the same time, the computer device 101 will reuse the service hidden vector across layers, that is, multiple service processing networks can cache the service hidden vector once, that is, cache the intermediate state introduced once, and multiple service processing networks reuse the intermediate state cached this time, so that the video memory space occupied by the key-value cache becomes 1 / d of the original, where d is a positive integer representing the number of service processing networks that reuse the same intermediate state, which also greatly reduces the video memory space occupied by the key-value cache. Among them, the key-value state is used to represent the key vector and the value vector in this application, and the intermediate state is used to represent the data to be cached, that is, the data added to the key-value cache, which is used to indicate the service hidden vector in this application and is the data introduced between the service hidden state and the key-value state.

[0107] Specifically, please refer to Figure 2 , Figure 2 which is a schematic diagram of a model processing scenario provided by an embodiment of this application. As Figure 2As shown, the business processing model may include multiple business processing networks. Among them, the business processing model can be regarded as an LLM model, and a business processing network can be regarded as a Transformer network. The LLM model refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. The large language model can handle various natural language tasks, such as text classification, question answering, dialogue, etc. The Transformer network is a sequence model based on the attention mechanism. Specifically, the business processing networks included in the business processing model can be divided into two categories, denoted as the first business processing network and the second business processing network respectively. The first business processing network is the business processing network that caches business hidden vectors, and the second business processing network is the business processing network that reuses the business hidden vectors in the first business processing network. The computer device can obtain the first business hidden state 202 corresponding to the business data in the first business processing network 201, and the first business hidden state 202 is used to represent the input data of the first business processing network 201; perform dimensional compression on the first business hidden state 202 using the first projection matrix to obtain business hidden vectors, and add the business hidden vectors to the key-value cache 203 to reduce the video memory occupancy of the key-value cache. Further, obtain business hidden vectors from the key-value cache 203, perform key-value conversion on the business hidden vectors using the second projection matrix to obtain the first key vector and the first value vector, and perform attention processing on the first key vector and the first value vector to obtain the second business hidden state, and the second business hidden state is used to represent the output data of the first business processing network. Further, the computer device can obtain cross-layer transfer parameters from data such as the first key vector, the first value vector, and the business hidden vectors 204, and transfer the cross-layer transfer parameters to the second business processing network 205; in the second business processing network 205, determine the second key vector and the second value vector based on the cross-layer transfer parameters, so that the second business processing network 205 does not need to cache intermediate states, realizing the reuse of business hidden vectors by multiple business processing networks, thereby greatly reducing the video memory occupancy of the key-value cache. The computer device can perform attention processing on the second key vector and the second value vector based on the second business hidden state to obtain the third business hidden state. When the second business processing network 205 is the last business processing network, perform information restoration on the third business hidden state to obtain the business processing result for the business data, realizing the prediction processing of the business data.

[0108] It can be understood that the computer device mentioned in the embodiments of the present application includes but is not limited to a terminal device or a server. In other words, the computer device can be a server or a terminal device, or a system composed of a server and a terminal device. Among them, the above-mentioned terminal device can be an electronic device, including but not limited to a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a vehicle-mounted device, an Augmented Reality / Virtual Reality (AR / VR) device, a head-mounted display, a smart TV, a wearable device, a smart speaker, a digital camera, a camera, and other mobile internet devices (MID) with network access capabilities, or terminal devices in scenarios such as trains, ships, and flights. As Figure 1 shown in the figure, the terminal device can be a laptop computer (as shown by service device 102b), a mobile phone (as shown by service device 102c), or a vehicle-mounted device (as shown by service device 102a), etc., Figure 1 only some of the devices are listed. Optionally, the service device 102a refers to a device located in the vehicle 103. The service device 102a can be used to run a business processing model and process business data 1021. Among them, the above-mentioned server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, vehicle-road collaboration, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0109] Optionally, the data involved in the embodiments of the present application can be stored in a computer device, or the data can be stored based on cloud storage technology or a blockchain network, which is not limited here.

[0110] Further, please refer to Figure 3 , Figure 3 which is a flowchart of a model processing method provided by the embodiments of the present application. As Figure 3 shown, the computer device can introduce an intermediate state between the business hidden state and the key-value state based on the Multi-Head Latent Attention (MLA) mechanism of the LLM model to achieve compression of the key-value state and reduce the video memory occupancy of the key-value cache. Specifically, the model processing process includes the following steps:

[0111] Step S301: In the service processing network, obtain the service hidden state corresponding to the service data, perform dimensionality compression on the service hidden state using the first projection matrix to obtain a service hidden vector, and add the service hidden vector to the key-value cache.

[0112] In the embodiments of the present application, reference can be made to Figure 4 , Figure 4 which is a schematic diagram of a model compression scenario provided by the embodiments of the present application. As Figure 4 shown, in the service processing network 401, the computer device can obtain the service hidden state 4011 corresponding to the service data. Among them, the service hidden state corresponding to the service data can be regarded as the input data of the service processing network 401. Further, the computer device can perform dimensionality compression on the service hidden state 4011 using the first projection matrix to obtain a service hidden vector, that is, project the service hidden state 4011 onto a hidden vector with a small hidden layer size (latent), and add the service hidden vector to the key-value cache 402. On the one hand, vector compression is performed, and on the other hand, the cache is changed from caching key vectors and value vectors to caching service hidden vectors, so that the video memory space occupied by the cache corresponding to a single vector element in the key-value cache is reduced from the sizes of the key vector and the value vector to the size of the service hidden vector, greatly reducing the video memory space occupied by the key-value cache. For example, if the sizes of the key vector and the value vector are both 4096 and the size of the latent is 512, then the video memory space occupied by the cache corresponding to a single vector element in the key-value cache is reduced from 2 * 4096 to 512.

[0113] Step S302: Obtain the service hidden vector from the key-value cache, perform key-value conversion on the service hidden vector using the second projection matrix to obtain a key vector and a value vector, and perform attention processing on the key vector and the value vector to obtain the service hidden state corresponding to the service processing network.

[0114] In the embodiments of the present application, as Figure 4As shown in the figure, the computer device can obtain the service hidden vector from the key-value cache 402, perform key-value conversion on the service hidden vector using the second projection matrix to obtain the key vector and the value vector, and perform attention processing on the key vector and the value vector to obtain the service hidden state 4012 corresponding to the service processing network 401. Among them, in a service processing network, the service hidden state corresponding to the service data is the input data of the service processing network, and the service hidden state corresponding to the service processing network is the output data of the service processing network. Further, if the service processing network is not the last service processing network, the service hidden state corresponding to the service processing network can be determined as the service hidden state corresponding to the service data in the next service processing network, and return to execute step S301. If the service processing network is the last service processing network, step S303 can be executed.

[0115] Step S303, when the service processing network is the last service processing network, restore the information of the service hidden state corresponding to the service processing network to obtain the service processing result for the service data.

[0116] In the embodiment of the present application, as Figure 4 shown, when the service processing network 401 is the last service processing network, restore the information of the service hidden state 4012 corresponding to the service processing network 401 to obtain the service processing result for the service data. Among them, the information restoration is determined based on the model processing task of the service processing model. For example, if the model processing task is a question-and-answer task, the information restoration can be to perform decoding processing on the service hidden state 4012 corresponding to the service processing network 401 to obtain the service processing result for the service data, and the service processing result is the answer to the service data, etc.

[0117] In the embodiment of the present application, the computer device can add a first projection matrix and a second projection matrix between the input (i.e., the service hidden state, Hidden States) and the key-value (KV) state of each layer of Transformer to introduce an intermediate state. First, project the input to a hidden vector with a small hidden layer size using the first projection matrix, and then restore it to the KV state to be used. When the model is deployed, the KV state is not retained, only the service hidden vector (latent) is retained, thereby reducing the video memory space occupancy of the KV-Cache. Among them, when the model is actually deployed, it is necessary to obtain the cached latent and then calculate the KV state, which is equivalent to exchanging a small amount of calculation for less video memory occupancy when the model is deployed, compressing 2H in the video memory space occupancy of the KV-Cache to the size of the latent, where H is larger than the size of the latent, to achieve a reduction in the video memory space occupancy of the KV-Cache.

[0118] Further, please refer toFigure 5 , Figure 5 is another flowchart of model processing provided by an embodiment of the present application. As shown in Figure 5 , on the basis of key-value cache compression, data cross-layer reuse is performed. At this time, the service processing network for data caching can be denoted as the first service processing network. The service hidden state corresponding to the service data in the first service processing network is denoted as the first service hidden state, and the service hidden state corresponding to the first service processing network can be denoted as the second service hidden state, etc. Specifically, the model processing process includes the following steps:

[0119] Step S501: In the first service processing network, obtain the first service hidden state corresponding to the service data, perform dimension compression on the first service hidden state using the first projection matrix to obtain a service hidden vector, and add the service hidden vector to the key-value cache.

[0120] In the embodiment of the present application, the computer device can obtain the first service hidden state corresponding to the service data in the first service processing network. Among them, when the first service processing network is the first service processing network in the service processing model, feature extraction can be performed on the service data to obtain the first service hidden state; when the first service processing network is not the first service processing network in the service processing model, the output data of the previous service processing network of the first service processing network can be determined as the first service hidden state.

[0121] Furthermore, the computer device can perform dimension compression on the first service hidden state using the first projection matrix to obtain a service hidden vector. Specifically, the computer device can multiply the first projection matrix (denoted as W_L) by the first service hidden state to obtain a first hidden vector (latent), and directly determine the first hidden vector as the service hidden vector. At this time, the service hidden vector can be denoted as latent. For example, referring to Figure 6 , Figure 6 is a schematic diagram of a cross-layer reuse scenario provided by an embodiment of the present application. As shown in Figure 6 , the computer device can perform dimension compression on the first service hidden state 602 in the first service processing network 601 to obtain a service hidden vector, and add the service hidden vector to the key-value cache 603.

[0122] Alternatively, the first service hidden state can be dimensionally compressed using the first projection matrix to obtain a first hidden vector; the first hidden vector can be quantized and compressed to obtain a service hidden vector. At this time, the service hidden vector can be denoted as Quant.

[0123] Among them, when quantizing and compressing the first hidden vector to obtain the service hidden vector, the computer device can group the first hidden vector to obtain N vector groups; N is a positive integer; each vector group includes vector elements, which can be denoted as [A / group_size, group_size], where A is used to represent the size of the first hidden vector, and group_size is used to represent the number of vector elements included in a vector group, and N = A / group_size. Obtain the quantization level M, which refers to the maximum value corresponding to the quantization technology. The quantization technology can be any technology for quantizing vectors, such as int4 quantization technology, int8 quantization technology or other quantization technologies. For example, when using int4 quantization technology, the quantization level can be 7, and when using int8 quantization technology, the quantization level can be 127, etc.; obtain the vector statistical values corresponding to the N vector groups respectively, and based on the quantization level and the vector statistical values corresponding to the N vector groups respectively, determine the quantization parameters corresponding to the N vector groups respectively, which can be denoted as Scale j = M / |X j |, where Scale j is used to represent the quantization parameter corresponding to the jth vector group, and X j is used to represent the vector statistical value corresponding to the jth vector group. For example, the absolute values of the vector elements included in the jth vector group can be obtained, and the largest absolute value can be determined as the vector statistical value corresponding to the jth vector group. Further, the quantization parameter corresponding to each vector group can be used to perform compression processing on the vector elements included in the vector group to obtain the service hidden vectors corresponding to the N vector groups respectively. For example, the product of the vector elements included in the jth vector group and the quantization parameter corresponding to the jth vector group can be determined as the initial element quantization value of the vector elements included in the jth vector group, and the quantization value with the smallest distance from the initial element quantization value in the quantization value range corresponding to the quantization technology can be determined as the service hidden vector of the vector elements included in the jth vector group. In this way, the video memory space occupancy of the key-value cache can be further reduced. For example, when using int4 quantization technology, the video memory space occupancy of the key-value cache can be further reduced by 4 times. If the initial video memory space occupancy is 5120, at this time, the video memory space occupancy can be compressed to 512 / (5120*2*4) = 1.25%, which shows that the video memory space occupancy of the key-value cache is greatly reduced.

[0124] For example, referring to Figure 7 , Figure 7 is a schematic diagram of a cross-layer reuse scenario after vector quantization provided by an embodiment of the present application. As Figure 7As shown in the figure, the computer device can perform dimensionality compression on the first service hidden state 702 in the first service processing network 701 to obtain a first hidden vector; perform quantization processing on the first hidden vector to obtain a service hidden vector, and add the service hidden vector to the key-value cache 703.

[0125] Step S502: Obtain the service hidden vector from the key-value cache, perform key-value conversion on the service hidden vector using the second projection matrix to obtain a first key vector and a first value vector, and perform attention processing on the first key vector and the first value vector to obtain a second service hidden state.

[0126] In the embodiment of the present application, the computer device can obtain the service hidden vector from the key-value cache, such as Figure 7 in the figure, obtain the service hidden vector from the key-value cache 703. Perform key-value conversion on the service hidden vector using the second projection matrix to obtain a first key vector and a first value vector. Specifically, when the service hidden vector is the first hidden vector, the product of the first key projection matrix (which can be denoted as W_K) in the second projection matrix and the service hidden vector can be determined as the first key vector, and the product of the first value projection matrix (which can be denoted as W_V) in the second projection matrix and the service hidden vector can be determined as the first value vector. Or, when the service hidden vector is obtained after quantization compression of the first hidden vector, the computer device can perform quantization restoration on the service hidden vector to obtain a second hidden vector; the product of the first key projection matrix in the second projection matrix and the second hidden vector can be determined as the first key vector, and the product of the first value projection matrix in the second projection matrix and the second hidden vector can be determined as the first value vector.

[0127] Among them, when performing quantization restoration on the service hidden vector to obtain a second hidden vector, the computer device can obtain the service hidden vectors and quantization parameters corresponding to N vector groups respectively, and use the quantization parameters corresponding to each vector group to perform quantization restoration processing on the service hidden vector corresponding to the vector group to obtain the restored hidden vectors corresponding to the N vector groups respectively, which can be denoted as X″ j = X′ j / Scale j where X′ j is used to represent the service hidden vector of the vector elements included in the j-th vector group, Scale j is used to represent the quantization parameter corresponding to the j-th vector group, and X″ j is used to represent the restored hidden vector corresponding to the j-th vector group. Further, the restored hidden vectors corresponding to the N vector groups are merged to obtain a second hidden vector.

[0128] Further, attention processing can be performed on the first key vector and the first value vector to obtain a second service hidden state. Specifically, the product of the first query matrix and the first service hidden state can be determined as the first query vector; the vector dimension of the first query vector is obtained, and feature fusion is performed on the first query vector and the first key vector based on the vector dimension to obtain vector attention; the first value vector is weighted by the vector attention to obtain a second service hidden state. As Figure 7 shown, attention processing is performed on the first key vector, the first value vector, and the first query vector to obtain a second service hidden state.

[0129] Optionally, when performing attention processing on the first key vector and the first value vector to obtain a second service hidden state, the position association information can be processed separately. Specifically, the position association information can be obtained based on the first service hidden state, and feature fusion is performed on the first key vector, the first value vector, and the position association information to obtain full-scale key-value data. Among them, the product of the position projection matrix (which can be denoted as W_rope) and the position-related data in the first service hidden state (which can be denoted as K_rope) can be determined as the position association information. The product of the first query matrix and the first service hidden state is determined as the first query vector; the first query vector is subjected to attention processing using the full-scale key-value data to obtain a second service hidden state.

[0130] Among them, the first service hidden state includes S first service sub-vectors, where S is a positive integer. If i is 1, obtain the i-th service hidden vector from the key-value cache, perform key-value conversion on the i-th service hidden vector using the second projection matrix to obtain the i-th first key vector and the i-th first value vector, and perform attention processing on the i-th first key vector and the i-th first value vector to obtain the i-th second service sub-vector; i is used to indicate the position of the i-th first service sub-vector among the S first service sub-vectors. If i is a positive integer greater than 1 and less than or equal to S, obtain the first query vector of the i-th first service sub-vector, and from the key-value cache, obtain the service hidden vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively, and perform key-value conversion on the first service hidden vector to the i-th service hidden vector respectively using the second projection matrix to obtain the first key vectors and the first value vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively; according to the first key vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively, and the first query vector corresponding to the i-th first service sub-vector, determine the sub-attentions corresponding to the first first service sub-vector to the i-th first service sub-vector respectively; use the sub-attentions corresponding to the first first service sub-vector to the i-th first service sub-vector respectively to perform weighted processing on the first value vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively to obtain the i-th second service sub-vector. When i is S, form the second service hidden state from the first second service sub-vector to the S-th second service sub-vector.

[0131] Step S503: Obtain the cross-layer transfer parameters from the first key vector, the first value vector, and the service hidden vector, and transfer the cross-layer transfer parameters to the second service processing network.

[0132] In the embodiment of the present application, as Figure 7 shown, the cross-layer transfer parameters can be obtained from the first key vector, the first value vector, and the service hidden vector 704. Among them, in a data reuse method ①, the computer device can determine the service hidden vector as the cross-layer transfer parameter. In a data reuse method ②, the computer device can determine the first key vector and the first value vector as the cross-layer transfer parameter. That is to say, the second service processing network can reuse the service hidden vector of the first service processing network, or directly reuse the key vector and the value vector of the first service processing network. Further, as Figure 6 shown, the computer device can transfer the cross-layer transfer parameters to the second service processing network 604, or as Figure 7 shown, the cross-layer transfer parameters can be transferred to the second service processing network 705.

[0133] Step S504, in the second service processing network, determine the second key vector and the second value vector based on the cross-layer transfer parameters, and perform attention processing on the second key vector and the second value vector based on the second service hidden state to obtain the third service hidden state.

[0134] In the embodiments of the present application, in the second service processing network, determine the second key vector and the second value vector based on the cross-layer transfer parameters. Specifically, in data multiplexing method ①, the computer device can use the third projection matrix in the second service processing network in the second service processing network to perform key-value conversion on the cross-layer transfer parameters to obtain the second key vector and the second value vector. In data multiplexing method ②, the computer device can obtain the second key vector and the second value vector from the cross-layer transfer parameters in the second service processing network. At this time, the second key vector is the first key vector, and the second value vector is the first value vector.

[0135] Further, as Figure 7 shown in, attention processing can be performed on the second key vector and the second value vector based on the second service hidden state to obtain the third service hidden state. The generation process of this third service hidden state can refer to the generation process of the second service hidden state in step S502, which will not be elaborated here.

[0136] Step S505, when the second service processing network is the last service processing network, perform information restoration on the third service hidden state to obtain the service processing result for the service data.

[0137] In the embodiments of the present application, when the second service processing network is the last service processing network, perform information restoration on the third service hidden state to obtain the service processing result for the service data.

[0138] Optionally, KV-Cache occupies video memory space (i.e., graphics processing unit video memory, abbreviated as GPU video memory). When using the LLM model, there may be situations such as switching contexts between different users and interruptions in model usage. KV-Cache will continuously read and write between the video memory space and other storages (such as hardware memories). Reducing the video memory space occupancy of KV-Cache can also improve the read and write efficiency of KV-Cache, reduce model deployment losses, and lower model deployment costs. For example, if it is detected that the first service object is in an offline state, the key-value cache corresponding to the first service object is stored in the hardware memory; the first service object is the service object that provides service data. When it is detected that the object state of the first service object changes from the offline state to the online state, the key-value cache is obtained from the hardware memory, and the process of obtaining the service hidden vector from the key-value cache is executed.

[0139] Among them, the service processing model may include multiple service processing networks, and the multiple service processing networks may be divided into a first service processing network and a second service processing network. For example, refer to Figure 8 , Figure 8 which is a schematic diagram of model division provided by an embodiment of the present application. As Figure 8 shown, the number of service processing networks included in the service processing model may be denoted as L, where L is a positive integer, that is, L service processing networks 802. Among them, d service processing networks share a service hidden vector, and d is a positive integer. At this time, service processing networks 1, d + 1, etc. will add the obtained service hidden vector to the key-value cache 801, and other service processing networks will reuse the cached service hidden vector. As Figure 8 shown in

[0140] shown in Figure 3 or Figure 5 service processing network 1 may add service hidden vector 1 to the key-value cache 801, and service processing networks 2 to d may obtain the cross-layer transfer parameters corresponding to service processing network 1; service processing network d + 1 may add service hidden vector d + 1 to the key-value cache 801, and service processing networks d + 2 to 2d may obtain the cross-layer transfer parameters corresponding to service processing network d + 1, etc. Among them, the division of multiple service processing networks may be performed during the training process of the service processing model. At this time, the first service processing network and the second service processing network can be directly determined based on the training situation. Alternatively, during the application process of the service processing model, multiple service processing networks can be divided, and based on the division result, the service processing model can be fine-tuned to obtain a fine-tuned service processing model. In the fine-tuned service processing model, execute Figure 3 or Figure 5Processes such as the above. Among them, model fine-tuning includes deleting neurons and connections, etc. Among them, when dividing multiple service processing networks, the computer device can obtain the number of networks L of the service processing networks included in the service processing model, obtain the range of parameter reuse layers, and this range of parameter reuse layers is used to represent the number range of service processing networks that can reuse service hidden vectors. By limiting the range of parameter reuse layers, while reducing the video memory space occupancy of the key-value cache, the performance of the model can be prevented from being overly lost. For example, this range of parameter reuse layers can be [2, 4], etc., and is used to balance between the compression rate of the video memory space occupancy and the model performance. From the range of parameter reuse layers, obtain the divisors of the number of networks, and determine the divisors of the number of networks as the reuse layers. Optionally, the reuse layers can also be provided by the management staff. Determine the (a*d + 1)-th service processing network in the service processing model as the first service processing network, and determine the service processing networks other than the first service processing network among the service processing networks included in the service processing model as the second service processing network; d is the reuse layer; a is a natural number, and a*d + 1 is less than or equal to L, that is, determine the 1st service processing network, the (d + 1)-th service processing network,... as the first service processing network.

[0141] In the embodiment of the present application, the computer device may obtain the first service hidden state corresponding to the service data in the first service processing network, compress the dimension of the first service hidden state using the first projection matrix to obtain a service hidden vector, and add the service hidden vector to the key-value cache; obtain the service hidden vector from the key-value cache, perform key-value conversion on the service hidden vector using the second projection matrix to obtain a first key vector and a first value vector, and perform attention processing on the first key vector and the first value vector to obtain a second service hidden state; obtain cross-layer transfer parameters from the first key vector, the first value vector, and the service hidden vector, and transfer the cross-layer transfer parameters to the second service processing network; in the second service processing network, determine a second key vector and a second value vector based on the cross-layer transfer parameters, and perform attention processing on the second key vector and the second value vector based on the second service hidden state to obtain a third service hidden state; when the second service processing network is the last service processing network, restore the information of the third service hidden state to obtain a service processing result for the service data. Through the above process, an intermediate state, that is, a service hidden vector, is introduced between the service hidden state and the key-value state (i.e., the key vector and the value vector). Instead of caching the key-value state, the introduced intermediate state is cached, and this intermediate state is obtained after dimension compression of the service hidden state, so that only one copy of data needs to be stored in the cache of the key-value state, and this data has undergone dimension compression, thereby reducing the size of the video memory space occupied by the key-value cache to less than half of the original, reducing the video memory occupancy of the model. The reduction in video memory occupancy can increase the batch size, thereby improving the model deployment efficiency, improving the video memory utilization rate, reducing the model deployment cost, and improving the model performance at the same time. Among them, the batch size refers to the amount of data used each time when training or inferring the model (i.e., the service processing model), that is, the number of data samples used each time.

[0142] Optionally, when cross-layer reuse is performed, the position association information can be processed separately. There can be various situations for the data of cross-layer reuse. For details, please refer to Figures 9a to 9d . Among them, for one situation ① of cross-layer reuse data, please refer to Figure 9a , Figure 9a is a schematic diagram of a cross-layer reuse scenario provided by the embodiment of the present application Figure 1 , such as Figure 9aAs shown, in the first service processing network 901, the dimensionality of the first service hidden state can be compressed to obtain a service hidden vector, which is determined as a cross-layer transfer parameter. The cross-layer transfer parameter is passed to the second service processing network 902. In the first service processing network 901, the first position association information can be obtained based on the first service hidden state, and feature fusion is performed on the first key vector, the first value vector, and the first position association information to obtain the first full-scale key-value data. Attention processing is performed on the first full-scale key-value data and the first query vector to obtain the second service hidden state. For specific details, please refer to Figure 5 the relevant description in step S502 of

[0143] A cross-layer reuse data case ②, please refer to Figure 9b , Figure 9b is a schematic diagram of a cross-layer reuse scenario provided by an embodiment of the present application Figure 2 , such as Figure 9b As shown, for cross-layer reuse of the key vector and the value vector, in the first service processing network 903, the dimensionality of the first service hidden state can be compressed to obtain a service hidden vector. The first key vector and the first value vector are determined as cross-layer transfer parameters, and the cross-layer transfer parameters are passed to the second service processing network 904. In the first service processing network 903, the first position association information can be obtained based on the first service hidden state, and feature fusion is performed on the first key vector, the first value vector, and the first position association information to obtain the first full-scale key-value data. Attention processing is performed on the first full-scale key-value data and the first query vector to obtain the second service hidden state. For specific details, please refer to Figure 5 the relevant description in step S502 of

[0144] A cross-layer reuse data case ③, please refer to Figure 9c , Figure 9c is a schematic diagram of a cross-layer reuse scenario provided by an embodiment of the present application Figure 3 , such as Figure 9cAs shown in the figure, when cross-layer reuse is performed on the service hidden vector and the first location association information, and the service hidden vector and the first location association information are obtained in the first service processing network 905, the service hidden vector can be determined as the cross-layer transfer parameter, and the cross-layer transfer parameter and the first location association information are transferred to the second service processing network 906. In the first service processing network 905, based on the first key vector, the first value vector, the first location association information, and the first query vector, a second service hidden state is obtained; in the second service processing network 905, based on the second key vector, the second value vector, the second location association information, and the second query vector, a third service hidden state is obtained.

[0145] A cross-layer reuse data situation ④, see Figure 9d , Figure 9d is a schematic diagram of a cross-layer reuse scenario provided by an embodiment of the present application Figure 4 , such as Figure 9d As shown in the figure, when cross-layer reuse is performed on the first key vector, the first value vector, and the first location association information, and the first key vector, the first value vector, and the first location association information are obtained in the first service processing network 907, the first key vector and the first value vector can be determined as the cross-layer transfer parameter, and the cross-layer transfer parameter and the first location association information are transferred to the second service processing network 908. In the first service processing network 907, based on the first key vector, the first value vector, the first location association information, and the first query vector, a second service hidden state is obtained; in the second service processing network 908, based on the second key vector, the second value vector, the second location association information, and the second query vector, a third service hidden state is obtained.

[0146] Further, please see Figure 10 , Figure 10 is another method flow chart of model processing provided by an embodiment of the present application. As shown in Figure 10 the figure, based on the key-value cache compression, vector quantization compression is performed. Specifically, the model processing process includes the following steps:

[0147] Step S1001, in the service processing network, obtain the service hidden state corresponding to the service data, and use the first projection matrix to perform dimensionality compression on the service hidden state to obtain the first hidden vector.

[0148] In the embodiment of the present application, this process can refer to Figure 5 the process of obtaining the first hidden vector in step S501 of

[0149] and will not be elaborated here.

[0150] In the embodiment of the present application, this process can refer toFigure 5 The relevant description in step S501 will not be elaborated here.

[0151] Step S1003: Obtain the service hidden vector from the key-value cache, and perform quantization reduction on the service hidden vector to obtain the second hidden vector.

[0152] In the embodiment of the present application, this process can refer to Figure 5 the relevant description in step S502, which will not be elaborated here.

[0153] Step S1004: Use the second projection matrix to perform key-value conversion on the second hidden vector to obtain a key vector and a value vector, and perform attention processing on the key vector and the value vector to obtain the service hidden state corresponding to the service processing network.

[0154] In the embodiment of the present application, this process can refer to Figure 5 the relevant description of the process of obtaining the second service hidden state in step S502, which will not be elaborated here. Further, if the service processing network is not the last service processing network, the service hidden state corresponding to the service processing network is determined as the service hidden state corresponding to the service data in the next service processing network, and step S1001 is returned for execution for the next service processing network. If the service processing network is the last service processing network, step S1005 is executed.

[0155] Step S1005: When the service processing network is the last service processing network, perform information restoration on the service hidden state corresponding to the service processing network to obtain the service processing result for the service data.

[0156] In the embodiment of the present application, this process can refer to Figure 5 the relevant description in step S505, which will not be elaborated here.

[0157] In the embodiment of the present application, by performing dimensionality compression and quantization compression on the service hidden state, the memory space occupied by the video memory corresponding to a single vector element in the cache is updated from the size of the key vector and the value vector to the size of the quantized and compressed service hidden vector, that is, from 2H to latent size / quantization compression multiple. Since H is greater than the latent size, the memory space occupied by the key-value cache in the video memory is greatly reduced, thereby improving the efficiency and performance of model deployment.

[0158] Further, please refer to Figure 11 , Figure 11 which is also a flowchart of a model processing method provided by the embodiment of the present application. As Figure 11 shown, based on the compression of the key-value cache, vector quantization compression and cross-layer reuse are performed. Specifically, the model processing process includes the following steps:

[0159] Step S1101: In the first service processing network, obtain the first service hidden state corresponding to the service data, and perform dimensionality compression on the first service hidden state using the first projection matrix to obtain the first hidden vector.

[0160] In the embodiments of the present application, reference may be made to Figure 5 the relevant description in step S501 of , and details will not be repeated here.

[0161] Step S1102: Quantize and compress the first hidden vector to obtain the service hidden vector, and add the service hidden vector to the key-value cache.

[0162] In the embodiments of the present application, reference may be made to Figure 5 the relevant description in step S501 of , and details will not be repeated here.

[0163] Step S1103: Obtain the service hidden vector from the key-value cache, and perform quantization reduction on the service hidden vector to obtain the second hidden vector.

[0164] In the embodiments of the present application, reference may be made to Figure 5 the relevant description in step S502 of , and details will not be repeated here.

[0165] Step S1104: Perform key-value conversion on the second hidden vector using the second projection matrix to obtain the first key vector and the first value vector, and perform attention processing on the first key vector and the first value vector to obtain the second service hidden state.

[0166] In the embodiments of the present application, reference may be made to Figure 5 the relevant description in step S502 of , and details will not be repeated here.

[0167] Step S1105: Obtain the cross-layer transfer parameter from the first key vector, the first value vector, and the service hidden vector, and transfer the cross-layer transfer parameter to the second service processing network.

[0168] In the embodiments of the present application, reference may be made to Figure 5 the relevant description in step S503 of , and details will not be repeated here. Optionally, the number of second service processing networks is d - 1, and the first service processing network can transfer the cross-layer transfer parameter to d - 1 second service processing networks.

[0169] Step S1106: In the second service processing network, determine the second key vector and the second value vector based on the cross-layer transfer parameter, and perform attention processing on the second key vector and the second value vector based on the second service hidden state to obtain the third service hidden state.

[0170] In the embodiments of the present application, reference may be made toFigure 5 For the relevant description in step S504, it will not be elaborated here. Optionally, when d - 1 is greater than 1, in the second service processing network, based on the cross - layer transfer parameters, determine the second key vector and the second value vector in the k - th second service processing network. Based on the second service hidden state corresponding to the k - th second service processing network, perform attention processing on the second key vector and the second value vector in the k - th second service processing network to obtain the third service hidden state corresponding to the k - th second service processing network, where k is a positive integer less than or equal to d - 1. When k is less than d - 1, determine the third service hidden state corresponding to the k - th second service processing network as the second service hidden state corresponding to the (k + 1) - th second service processing network, and determine the third service hidden state corresponding to the (k + 1) - th second service processing network. When k is d - 1, execute step S1107.

[0171] Step S1107, the second service processing network is the last service processing network.

[0172] In the embodiments of the present application, if the second service processing network is not the last service processing network, then execute step S1108; if the second service processing network is the last service processing network, then execute step S1109.

[0173] Step S1108, enter the next service processing network.

[0174] Step S1109, restore the information of the third service hidden state to obtain the service processing result for the service data.

[0175] In the embodiments of the present application, reference can be made to Figure 5 the relevant description in step S505, which will not be elaborated here.

[0176] In the embodiments of the present application, it can be denoted as the Cross Layer Latent Attention (CLLA) plus Quantization (Quant) mechanism, which can greatly reduce the video memory space occupancy of the key - value cache. It can compress the key - value cache to 1% of the original, significantly reducing the video memory space occupancy. The reduction of the video memory space occupancy can increase the batchsize, thereby improving the model deployment efficiency. Moreover, the more efficient video memory space utilization can reduce the model deployment cost, improving the model performance while saving costs.

[0177] Further, please refer to Figure 12 , Figure 12 which is a flowchart of the training process method for model processing provided by the embodiments of the present application. As Figure 12 shown, the model processing process includes the following steps:

[0178] In step S1201, in the first initial processing network of the initial processing model, obtain the first sample hidden state corresponding to the service sample, compress the dimension of the first sample hidden state using the first initial projection matrix to obtain a sample hidden vector, and add the sample hidden vector to the sample key-value cache.

[0179] In the embodiments of the present application, this process can refer to Figure 5 in step S501 of

[0180] Optionally, the computer device can directly obtain the initial processing model. Or, it can obtain the basic model, and determine the first initial processing network and the second initial processing network from the basic model; based on the model architectures of the first initial processing network, the second initial processing network, and the basic model, determine the network dependency relationship and the topological dependency relationship in the basic model. The network dependency relationship is used to represent the dependency relationship between adjacent initial processing networks, and the topological dependency relationship is used to represent the dependency relationship of residual structures, etc. in the model; based on the network dependency relationship and the topological dependency relationship, obtain the neurons to be optimized from the basic model, and perform sparsification processing on the neurons to be optimized in the basic model to obtain the initial processing model. Among them, the sparsification processing refers to the process of deleting the neurons to be optimized and the connections associated with the neurons to be optimized.

[0181] Or, it can obtain the basic model and the associated model of the basic model, and obtain the associated model parameters in the associated model; the task similarity between the associated model task corresponding to the associated model and the basic model task corresponding to the basic model is greater than or equal to the model association threshold; perform parameter initialization on the basic model parameters in the basic model based on the associated model parameters to obtain the initial processing model. It can be considered that the associated model and the initial processing model form a teacher-student model framework.

[0182] That is to say, based on the model compression technology, the basic model can be compressed to obtain the initial processing model, and this model compression technology can include but is not limited to methods such as model sparsification or knowledge distillation.

[0183] Optionally, the first initial processing network and the second initial processing network in the initial processing model can be determined. The determination process of the first initial processing network and the second initial processing network can refer to Figure 5 the division process of the first service processing network and the second service processing network in

[0184] Step S1202: Obtain the sample hidden vector from the sample key-value cache, perform key-value conversion on the sample hidden vector using the second initial projection matrix to obtain the first sample key vector and the first sample value vector, and perform attention processing on the first sample key vector and the first sample value vector to obtain the second sample hidden state.

[0185] In the embodiment of the present application, this process can refer to Figure 5 the process of generating and caching the service hidden vector in step S502, which will not be elaborated here.

[0186] Step S1203: Obtain the sample transfer parameters from the first sample key vector, the first sample value vector, and the sample hidden vector, and transfer the sample transfer parameters to the second initial processing network.

[0187] In the embodiment of the present application, this process can refer to Figure 5 the relevant description in step S503, which will not be elaborated here.

[0188] Step S1204: In the second initial processing network, determine the second sample key vector and the second sample value vector based on the sample transfer parameters, and perform attention processing on the second sample key vector and the second sample value vector based on the second sample hidden state to obtain the third sample hidden state.

[0189] In the embodiment of the present application, this process can refer to Figure 5 the process of generating and caching the service hidden vector in step S504, which will not be elaborated here.

[0190] Step S1205: When the second initial processing network is the last initial processing network, restore the information of the third sample hidden state to obtain the sample processing result for the service sample.

[0191] In the embodiment of the present application, this process can refer to Figure 5 the process of generating and caching the service hidden vector in step S505, which will not be elaborated here.

[0192] Step S1206: Adjust the parameters of the initial processing model based on the sample processing result to obtain the service processing model with converged parameters.

[0193] In the embodiment of the present application, the service processing model includes the first projection matrix corresponding to the first initial projection matrix and the second projection matrix corresponding to the second initial projection matrix.

[0194] Among them, the computer device for training the service processing model and the computer device for running the service processing model can be the same device or different devices.

[0195] In the embodiments of the present application, the input data can be first projected onto a hidden vector of a small hidden layer size and restored to the KV state to be used when needed, so that only the projected hidden vector needs to be retained during deployment, thereby reducing the video memory space occupancy of the key-value cache. The original 2H in the video memory space occupancy is compressed to the size of the hidden vector, greatly reducing the video memory space occupancy of the key-value cache. At the same time, cross-layer data reuse is performed on the service processing network, so that multiple service processing networks only need to cache the service hidden vector once, further compressing the video memory space occupancy of the key-value cache. During the autoregressive generation process of the service processing model, the required data can be directly obtained from the key-value cache, that is, less computing resources are used in exchange for more quantization space resources, thereby improving the model deployment efficiency.

[0196] Among them, compared with the prior art, the performance of the model in the present application is shown in Table 1 below:

[0197] Table 1

[0198]

[0199]

[0200]

[0201] As shown in Table 1, the network impact index is used to represent the impact of network congestion on the model performance. BoolQ is used to evaluate the model's ability to handle simple fact-based queries. commonSenseQA is used to represent the correctness of answer prediction in the common sense question and answer text dataset. PIQA is used to represent the model's ability to solve physics-related problems. It can be seen from Table 1 that compared with the existing grouped query attention-based model, each model index of the model implemented in the present application is optimized and the model performance is improved.

[0202] Among them, reference can also be made to Figure 13 , Figure 13 which is a schematic diagram of the model operation performance provided by the embodiments of the present application. As Figure 13 shown, it is used to describe the model accuracy index (Accuracy, ACC) of the existing grouped query attention-based model and the model implemented in the present application, that is, the ACC values of each model under different batch sizes. It can be seen that the model performance in the present application is improved.

[0203] Furthermore, please refer to Figure 14 , Figure 14This is a schematic diagram of a model processing device provided by an embodiment of the present application. The model processing device can be a computer program (including program code, etc.) running on a computer device. For example, the model processing device can be an application software. The device can be used to execute the corresponding steps in the method provided by the embodiment of the present application. As Figure 14 shown, the model processing device 1400 can be used for Figure 5 the computer device in the corresponding embodiment. Specifically, the device can include: a parameter compression module 11, a key-value processing module 12, a cross-layer transfer module 13, and a service prediction module 14.

[0204] The parameter compression module 11 is used to obtain the first service hidden state corresponding to the service data in the first service processing network, compress the dimension of the first service hidden state by using the first projection matrix to obtain a service hidden vector, and add the service hidden vector to the key-value cache;

[0205] The key-value processing module 12 is used to obtain the service hidden vector from the key-value cache, perform key-value conversion on the service hidden vector by using the second projection matrix to obtain a first key vector and a first value vector, and perform attention processing on the first key vector and the first value vector to obtain a second service hidden state;

[0206] The cross-layer transfer module 13 is used to obtain cross-layer transfer parameters from the first key vector, the first value vector, and the service hidden vector, and transfer the cross-layer transfer parameters to the second service processing network;

[0207] The key-value processing module 12 is further used to determine a second key vector and a second value vector based on the cross-layer transfer parameters in the second service processing network, and perform attention processing on the second key vector and the second value vector based on the second service hidden state to obtain a third service hidden state;

[0208] The service prediction module 14 is used to restore the information of the third service hidden state to obtain a service processing result for the service data when the second service processing network is the last service processing network.

[0209] Wherein, when compressing the dimension of the first service hidden state by using the first projection matrix to obtain a service hidden vector, the parameter compression module 11 can be used for:

[0210] Compress the dimension of the first service hidden state by using the first projection matrix to obtain a first hidden vector;

[0211] Perform quantization compression on the first hidden vector to obtain a service hidden vector;

[0212] When performing key-value conversion on the service hidden vector by using the second projection matrix to obtain a first key vector and a first value vector, the key-value processing module 12 can be used for:

[0213] Quantize and restore the service hidden vector to obtain a second hidden vector;

[0214] Determine the first key vector as the product of the first key projection matrix in the second projection matrix and the second hidden vector, and determine the first value vector as the product of the first value projection matrix in the second projection matrix and the second hidden vector.

[0215] Wherein, when quantizing and compressing the first hidden vector to obtain the service hidden vector, the parameter compression module 11 can be used for:

[0216] Group the first hidden vector to obtain N vector groups; N is a positive integer; each vector group includes vector elements;

[0217] Obtain the quantization levels, obtain the vector statistical values respectively corresponding to the N vector groups, and determine the quantization parameters respectively corresponding to the N vector groups based on the quantization levels and the vector statistical values respectively corresponding to the N vector groups;

[0218] Use the quantization parameter corresponding to each vector group to perform compression processing on the vector elements included in the vector group to obtain the service hidden vectors respectively corresponding to the N vector groups.

[0219] Wherein, when quantizing and restoring the service hidden vector to obtain the second hidden vector, the key-value processing module 12 can be used for:

[0220] Obtain the service hidden vectors and quantization parameters respectively corresponding to the N vector groups, and use the quantization parameter corresponding to each vector group to perform quantization restoration processing on the service hidden vector corresponding to the vector group to obtain the restored hidden vectors respectively corresponding to the N vector groups;

[0221] Merge the restored hidden vectors respectively corresponding to the N vector groups to obtain the second hidden vector.

[0222] Wherein, when performing attention processing on the first key vector and the first value vector to obtain the second service hidden state, the key-value processing module 12 can be used for:

[0223] Determine the first query vector as the product of the first query matrix and the first service hidden state;

[0224] Obtain the vector dimension of the first query vector, and perform feature fusion on the first query vector and the first key vector based on the vector dimension to obtain vector attention;

[0225] Use the vector attention to perform weighted processing on the first value vector to obtain the second service hidden state.

[0226] Wherein, the first service hidden state includes S first service sub-vectors, and S is a positive integer;

[0227] The key-value processing module 12 can be used for:

[0228] If i is 1, obtain the i-th service hidden vector from the key-value cache, perform key-value conversion on the i-th service hidden vector using the second projection matrix to obtain the i-th first key vector and the i-th first value vector, and perform attention processing on the i-th first key vector and the i-th first value vector to obtain the i-th second service sub-vector; i is used to indicate the position of the i-th first service sub-vector among the S first service sub-vectors;

[0229] If i is a positive integer greater than 1 and less than or equal to S, obtain the first query vector of the i-th first service sub-vector, obtain the service hidden vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively from the key-value cache, perform key-value conversion on the first service hidden vector to the i-th service hidden vector respectively using the second projection matrix to obtain the first key vectors and the first value vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively, determine the sub-attention corresponding to the first first service sub-vector to the i-th first service sub-vector respectively according to the first key vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively and the first query vector corresponding to the i-th first service sub-vector, and use the sub-attention corresponding to the first first service sub-vector to the i-th first service sub-vector respectively to perform weighted processing on the first value vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively to obtain the i-th second service sub-vector;

[0230] When i is S, form the second service hidden state from the first second service sub-vector to the S-th second service sub-vector.

[0231] Among them, when obtaining the cross-layer transfer parameter from the first key vector, the first value vector and the service hidden vector, the cross-layer transfer module 13 can be used for:

[0232] Determine the service hidden vector as the cross-layer transfer parameter;

[0233] In the second service processing network, when determining the second key vector and the second value vector based on the cross-layer transfer parameter, the key-value processing module 12 can be used for:

[0234] In the second service processing network, perform key-value conversion on the cross-layer transfer parameter using the third projection matrix in the second service processing network to obtain the second key vector and the second value vector.

[0235] Among them, when obtaining the cross-layer transfer parameter from the first key vector, the first value vector and the service hidden vector, the cross-layer transfer module 13 can be used for:

[0236] Determine the first key vector and the first value vector as cross-layer transfer parameters;

[0237] In the second service processing network, when determining the second key vector and the second value vector based on the cross-layer transfer parameters, the key-value processing module 12 can be used for:

[0238] In the second service processing network, obtain the second key vector and the second value vector from the cross-layer transfer parameters.

[0239] Among them, when performing attention processing on the first key vector and the first value vector to obtain the second service hidden state, the key-value processing module 12 can be used for:

[0240] Obtain position correlation information based on the first service hidden state, perform feature fusion on the first key vector, the first value vector, and the position correlation information to obtain full-scale key-value data;

[0241] Determine the product of the first query matrix and the first service hidden state as the first query vector;

[0242] Perform attention processing on the first query vector using the full-scale key-value data to obtain the second service hidden state.

[0243] Among them, the device 1400 further includes:

[0244] The data acquisition module 15 is used to obtain the number of networks in the service processing network included in the service processing model and obtain the range of parameter reuse layers;

[0245] The layer determination module 16 is used to obtain the divisor of the number of networks from the range of parameter reuse layers and determine the divisor of the number of networks as the reuse layer;

[0246] The network determination module 17 is used to determine the (a * d + 1)-th service processing network in the service processing model as the first service processing network, and determine the service processing networks other than the first service processing network in the service processing networks included in the service processing model as the second service processing network; d is the reuse layer; a is a natural number.

[0247] Among them, the device 1400 further includes:

[0248] The storage transfer module 18 is used to store the key-value cache corresponding to the first service object in the hardware memory if it is detected that the first service object is in an offline state; the first service object is a service object that provides service data;

[0249] The storage transfer module 18 is further used to, when it is detected that the object state of the first service object changes from the offline state to the online state, obtain the key-value cache from the hardware memory and execute the process of obtaining the service hidden vector from the key-value cache.

[0250] An embodiment of the present application provides a model processing device. In a first service processing network, the device can obtain a first service hidden state corresponding to service data, compress the dimension of the first service hidden state using a first projection matrix to obtain a service hidden vector, and add the service hidden vector to a key-value cache; obtain the service hidden vector from the key-value cache, perform key-value conversion on the service hidden vector using a second projection matrix to obtain a first key vector and a first value vector, perform attention processing on the first key vector and the first value vector to obtain a second service hidden state; obtain cross-layer transfer parameters from the first key vector, the first value vector, and the service hidden vector, and transfer the cross-layer transfer parameters to a second service processing network; in the second service processing network, determine a second key vector and a second value vector based on the cross-layer transfer parameters, and perform attention processing on the second key vector and the second value vector based on the second service hidden state to obtain a third service hidden state; when the second service processing network is the last service processing network, restore information of the third service hidden state to obtain a service processing result for the service data. Through the above process, an intermediate state, that is, a service hidden vector, is introduced between the service hidden state and the key-value state (i.e., the key vector and the value vector). Instead of caching the key-value state, the introduced intermediate state is cached, and the intermediate state is obtained after dimension compression of the service hidden state, so that only one piece of data needs to be stored in the cache of the key-value state, and the data has undergone dimension compression, thereby reducing the size of the video memory space occupied by the key-value cache to less than half of the original, reducing the video memory occupancy of the model. The reduction in video memory occupancy can increase the batch size, thereby improving the model deployment efficiency, improving the video memory utilization rate, reducing the model deployment cost, and improving the model performance at the same time.

[0251] Further, please refer to Figure 15 , Figure 15 which is a schematic diagram of another model processing device provided by an embodiment of the present application. The model processing device can be a computer program (including program code, etc.) running in a computer device. For example, the model processing device can be an application software; the device can be used to execute the corresponding steps in the method provided by the embodiment of the present application. As Figure 15 shown, the model processing device 1500 can be used for Figure 12 the computer device in the corresponding embodiment. Specifically, the device can include: a sample compression module 21, a sample processing module 22, a parameter processing module 23, a sample prediction module 24, and a model training module 25.

[0252] The sample compression module 21 is used to obtain the first sample hidden state corresponding to the service sample in the first initial processing network of the initial processing model, compress the dimension of the first sample hidden state by using the first initial projection matrix to obtain a sample hidden vector, and add the sample hidden vector to the sample key-value cache;

[0253] The sample processing module 22 is used to obtain the sample hidden vector from the sample key-value cache, perform key-value conversion on the sample hidden vector by using the second initial projection matrix to obtain a first sample key vector and a first sample value vector, and perform attention processing on the first sample key vector and the first sample value vector to obtain a second sample hidden state;

[0254] The parameter processing module 23 is used to obtain sample transfer parameters from the first sample key vector, the first sample value vector and the sample hidden vector, and transfer the sample transfer parameters to the second initial processing network in the initial processing model;

[0255] The sample processing module 22 is further used to determine a second sample key vector and a second sample value vector based on the sample transfer parameters in the second initial processing network, and perform attention processing on the second sample key vector and the second sample value vector based on the second sample hidden state to obtain a third sample hidden state;

[0256] The sample prediction module 24 is used to restore the information of the third sample hidden state when the second initial processing network is the last initial processing network to obtain a sample processing result for the service sample;

[0257] The model training module 25 is used to adjust the parameters of the initial processing model based on the sample processing result to obtain a service processing model with converged parameters; the service processing model includes a first projection matrix corresponding to the first initial projection matrix and a second projection matrix corresponding to the second initial projection matrix.

[0258] Wherein, the device 1500 further includes:

[0259] The model acquisition module 26 is used to acquire a base model and determine the first initial processing network and the second initial processing network from the base model;

[0260] The dependency acquisition module 27 is used to determine the network dependency relationship and the topological dependency relationship in the base model based on the model architectures of the first initial processing network, the second initial processing network and the base model;

[0261] The sparse processing module 28 is used to obtain neurons to be optimized from the base model based on the network dependency relationship and the topological dependency relationship, and perform sparse processing on the neurons to be optimized in the base model to obtain an initial processing model.

[0262] Wherein, the device 1500 further includes:

[0263] An associated acquisition module 29, configured to acquire a base model and an associated model of the base model, and acquire associated model parameters in the associated model; the task similarity between the associated model task corresponding to the associated model and the base model task corresponding to the base model is greater than or equal to a model association threshold.

[0264] A parameter initialization module 30, configured to perform parameter initialization on the base model parameters in the base model based on the associated model parameters to obtain an initial processing model.

[0265] In an embodiment of the present application, the input data can be first projected into a hidden vector with a small hidden layer size, and restored to the KV state to be used when needed, so that only the projected hidden vector needs to be retained during deployment, thereby reducing the video memory space occupancy of the key-value cache. The original 2H in the video memory space occupancy is compressed to the size of the hidden vector, greatly reducing the video memory space occupancy of the key-value cache. At the same time, cross-layer data reuse is performed on the service processing network, so that multiple service processing networks only need to cache the service hidden vector once, further compressing the video memory space occupancy of the key-value cache. During the autoregressive generation process of the service processing model, the required data can be directly obtained from the key-value cache, that is, less computing resources are used in exchange for more quantization space resources, thereby improving the model deployment efficiency.

[0266] See Figure 16 , Figure 16 is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 16 shown, the computer device in the embodiment of the present application may include: one or more processors 1601, a memory 1602, and an input / output interface 1603. The processor 1601, the memory 1602, and the input / output interface 1603 are connected through a bus 1604. The memory 1602 is used to store a computer program, and the computer program includes program instructions. The input / output interface 1603 is used to receive data and output data, such as for data interaction between the computer device and a service device; the processor 1601 is used to execute the program instructions stored in the memory 1602.

[0267] Wherein, when the processor 1601 runs the service processing model, the following operations may be performed:

[0268] In a first service processing network, acquire a first service hidden state corresponding to service data, compress the dimension of the first service hidden state by using a first projection matrix to obtain a service hidden vector, and add the service hidden vector to the key-value cache;

[0269] Obtain the business hidden vector from the key-value cache, perform key-value conversion on the business hidden vector using the second projection matrix to obtain the first key vector and the first value vector, and perform attention processing on the first key vector and the first value vector to obtain the second business hidden state;

[0270] Obtain the cross-layer transfer parameter from the first key vector, the first value vector, and the business hidden vector, and pass the cross-layer transfer parameter to the second business processing network;

[0271] In the second business processing network, determine the second key vector and the second value vector based on the cross-layer transfer parameter, and perform attention processing on the second key vector and the second value vector based on the second business hidden state to obtain the third business hidden state;

[0272] When the second business processing network is the last business processing network, perform information restoration on the third business hidden state to obtain the business processing result for the business data.

[0273] Among them, when training the model, the processor 1601 can perform the following operations:

[0274] In the first initial processing network of the initial processing model, obtain the first sample hidden state corresponding to the business sample, perform dimensionality compression on the first sample hidden state using the first initial projection matrix to obtain the sample hidden vector, and add the sample hidden vector to the sample key-value cache;

[0275] Obtain the sample hidden vector from the sample key-value cache, perform key-value conversion on the sample hidden vector using the second initial projection matrix to obtain the first sample key vector and the first sample value vector, and perform attention processing on the first sample key vector and the first sample value vector to obtain the second sample hidden state;

[0276] Obtain the sample transfer parameter from the first sample key vector, the first sample value vector, and the sample hidden vector, and pass the sample transfer parameter to the second initial processing network in the initial processing model;

[0277] In the second initial processing network, determine the second sample key vector and the second sample value vector based on the sample transfer parameter, and perform attention processing on the second sample key vector and the second sample value vector based on the second sample hidden state to obtain the third sample hidden state;

[0278] When the second initial processing network is the last initial processing network, perform information restoration on the third sample hidden state to obtain the sample processing result for the business sample;

[0279] Adjust the parameters of the initial processing model based on the sample processing results to obtain a business processing model with converged parameters; the business processing model includes a first projection matrix corresponding to a first initial projection matrix and a second projection matrix corresponding to a second initial projection matrix.

[0280] In some possible implementation manners, the processor 1601 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0281] The memory 1602 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1601 and the input / output interface 1603. A part of the memory 1602 may also include a non-volatile random access memory. For example, the memory 1602 may also store information about the device type.

[0282] In a specific implementation, the computer device may execute, through each of its built-in functional modules, the implementation manners provided in each step of the Figure 5 or Figure 12 The implementation manners provided in each step of the Figure 5 or Figure 12 The implementation manners provided in each step of the

[0283] The embodiment of the present application provides a computer device, including: a processor, an input / output interface, and a memory. The processor obtains a computer program in the memory and executes the Figure 5For each step of the method shown in the figure, perform model processing operations. In the embodiment of the present application, in the first service processing network, the first service hidden state corresponding to the service data is obtained, the dimension of the first service hidden state is compressed by using the first projection matrix to obtain a service hidden vector, and the service hidden vector is added to the key-value cache; the service hidden vector is obtained from the key-value cache, the key-value conversion is performed on the service hidden vector by using the second projection matrix to obtain a first key vector and a first value vector, and the attention processing is performed on the first key vector and the first value vector to obtain a second service hidden state; the cross-layer transfer parameter is obtained from the first key vector, the first value vector and the service hidden vector, and the cross-layer transfer parameter is passed to the second service processing network; in the second service processing network, the second key vector and the second value vector are determined based on the cross-layer transfer parameter, and the attention processing is performed on the second key vector and the second value vector based on the second service hidden state to obtain a third service hidden state; when the second service processing network is the last service processing network, the information is restored for the third service hidden state to obtain a service processing result for the service data. Through the above process, an intermediate state, that is, a service hidden vector, is introduced between the service hidden state and the key-value state (that is, the key vector and the value vector). It is not necessary to cache the key-value state, but to cache the introduced intermediate state, and this intermediate state is obtained after dimension compression of the service hidden state, so that only one piece of data needs to be stored in the cache of the key-value state, and this data has been dimension-compressed, thereby reducing the size of the video memory space occupied by the key-value cache to less than half of the original, reducing the video memory occupancy of the model. The reduction of the video memory occupancy can increase the batch size, thereby improving the model deployment efficiency, improving the video memory utilization rate, reducing the model deployment cost, and at the same time improving the model performance.

[0284] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded and executed by the processor Figure 5 or Figure 12 each step of the model processing method provided in the figure, specifically, refer to the Figure 5 or Figure 12 implementation manners provided in each step in the figure, which will not be elaborated here. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. For the technical details not disclosed in the embodiment of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiment of the present application. As an example, the computer program can be deployed to be executed on a computer device, or on multiple computer devices located at one place, or on multiple computer devices distributed at multiple places and interconnected through a communication network.

[0285] The computer-readable storage medium may be the model processing device provided in any of the foregoing embodiments or the internal storage unit of the computer device, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.

[0286] An embodiment of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 5 or Figure 12 The method provided in various alternative manners in realizes the introduction of an intermediate state, that is, a service hidden vector, between the service hidden state and the key-value state (i.e., the key vector and the value vector), without caching the key-value state, but caching the introduced intermediate state, and the intermediate state is obtained after dimension compression of the service hidden state, so that only one copy of data needs to be stored in the cache of the key-value state, and the data has been dimension-compressed, so that the size of the video memory space occupied by the key-value cache is reduced to less than half of the original, reducing the video memory occupancy of the model. The reduction of the video memory occupancy can increase the batch size, thereby improving the model deployment efficiency, improving the video memory utilization rate, reducing the model deployment cost, and at the same time improving the model performance.

[0287] The terms "first", "second", etc. in the description, claims and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other step units inherent to these processes, methods, devices, products or equipment.

[0288] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0289] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in this description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0290] The methods and related devices provided in the embodiments of the present application are described with reference to the method flowcharts and / or structural schematic diagrams provided in the embodiments of the present application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processing machine, or other programmable model processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable model processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can guide the computer or other programmable model processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto the computer or other programmable model processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.

[0291] The steps in the method of the embodiments of the present application can be adjusted in order, combined, and deleted according to actual needs.

[0292] The modules in the device of the embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0293] The foregoing disclosure is only for the preferred embodiments of the present application, and of course cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A model processing method, characterized in that, The method includes: In a first service processing network, obtain a first service hidden state corresponding to service data, perform dimensionality compression on the first service hidden state using a first projection matrix to obtain a first hidden vector, perform quantization compression on the first hidden vector to obtain a service hidden vector, and add the service hidden vector to a key-value cache; wherein, the service data is text data; Obtain the service hidden vector from the key-value cache, perform quantization restoration on the service hidden vector to obtain a second hidden vector, determine a first key vector as the product of a first key projection matrix in a second projection matrix and the second hidden vector, determine a first value vector as the product of a first value projection matrix in the second projection matrix and the second hidden vector, and perform attention processing on the first key vector and the first value vector to obtain a second service hidden state; Obtain cross-layer transfer parameters from the first key vector, the first value vector, and the service hidden vector, and transfer the cross-layer transfer parameters to a second service processing network; In the second service processing network, determine a second key vector and a second value vector based on the cross-layer transfer parameters, and perform attention processing on the second key vector and the second value vector based on the second service hidden state to obtain a third service hidden state; When the second service processing network is the last service processing network, perform information restoration on the third service hidden state to obtain a service processing result for the service data.

2. The method according to claim 1, characterized in that, The performing quantization compression on the first hidden vector to obtain a service hidden vector includes: Group the first hidden vector to obtain N vector groups; N is a positive integer; each vector group includes vector elements; Obtain quantization levels, obtain vector statistical values respectively corresponding to the N vector groups, and determine quantization parameters respectively corresponding to the N vector groups based on the quantization levels and the vector statistical values respectively corresponding to the N vector groups; Use the quantization parameter corresponding to each vector group to perform compression processing on the vector elements included in the vector group to obtain service hidden vectors respectively corresponding to the N vector groups.

3. The method according to claim 2, wherein The performing quantization restoration on the service hidden vector to obtain a second hidden vector includes: Obtain the service hidden vectors and quantization parameters respectively corresponding to the N vector groups, and use the quantization parameter corresponding to each vector group to perform quantization restoration processing on the service hidden vector corresponding to the vector group to obtain restored hidden vectors respectively corresponding to the N vector groups; Merge the restored hidden vectors respectively corresponding to the N vector groups to obtain a second hidden vector.

4. The method according to claim 1, wherein The performing attention processing on the first key vector and the first value vector to obtain a second service hidden state includes: Determine a first query vector as the product of a first query matrix and the first service hidden state; Obtain the vector dimension of the first query vector, and perform feature fusion on the first query vector and the first key vector based on the vector dimension to obtain vector attention; Use the vector attention to perform weighted processing on the first value vector to obtain a second service hidden state.

5. The method according to claim 1, wherein The first service hidden state includes S first service sub-vectors, where S is a positive integer; Obtain the service hidden vector from the key-value cache, perform quantization restoration on the service hidden vector to obtain a second hidden vector, determine the product of the first key projection matrix in the second projection matrix and the second hidden vector as the first key vector, and determine the product of the first value projection matrix in the second projection matrix and the second hidden vector as the first value vector. Perform attention processing on the first key vector and the first value vector to obtain a second service hidden state, including: If i is 1, obtain the i-th service hidden vector from the key-value cache, perform quantization restoration on the i-th service hidden vector to obtain the i-th second hidden vector, determine the product of the first key projection matrix in the second projection matrix and the i-th second hidden vector as the i-th first key vector, and determine the product of the first value projection matrix in the second projection matrix and the i-th second hidden vector as the i-th first value vector. Perform attention processing on the i-th first key vector and the i-th first value vector to obtain the i-th second service sub-vector; i is used to indicate the position of the i-th first service sub-vector in the S first service sub-vectors; If i is a positive integer greater than 1 and less than or equal to S, obtain the first query vector of the i-th first service sub-vector, and from the key-value cache, obtain the service hidden vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively. Perform quantization restoration on each service hidden vector to obtain the second hidden vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively. Determine the products of the first key projection matrix in the second projection matrix and each second hidden vector as the first key vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively, and determine the products of the first value projection matrix in the second projection matrix and each second hidden vector as the first value vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively. Determine the sub-attention corresponding to the first first service sub-vector to the i-th first service sub-vector respectively according to the first key vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively and the first query vector corresponding to the i-th first service sub-vector. Use the sub-attention corresponding to the first first service sub-vector to the i-th first service sub-vector respectively to perform weighted processing on the first value vectors corresponding to the first first service sub-vector to the i-th first service sub-vector respectively to obtain the i-th second service sub-vector; When i is S, combine the first second service sub-vector to the S-th second service sub-vector to form the second service hidden state.

6. The method according to claim 1, characterized in that, The obtaining of the cross-layer transfer parameter from the first key vector, the first value vector and the service hidden vector includes: Determine the service hidden vector as the cross-layer transfer parameter; In the second service processing network, determining a second key vector and a second value vector based on the cross-layer transfer parameter includes: In the second service processing network, using a third projection matrix in the second service processing network to perform key-value conversion on the cross-layer transfer parameter to obtain a second key vector and a second value vector.

7. The method according to claim 1, characterized in that Obtaining a cross-layer transfer parameter from the first key vector, the first value vector, and the service hidden vector includes: Determining the first key vector and the first value vector as the cross-layer transfer parameter; In the second service processing network, determining a second key vector and a second value vector based on the cross-layer transfer parameter includes: In the second service processing network, obtaining a second key vector and a second value vector from the cross-layer transfer parameter.

8. The method according to claim 1, characterized in that Performing attention processing on the first key vector and the first value vector to obtain a second service hidden state includes: Obtaining position association information based on the first service hidden state, performing feature fusion on the first key vector, the first value vector, and the position association information to obtain full-scale key-value data; Determining the product of a first query matrix and the first service hidden state as a first query vector; Performing attention processing on the first query vector using the full-scale key-value data to obtain a second service hidden state.

9. The method according to claim 1, characterized in that, The method further includes: Obtaining the number of networks in the service processing network included in the service processing model, and obtaining a parameter reuse layer range; Obtaining a divisor of the number of networks from the parameter reuse layer range, and determining the divisor of the number of networks as the reuse layer; Determining the (a*d + 1)-th service processing network in the service processing model as the first service processing network, and determining service processing networks other than the first service processing network in the service processing network included in the service processing model as the second service processing network; d is the reuse layer; a is a natural number.

10. The method according to claim 1, characterized in that The method further includes: If it is detected that the first service object is in an offline state, storing the key-value cache corresponding to the first service object in a hardware memory; the first service object is a service object providing the service data; When it is detected that the object state of the first service object changes from the offline state to the online state, obtaining the key-value cache from the hardware memory, and performing the process of obtaining the service hidden vector from the key-value cache.

11. A model processing method, characterized in that, The method includes: In a first initial processing network of an initial processing model, obtaining a first sample hidden state corresponding to a service sample, performing dimension compression on the first sample hidden state using a first initial projection matrix to obtain a first hidden vector, performing quantization compression on the first sample hidden vector to obtain a sample hidden vector, and adding the sample hidden vector to a sample key-value cache; wherein the service sample is a text sample. Obtain the sample hidden vector from the sample key-value cache, perform quantization restoration on the sample hidden vector to obtain a second hidden vector, and determine the product of the first key projection matrix in the second initial projection matrix and the second hidden vector as the first sample key vector. Determine the product of the first value projection matrix in the second initial projection matrix and the second hidden vector as the first sample value vector, and perform attention processing on the first sample key vector and the first sample value vector to obtain a second sample hidden state; Obtain a sample transfer parameter from the first sample key vector, the first sample value vector, and the sample hidden vector, and transfer the sample transfer parameter to a second initial processing network in the initial processing model; In the second initial processing network, determine a second sample key vector and a second sample value vector based on the sample transfer parameter, and perform attention processing on the second sample key vector and the second sample value vector based on the second sample hidden state to obtain a third sample hidden state; When the second initial processing network is the last initial processing network, perform information restoration on the third sample hidden state to obtain a sample processing result for the service sample; Based on the sample processing result, adjust the parameters of the initial processing model to obtain a service processing model with converged parameters; the service processing model includes a first projection matrix corresponding to the first initial projection matrix and a second projection matrix corresponding to the second initial projection matrix.

12. The method according to claim 11, wherein The method further includes: Obtain a base model, and determine the first initial processing network and the second initial processing network from the base model; Based on the model architectures of the first initial processing network, the second initial processing network, and the base model, determine the network dependency relationship and the topological dependency relationship in the base model; Based on the network dependency relationship and the topological dependency relationship, obtain neurons to be optimized from the base model, and perform sparsification processing on the neurons to be optimized in the base model to obtain the initial processing model.

13. The method according to claim 11, wherein The method further includes: Obtain a base model and an associated model of the base model, and obtain associated model parameters in the associated model; the task similarity between the associated model task corresponding to the associated model and the base model task corresponding to the base model is greater than or equal to a model association threshold; Based on the associated model parameters, perform parameter initialization on the base model parameters in the base model to obtain the initial processing model.

14. A model processing device, characterized in that, The device includes: A parameter compression module, configured to obtain a first service hidden state corresponding to service data in a first service processing network, perform dimension compression on the first service hidden state using a first projection matrix to obtain a first hidden vector, perform quantization compression on the first hidden vector to obtain a service hidden vector, and add the service hidden vector to a key-value cache; wherein the service data is text data; The key-value processing module is used to obtain the service hidden vector from the key-value cache, perform quantization restoration on the service hidden vector to obtain a second hidden vector, determine the product of the first key projection matrix in the second projection matrix and the second hidden vector as the first key vector, determine the product of the first value projection matrix in the second projection matrix and the second hidden vector as the first value vector, and perform attention processing on the first key vector and the first value vector to obtain a second service hidden state; The cross-layer transfer module is used to obtain cross-layer transfer parameters from the first key vector, the first value vector, and the service hidden vector, and transfer the cross-layer transfer parameters to the second service processing network; The key-value processing module is further used to determine a second key vector and a second value vector based on the cross-layer transfer parameters in the second service processing network, and perform attention processing on the second key vector and the second value vector based on the second service hidden state to obtain a third service hidden state; The service prediction module is used to perform information restoration on the third service hidden state to obtain a service processing result for the service data when the second service processing network is the last service processing network.

15. A model processing device, characterized in that, The device includes: The sample compression module is used to obtain a first sample hidden state corresponding to a service sample in the first initial processing network of the initial processing model, perform dimensionality compression on the first sample hidden state using the first initial projection matrix to obtain a first hidden vector, perform quantization compression on the first sample hidden vector to obtain a sample hidden vector, and add the sample hidden vector to the sample key-value cache; wherein, the service sample is a text sample; The sample processing module is used to obtain the sample hidden vector from the sample key-value cache, perform quantization restoration on the sample hidden vector to obtain a second hidden vector, determine the product of the first key projection matrix in the second initial projection matrix and the second hidden vector as the first sample key vector, determine the product of the first value projection matrix in the second initial projection matrix and the second hidden vector as the first sample value vector, and perform attention processing on the first sample key vector and the first sample value vector to obtain a second sample hidden state; The parameter processing module is used to obtain sample transfer parameters from the first sample key vector, the first sample value vector, and the sample hidden vector, and transfer the sample transfer parameters to the second initial processing network in the initial processing model; The sample processing module is further used to determine a second sample key vector and a second sample value vector based on the sample transfer parameters in the second initial processing network, and perform attention processing on the second sample key vector and the second sample value vector based on the second sample hidden state to obtain a third sample hidden state; The sample prediction module is used to perform information restoration on the third sample hidden state to obtain a sample processing result for the service sample when the second initial processing network is the last initial processing network; A model training module, configured to adjust parameters of the initial processing model based on the sample processing result to obtain a service processing model with converged parameters; the service processing model includes a first projection matrix corresponding to the first initial projection matrix and a second projection matrix corresponding to the second initial projection matrix.

16. A computer device, characterized in that, It includes a processor, a memory, and an input / output interface; The processor is respectively connected to the memory and the input / output interface. Among them, the input / output interface is used to receive and output data, the memory is used to store computer programs, and the processor is used to call the computer programs so that the computer device executes the method according to any one of claims 1-10, or executes the method according to any one of claims 11-13.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is adapted to be loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-10, or executes the method according to any one of claims 11-13.

18. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the method according to any one of claims 1-10 is implemented, or the method according to any one of claims 11-13 is executed.