Inference methods, devices, media, and program products for attention models
By generating key-value data summaries through collaborative training and combining hierarchical storage and data prefetching mechanisms, the memory and bandwidth bottlenecks of attention models in ultra-long text processing are solved, achieving efficient and accurate inference operations.
Patent Information
- Application Number
- CN202511212955.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-28
AI Technical Summary
When attention models process very long texts, the storage volume of key-value caches grows exponentially, leading to memory capacity and bandwidth bottlenecks, which affect inference efficiency and accuracy.
A first model trained collaboratively is used to generate key-value data summaries. Key-value cache data blocks are dynamically matched by query encoder and summary encoder. Combined with hierarchical storage and data prefetching mechanism, access and management of key-value cache are optimized.
It improves the inference efficiency and accuracy of the attention model when processing ultra-long texts, breaks through the GPU memory limit, and ensures high-performance operation of the model in ultra-long contexts.
Smart Images

Figure CN120723895B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically to a reasoning method, device, medium, and program product for an attention model. Background Technology
[0002] Attention models employ a Transformer architecture, using an autoregressive generation paradigm to predict and synthesize new lexical units one by one. To reduce the repetitive encoding of historical sequences in each generation step, relevant examples typically use a key-value caching mechanism to optimize the inference efficiency of attention models.
[0003] However, when attention models need to process extremely long texts, the storage volume of the key-value cache grows exponentially. How to achieve efficient access to the key-value cache data during inference becomes a decisive factor limiting the inference efficiency of attention models. Summary of the Invention
[0004] In view of the above problems, this application provides an inference method, device, medium and program product for attention models that improves the efficiency of accessing key-value cached data during the inference process of attention models.
[0005] According to a first aspect of this application, an inference method for an attention model is provided, comprising: processing a first text block using a query encoder of a first model to generate first query information; the first query information having the same dimension as a key-value data digest; wherein the first text block is generated by the attention model performing an nth inference operation; the key-value data digest is obtained by processing a key-value cached data block using a digest encoder of the first model; the key-value data digest indicates the importance of the key-value cached data block to the text block in the inference process of the attention model; the key-value cached data block is obtained by performing attention calculation on multiple text blocks in the prompt text using the attention model; the query encoder and the digest encoder of the first model are co-trained using the importance of the sample key-value cached data block to the sample text block in the historical inference process generated by the attention model during the execution of a historical inference task as a first label; determining first key-value data matching the first text block from multiple key-value cached data blocks based on the first query information and multiple key-value data digests; and performing an (n+1)th inference operation using the attention model based on the first key-value data, where n is an integer greater than or equal to 1.
[0006] A second aspect of this application provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0007] A third aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0008] A fourth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method. Attached Figure Description
[0009] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments of this application with reference to the accompanying drawings.
[0010] Figure 1 The diagram illustrates application scenarios of inference methods, apparatuses, devices, media, and program products based on the attention model according to embodiments of this application.
[0011] Figure 2 A flowchart of an inference method for an attention model according to an embodiment of this application is shown.
[0012] Figure 3 A schematic diagram of key-value cache data according to an embodiment of this application is shown.
[0013] Figure 4 A schematic diagram of a training method for a first model according to an embodiment of this application is shown.
[0014] Figure 5 A schematic diagram illustrating the determination of first key value data according to an embodiment of this application is shown.
[0015] Figure 6 A schematic diagram of an inference method for an attention model according to another embodiment of this application is shown.
[0016] Figure 7 A schematic diagram of a training method for a second model according to an embodiment of this application is shown.
[0017] Figure 8 A schematic diagram of prefetch key-value cache data according to an embodiment of this application is shown.
[0018] Figure 9 A structural block diagram of an inference apparatus for an attention model according to an embodiment of this application is shown.
[0019] Figure 10 A block diagram of an electronic device suitable for implementing an attention model inference method according to an embodiment of this application is shown. Detailed Implementation
[0020] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0021] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0023] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0024] Attention models are based on the Transformer model and have become a key component of artificial intelligence. Their architecture typically uses the Transformer model, employing an autoregressive generation paradigm to predict and synthesize new tokens (also known as text blocks) one by one. When generating the (N+1)th token in the sequence, the model must refer to and integrate all the information contained in the previous N tokens; this process is driven by the core self-attention mechanism. To avoid repetitive and computationally expensive encoding operations on the historical sequence in each generation step, inference systems commonly use a key-value cache (KV Cache) mechanism to optimize computational efficiency. Here, K and V represent "key" and "value," respectively, serving as intermediate representations in the attention mechanism. The key represents the importance or relevance of each position in the input sequence, while the value represents the actual information or content of each position. The key-value cache mechanism caches the key and value vectors of each historical token. Therefore, when generating a new token, it is only necessary to compute the "Query" vector of the current token and perform attention operations on it with all historical key-value pairs in the cache.
[0025] However, when the context length processed by attention models expands to hundreds of thousands or even millions of tokens, the storage volume of key-value cache data exhibits an exponential growth trend. For example, a model with 7 bytes of parameters processing 128KB of context requires over 60GB of storage space for its key-value cache data, significantly exceeding the capacity of the high-bandwidth memory (HBM) on current mainstream GPUs. This phenomenon raises two core technical challenges: 1) Memory capacity bottleneck: HBM cannot fully accommodate the key-value cache data corresponding to ultra-long contexts; 2) Memory bandwidth bottleneck: Even with a moderate context length and HBM's capacity to accommodate key-value cache data, each decoding iteration still requires loading massive amounts of data from HBM, resulting in input / output latency becoming a major performance constraint. Therefore, efficient management and access to this large-scale key-value cache data constitute a decisive factor in achieving high-performance inference for ultra-long context attention models.
[0026] In related examples, query-agnostic sparsity methods can be used to determine whether to retain or discard key-value cache data based on fixed rules or historical information. The main limitation of this type of method is that its discard decisions are permanent, and a token that was historically less important may be crucial for a specific future query. Once a token is incorrectly discarded, the model will permanently lose relevant information, potentially leading to a significant performance degradation in complex tasks requiring long-running dependencies.
[0027] In view of this, embodiments of this application provide an inference method for an attention model. A co-trained first model generates key-value data summaries for each key-value cache data block. These key-value data summaries accurately indicate the importance of the key-value cache data block to the text block during the inference process of the attention model. During the attention model inference process, the matching degree between any key-value data summary and the first query information corresponding to the first text block can be dynamically determined based on the first text block generated in the previous inference operation. This accurately selects the first key-value data that matches the first text block for the next inference operation. Under the same sparsity, this method maximizes the preservation of key information in the cached key-value data blocks, further improving the inference efficiency and accuracy of the attention model.
[0028] Figure 1 The diagram illustrates application scenarios of inference methods, devices, media, and program products based on the attention model according to embodiments of this application.
[0029] like Figure 1 As shown, the GPU (Graphics Processing Unit) is primarily used to perform training and inference operations for the attention model. The GPU may include a GPU computing module 101 and high-bandwidth memory on the GPU. The CPU scheduling module 103 is used to dynamically schedule key-value cache data blocks.
[0030] Before inference, a first cache region 110 and a second cache region 120 can be partitioned in the high-bandwidth memory of the GPU. The second cache region 120 is used to store key-value cache data. The key-value cache data can be divided into multiple key-value cache data blocks according to a preset size, and any key-value cache data block is called a Page. Any key-value cache data block includes key data (key vector) and value data (value vector). For example: K Page I1 represents the first key vector in the I-th key-value cache data block. V Page I1 represents the first value vector in the I-th key-value cache data block. The first cache region 110 is used to store the key-value data digest corresponding to any key-value cache data block. The key-value data digest is generated during the pre-filling stage of the attention model by processing any key-value cache data block using the digest encoder of the first model 102.
[0031] During the inference process of the attention model, based on the first text block generated by the GPU computing module 101 during the nth inference operation, the query encoder of the first model 102 generates first query information. n is an integer greater than or equal to 1. Then, the first query information is matched with the key-value data digest stored in the first cache area 110. The matching result can be fed back to the CPU scheduling module 103, so that the CPU scheduling module can load the first key-value data matching the first text block from the target storage area 130 to the second cache area 120 as needed based on the matching result.
[0032] The GPU computing module 101 can read the first key-value data from the second cache area 120 and perform the (n+1)th inference operation of the attention model.
[0033] Because the high-bandwidth memory storage space on the GPU is limited, the number of key-value cache data blocks stored in the second cache area 120 increases exponentially with the number of inference operations performed iteratively by the attention model. To reduce the GPU's memory usage, some key-value cache data blocks can be offloaded to the target storage area 130. The target storage area can be the CPU's system memory or the CPU's extended memory, such as a hard disk.
[0034] After the (n+1)th inference operation is completed, the new key-value cache data block generated for the first text block can be stored in the second cache area or the target storage area. Alternatively, the digest encoder of the first model 102 can be used to process the new key-value cache data block to generate a new key-value data digest, which is then stored in the first target cache area. This allows for continuous updating of the key-value cache data block and the key-value data digest during the inference process of the attention model.
[0035] The following will be based on Figure 1 The described scene, through Figures 2-8 The inference method of the attention model in the embodiments of this application will be described in detail.
[0036] Figure 2 A flowchart of an inference method for an attention model according to an embodiment of this application is shown.
[0037] like Figure 2 As shown, the inference method of the attention model in this embodiment includes operations S210 to S230.
[0038] In operation S210, the query encoder of the first model is used to process the first text block to generate the first query information.
[0039] In operation S220, based on the first query information and multiple key-value data digests, the first key-value data that matches the first text block is determined from multiple key-value cache data blocks.
[0040] In operation S230, the attention model is used to perform the (n+1)th inference operation based on the first key-value data.
[0041] Before the attention model performs inference operations, it first enters a pre-filling and initialization phase. This pre-filling and initialization phase is executed only once when the attention model processes a new inference request.
[0042] During the pre-filling and initialization phases, the attention model performs attention calculations on multiple text blocks in the prompt text in parallel, calculating the key vector and value vector of each text block (token) in all input prompt texts at once, forming the initial key-value cache data.
[0043] The initial key-value cache data can be divided into multiple key-value cache data blocks (Pages) according to a predetermined size. Therefore, the key-value cache data blocks are obtained by using an attention model to perform attention calculations on multiple text blocks in the prompt text.
[0044] Figure 3 A schematic diagram of key-value cache data according to an embodiment of this application is shown.
[0045] like Figure 3 As shown, the initial key-value cache data is divided into 3 key-value cache data blocks. Each key-value cache data block includes a key vector and a value vector for 4 text blocks.
[0046] In self-attention mechanisms, the importance of the key vector of a historical text block to the current text block can be determined by the analysis terms in the Softmax function. Therefore, the importance of any key-value cached data block to the current text block can be quantified by the sum of the contributions of the text blocks corresponding to any key-value cached data block, that is: .
[0047] Because directly calculating the sum of the contributions of the text blocks corresponding to any key-value cache data block requires excessive computational resources, this embodiment approximates the sum of the contributions of the text blocks corresponding to any key-value cache data block by using the dot product of the key-value data digest output by the digest encoder and the query information output by the query encoder. This is achieved through a collaboratively trained learnable digest encoder. and query encoder With relatively low computational resources, it can accurately approximate the importance of any key-value cache data block to the current text block.
[0048] (1)
[0049] Where p represents a key-value cache data block; q represents a text block; k represents the key vector in the key-value cache data block; and d represents the dimension of the key vector. This represents the cumulative sum of the contributions of the text block corresponding to any key-value cached data block i; This represents the dot product of the key-value data digest output by the digest encoder and the query information output by the query encoder. This indicates the query information output by the encoder; This represents the key-value data digest of the key-value cached data block output by the digest encoder.
[0050] In order to retain the key information in each key-value cache data block to the greatest extent under the same sparsity, in the embodiments of this application, the query encoder and the summary encoder of the first model are obtained by co-training using the importance of the sample text block in the historical reasoning process of the sample key-value cache data block generated by the attention model during the execution of the historical reasoning task as the first label.
[0051] In this embodiment, the first model's summary encoder can adopt a feature fusion network architecture of the Transformer model, and the first model's query encoder can adopt a multilayer perceptron. The importance of the sample text block in the historical reasoning process generated by the attention model during the execution of the historical reasoning task is used as the first label for supervised collaborative training to obtain the trained first model.
[0052] The summary encoder and query encoder of the first model can also adopt other neural network architectures, and this application embodiment does not specifically limit them.
[0053] For any key-value cached data block, a key-value data summary is generated using the summary encoder of the trained first model. This key-value data summary is a fixed-length summary vector that reflects the content of the key-value cached data block with high fidelity. The key-value data summary indicates the importance of the key-value cached data block to the text block in the inference process of the attention model.
[0054] During the inference phase of the attention model, iterative inference operations are performed. For example, the first text block generated by the nth inference operation of the attention model, i.e., the new token, can be processed by the query encoder of the first model to generate the first query. The purpose of this step is to transform the first text block into the first query information of the same dimension as the key-value data digest, so as to perform subsequent dot product operations.
[0055] Then, the first query information can be batch-multiplied with multiple key-value data digests according to formula (1), and the first key-value data digest matching the first text block can be determined based on the result of the dot product operation. Then, based on the first key-value data digest, the first key-value data matching the first text block can be determined from multiple key-value cache data blocks.
[0056] Next, the attention model is used to perform the (n+1)th inference operation based on the first key-value data.
[0057] This application provides an inference method for an attention model. A co-trained first model generates key-value data summaries for each key-value cache data block. These summaries accurately indicate the importance of the key-value cache data block to the text block during the attention model's inference process. During the attention model's inference process, the matching degree between any key-value data summary and the first query information corresponding to the first text block can be dynamically determined based on the first text block generated in the previous inference operation. This accurately selects the first key-value data that matches the first text block for the next inference operation. Under the same sparsity, this method maximizes the preservation of key information in the cached key-value data blocks, further improving the inference efficiency and accuracy of the attention model.
[0058] The following combination Figure 4 The training method of the first model in the embodiments of this application will be described in detail.
[0059] Figure 4 A schematic diagram of a training method for a first model according to an embodiment of this application is shown.
[0060] like Figure 4 As shown, the first model 102 may include a summary encoder 1021 and a query encoder 1022. The summary encoder includes a feature fusion module and a linear mapping module. The query encoder includes a multilayer perceptron module.
[0061] According to an embodiment of this application, the first model 102 is trained in the following manner: using a feature fusion module, feature fusion is performed by capturing the correlation between key data in any sample key-value cache data block to generate sample fusion features between key data in any sample key-value cache data block; using a linear mapping module, the sample fusion features between key data in any sample key-value cache data block are linearly mapped to generate a sample key-value data summary of any sample key-value cache data block; using a multilayer perception module, the sample text block is processed to generate sample query information; based on a first loss function, the first initial model is trained according to the sample key-value data summary, sample query information and first label of any key-value cache data block to generate the first model.
[0062] In this embodiment, the first initial model has the same model structure as the first model, but different model parameters. Through continuous iterative training, the model parameters of the first initial model are adjusted until the convergence condition is met, thus generating the first model. The convergence condition includes, but is not limited to, loss value convergence or reaching a predetermined number of iterations.
[0063] Sample data can be collected from inference logs generated by the attention model during multiple historical inference tasks. Sample data may include multiple sample pairs consisting of sample query information and sample key-value pairs. Historical inference tasks include, but are not limited to, code generation tasks and long-text question answering tasks.
[0064] like Figure 4 As shown, firstly, sample Page1 is input into digest encoder 1021, which outputs a key-value data digest of sample Page1. Simultaneously, sample Token1 can be input into query encoder 1022, which outputs sample query information.
[0065] Then, according to formula (1), the key-value data summary of sample Page1 and the sample query information are multiplied by a dot product to generate the sample importance.
[0066] Next, the loss value between sample importance and the first label is calculated using the first loss function. Based on the loss value, the model parameters of the first initial model are adjusted until the convergence condition is met, resulting in the first model 102.
[0067] In some embodiments, the first loss function may be the mean squared error loss function. To optimize the model parameters of the first initial model. Among them, This represents the expected parameter, for example, it could be the reciprocal of the number of key-value cache data blocks. This represents the sum of contributions of the text block corresponding to any key-value cache data block i. During training, this value can characterize the importance of the sample key-value cache data blocks generated by the attention model to the sample text blocks in the historical reasoning process.
[0068] In some embodiments, the feature fusion module can be any neural network unit capable of capturing the correlation between key data in any sample key-value cache data block and performing feature fusion. For example, a multi-size feature fusion network (FPN).
[0069] In some embodiments, the feature fusion module may also be a neural network unit that performs feature fusion based on an attention mechanism to capture the correlation between key data in any sample key-value cache data block. For example, the feature fusion module may include: a multi-head self-attention unit, a feedforward network unit, a layer normalization unit, and a global average pooling unit.
[0070] The input to a Multi-Head Self-Attention (MHSA) unit can be the key data in any sample key-value cache block; that is, the set of key vectors included in the sample key-value cache block can be represented as a matrix. Where N is the size of the sample key-value cache data block; d is the dimension of the key vector. A multi-head self-attention mechanism is used to capture the relationships between the key vectors within this sample key-value cache data block. , This indicates the correlation characteristics between the key data within the sample key-value cache data block.
[0071] The Feed-Forward Network (FFN) unit performs a nonlinear transformation on the correlation features between key data within the sample key-value cache data block output by the multi-head self-attention unit, generating global features for each key data in any sample key-value cache data block. , This represents the global feature of each key data in any sample key-value cache data block. This global feature enhances the details of each key data that are ignored by the multi-head self-attention unit.
[0072] The Layer Normalization (LayerNorm) unit is used to normalize the global features of any sample key-value cache data block and each key data in any sample key-value cache data block, so as to stabilize the training process, prevent gradient vanishing during training, and generate normalized features. , This represents the normalized feature.
[0073] The global average pooling unit aggregates the normalized features into a single vector, generating sample fusion features for each key-value data in any sample key-value cache data block.
[0074] In some embodiments, the linear mapping module is used to perform linear mapping on the sample fusion features between each key data in any sample key-value cache data block to generate a sample key-value data summary for any sample key-value cache data block.
[0075] (2)
[0076] in, This represents the digest of the sample key-value data; N represents the size of the sample key-value cache data block; represents the learnable projection matrix, and h represents the dimension of the key-value data summary.
[0077] therefore, .
[0078] In some embodiments, the query encoder A multilayer perceptron can be used to map the input sample text blocks (tokens) to the same feature space as the key-value data digest, generating query information with the same dimensions as the key-value data digest.
[0079] For example: (3)
[0080] in, Indicates sample query information; Represents the weight matrix. and For bias vectors, is the dimension of the hidden layer, and ReLU is the modified linear unit activation function.
[0081] In the pre-filling stage, the key-value cache data block is processed using the summary encoder of the first model to generate a key-value data summary, which is the same operation as generating sample key-value data summaries during the training of the first model.
[0082] For example, the feature fusion module is used to capture the correlation between the key data in any key-value cache data block and perform feature fusion to generate the fused features between the key data in any key-value cache data block; and the linear mapping module is used to perform linear mapping on the fused features between the key data in any key-value cache data block to generate the key-value data summary of any key-value cache data block.
[0083] By capturing the correlation between the key data in any key-value cache data block and performing feature fusion, the key information in the corresponding text block is preserved to the greatest extent possible under the condition that the key-value cache data blocks have the same sparsity. This information is then used to perform model inference operations, further improving the accuracy of model inference.
[0084] By using the sample key-value cache data blocks generated by the attention model during the historical reasoning task as the first label for the importance of the sample text blocks in the historical reasoning process, the summary encoder and query encoder are trained in a collaborative manner. This allows the output results of the trained summary encoder and query encoder to accurately represent the true distribution of the attention calculation results. As a result, key-value data can be dynamically matched to newly generated text blocks during the model reasoning process, which improves the reasoning efficiency of the attention model while ensuring the reasoning accuracy of the attention model.
[0085] According to an embodiment of this application, determining the first key-value data matching the first text block from multiple key-value cache data blocks based on first query information and multiple key-value data digests may include the following operations: obtaining the matching degree between the first query information and any key-value data digest from the first query information and multiple key-value data digests; determining the first key-value data digest matching the first query information from multiple key-value data digests based on multiple matching degrees; and determining the key-value data corresponding to the first key-value data digest from the key-value cache data as the first key-value data.
[0086] Figure 5 A schematic diagram illustrating the determination of first key value data according to an embodiment of this application is shown.
[0087] like Figure 5 As shown, for the first text block Token N generated by the nth inference operation, the query encoder 1022 is used to process Token N to generate the first query information.
[0088] Then, based on the first query information, a batch dot product operation can be performed on all key-value data digests stored in the first cache area 110 to obtain the importance corresponding to each key-value data digest. This importance characterizes the matching degree between the first query information and any key-value data digest. For example, the first query information and the key-value data digest S1 of Page1 are multiplied by a dot product to obtain the importance of Page1.
[0089] Next, based on the importance of each key-value data digest, the Page that matches the first query information can be determined from the first cache area 110. i Key-value data digest S i .
[0090] There is a correspondence between key-value data digests and key-value cache data blocks. Based on this correspondence, the key-value data corresponding to the first key-value data digest can be determined from the key-value cache data stored in the target storage area 130. For example, the first key-value data could be a Page. i Key-value cached data.
[0091] Since the key-value data digests of each key-value cache data block accurately reflect the true attention contribution of the corresponding text block, the key-value data digests can be dynamically matched from the first cache data using the first query information. This allows for accurate determination of the key-value cache data block corresponding to the key-value data digest, further improving the output accuracy of the attention model.
[0092] Because GPUs have limited high-bandwidth memory storage space, when attention models process very long texts, the storage volume of key-value cache data will grow exponentially during iterative inference operations.
[0093] Therefore, this application embodiment adopts a hierarchical storage mechanism, storing the key-value data digest in the first cache area; storing the key-value cache data block in the second cache area; and storing the target data block in the key-value cache data block in the target storage area based on the unloading strategy when the available storage space in the second cache area is less than or equal to a predetermined threshold.
[0094] In some embodiments, since the key-value data digest occupies a small amount of storage space, the key-value data digest is stored in a first cache area so that the attention model can directly read the key-value data digest from the first cache area and perform a dot product operation with the first query information each time it performs an inference operation, thereby improving retrieval efficiency.
[0095] In some embodiments, key-value cache data blocks may be stored in a second cache area first. When the available storage space in the second cache area is less than or equal to a predetermined threshold, the target data block in the key-value cache data block is stored in the target storage area based on an unloading strategy.
[0096] In this embodiment of the application, the uninstallation strategy can be configured based on the needs of the application scenario, and this embodiment of the application does not impose any specific limitations on it.
[0097] For example, in text generation tasks, for a text sequence, the key-value cache data block corresponding to the first text block provides a structural anchor point for the global context of the entire text sequence, while the key-value cache data block corresponding to the last text block has a high relevance to the generated text block. In subsequent iterations, these key-value cache data blocks are more likely to be selected. Therefore, the key-value cache data blocks corresponding to the first and last text blocks can be stored in a second cache area, which can be called the hot page cache area of the GPU's high-bandwidth memory. This ensures the core performance of the attention model with minimal storage cost.
[0098] The key-value cache data block corresponding to the text block in the middle usually contains a lot of redundant or easily replaceable information. Therefore, it can be stored in the target storage area as the target data block.
[0099] For example, the unloading strategy can also be determined based on the frequency with which key-value cached data blocks within the second target cache region are accessed within a predetermined time period. For instance, target data blocks that are accessed less frequently than a predetermined frequency threshold within the predetermined time period can be preferentially unloaded.
[0100] In some embodiments, during the initialization phase of the attention model, the key-value cache data blocks corresponding to the text blocks located at the beginning and end of the prompt text can be directly stored in the second target cache area, and the key-value cache data blocks corresponding to the text blocks located in the middle of the prompt text can be stored in the target storage area. This establishes a complete initial state for the autoregressive decoding phase of the attention model.
[0101] This application employs a hierarchical caching mechanism of key-value data digests and key-value cached data blocks. Only lightweight key-value data digests of all key-value cached data blocks are stored in the high-bandwidth memory of the GPU, while the vast majority of complete key-value cached data block pages are offloaded to CPU memory or CXL (Compute Express Link) extended memory. This enables the system to handle extremely long context sequences far exceeding the capacity of high-bandwidth memory, and ensures that even if the key-value cached data blocks are not on the GPU, they can still be dynamically and quickly matched with query information using lightweight key-value data digests, overcoming the storage limitations of the high-bandwidth memory of graphics processors.
[0102] According to an embodiment of this application, storing a target data block in a key-value cache data block in a target storage area may include the following operations: quantizing any target data block based on the quantization parameters of any target data block; and performing an unloading operation on any quantized target data block and storing the quantized target data block in the target storage area.
[0103] In this embodiment, compressing the target data block before unloading it can improve data transmission speed. When there are multiple target data blocks to be unloaded, quantization parameters can be calculated individually for each target data block. These quantization parameters include, but are not limited to, scaling factors and zero points. Then, each target data block is quantized according to its corresponding quantization parameters. Finally, the unloading operation is performed on the quantized target data blocks.
[0104] Based on the quantization parameters of any target data block, quantization is performed on any target data block, avoiding outlier dimensions caused by the large range of a few key-value cache data, which reduces the overall quantization accuracy. Compared with global quantization, it can more finely adapt to the data distribution of each key-value cache data block, achieving a better balance between compression efficiency and information guarantee, thereby further improving the information accuracy of key-value cache data blocks under the same compression rate.
[0105] Before performing the (n+1)th inference operation based on the first key-value data using the attention model, the storage location of the first key-value data can be determined first. When the first key-value data is located in the second cache area, the attention model can directly read the first key-value data from the second cache area to perform the inference operation.
[0106] In response to determining that the first key-value data is located in the target storage area, the first key-value data is loaded from the target storage area into the second cache area so that the attention model can read the first key-value data from the second cache area.
[0107] Since the key-value cache data blocks stored in the target storage area are all quantized, an inverse quantization operation needs to be performed first when determining that the first key-value data comes from the target storage area.
[0108] After loading the first key-value data into the second cache area, the above method may further include the following operation: dequantizing the quantized first key-value data based on the quantization parameters used for the first key-value data to generate the first key-value data.
[0109] For example, you can call the dequantization kernel function, which calls the quantization parameters used to quantize the first key-value data, and dequantize the quantized first key-value data to restore it to its original precision.
[0110] For key-value cache data blocks from the target storage area, due to the need to perform dequantization, the GPU needs to wait for the dequantization operation to complete before it can read the key-value cache data blocks from the second cache area and then perform inference operations, compared to key-value cache data blocks from the second cache area.
[0111] The hierarchical storage mechanism proposed in this application overcomes the storage limitations of the GPU's high-bandwidth memory; however, it also introduces latency in cross-level data transfer. Therefore, this application proposes a data prefetching mechanism, which allocates a third cache region on the GPU's high-bandwidth memory to store prefetched key-value cache data blocks. The purpose of the data prefetching mechanism is to preload the key-value cache data blocks required for the (n+2)th inference operation into the third cache region during the (n+1)th inference operation. This reduces the time the GPU spends waiting for the key-value cache data blocks to be transferred from the target storage region to the second cache region and the time spent waiting for the inverse quantization operation of the key-value cache data blocks during the (n+2)th inference operation.
[0112] The following will combine Figures 6-8 The data prefetching mechanism is explained in detail.
[0113] Figure 6 A schematic diagram of an inference method for an attention model according to another embodiment of this application is shown.
[0114] like Figure 6 As shown, with Figure 1 Compared to the illustrated embodiment, this embodiment adds a third cache area 140 for storing key-value cache data blocks loaded from the target storage area 130. In the data prefetching mechanism, the attention model has not yet completed the (n+1)th inference operation and cannot determine the required key-value cache data blocks to be prefetched based on the generated text block query key-value data digest.
[0115] Because the semantic focus of the output text sequence evolves smoothly during the decoding process of the attention model, it means that consecutive query vectors form a smooth trajectory in the high-order embedding space. Therefore, the historical key-value cache data blocks that the query vectors focus on also have a high degree of overlap and predictability.
[0116] Based on this, embodiments of this application utilize the temporal locality of attention distribution exhibited by the attention model when generating coherent text to train a second model. The trained second model is then used to predict the key-value cache data blocks required for the next inference operation based on the historical query sequence, thereby achieving data prefetching.
[0117] According to embodiments of this application, the above method may further include the following operations: using a second model, determining second key-value data from multiple key-value cache data blocks based on the historical query sequence in the previous n+1 inference operations; and during the n+1 inference operation performed by the attention model, preloading the second key-value data into a third cache area, so that after the attention model completes the n+1 inference operation, it reads the second key-value data from the third cache area and performs the n+2 inference operation based on the second key-value data.
[0118] In this embodiment, the data prefetching operation can be performed in parallel with the batch dot product operation of the data key-value summary and query information.
[0119] like Figure 6 As shown, the historical query sequence can be input into the second model 104, which outputs the probability that each key-value cache data block will be selected in the (n+2)th inference operation. Then, the CPU scheduling module 103 can determine the second key-value data from the target storage area 130 based on this probability, and preload the second key-value data into the third cache area 140 during the (n+1)th inference operation performed by the attention model based on the first key-value data.
[0120] After the second key-value data is loaded into the third cache area 140, the dequantization function can be called to dequantize the quantized second key-value data. At this time, the GPU is performing the (n+1)th inference operation. When the GPU starts performing the (n+2)th inference operation, the dequantization operation on the quantized second key-value data has been completed, thus actively hiding the data transmission latency inherent in the hierarchical storage architecture.
[0121] During the (n+1)th inference operation, the key-value cache data block required for the (n+2)th inference operation is preloaded into the third cache area. This can reduce the time the GPU spends waiting for the key-value cache data block to be transferred from the target storage area to the second cache area and the time spent waiting for the inverse quantization operation of the key-value cache data block during the (n+2)th inference operation.
[0122] The following combination Figure 7 The training method of the second model in the embodiments of this application will be described in detail.
[0123] Figure 7 A schematic diagram of a training method for a second model according to an embodiment of this application is shown.
[0124] like Figure 7 As shown, the second initial model may include a gated loop unit and a fully connected layer.
[0125] According to an embodiment of this application, the second model is trained in the following manner: a gated recurrent unit is used to process the sample query sequence to generate sample temporal dependency features; a fully connected layer is used to process the sample temporal dependency features to generate the expected matching probability of sample key-value data when the attention model performs inference operations on the sample query sequence; based on the second loss function, the second initial model is trained according to the expected matching probability and the second label to obtain the second model; the second label is the sample key-value data required by the attention model to perform inference operations on the sample query sequence.
[0126] In this embodiment, the training samples may include: a set of historical query sequences for each decoding step collected from the inference log and an index set of the key-value cache data blocks actually selected for each decoding step, obtained by performing a complete inference operation using an attention model on a benchmark task. .
[0127] Before training the second model, you can Transform into a multidimensional, multi-label binary vector When key-value cache data block j belongs to When =1, otherwise 0.
[0128] like Figure 7 As shown, the sample query sequence from the first t inference operations can be input into the second initial model, which outputs the expected matching probability. The loss value is calculated by comparing the actual matching probability and the expected matching probability of the sample key-value data required for the (t+1)th inference operation. This loss value is then used to train the second initial model until convergence is achieved, resulting in the second model. Convergence conditions include, but are not limited to, loss value convergence or reaching the maximum number of iterations. t is an integer greater than or equal to 1.
[0129] For example, the second initial model may include a gated loop unit 1041 and a fully connected layer 1042.
[0130] First, the sample query sequence of the previous t inference operations can be input into the gated recurrent unit 1041. The temporal dependency of the sample query sequence is captured by iterative updates, and the sample temporal dependency features that can summarize the temporal dependency information of the sample query sequence of the previous t inference operations are output.
[0131] Then, the temporal dependency features of the samples are passed through a fully connected layer, and the Sigmoid activation function is used to generate a probability vector with the dimension of the number of key-value cache data blocks, which is the expected matching probability of each key-value cache data block being selected in the (t+1)th inference operation.
[0132] (4)
[0133] in, and These are learnable parameters. yes function. Indicates the number of key-value cache data blocks; Represents the temporal dependency features of the samples; Indicates the dimension of the key-value cache data block. This represents the expected matching probability that the j-th key-value cache data block will be selected in the (t+1)-th inference operation.
[0134] In this embodiment of the application, the second loss function may be a binary cross-entropy loss function. As shown in equation (5):
[0135] (5)
[0136] in, Indicates the number of key-value cache data blocks; This represents the expected matching probability that the j-th key-value cache data block will be selected in the (t+1)-th inference operation; This represents the actual matching probability that the j-th key-value cache data block is selected in the (t+1)-th inference operation.
[0137] During training, the model parameters of the second initial model are adjusted based on the loss value until the loss value is minimized, thus obtaining the second model.
[0138] The second initial model has the same model structure as the second model, only the model parameters are different.
[0139] The operation of processing the historical query sequence in the first n+1 inference operations using the second model is the same as the operation of processing the sample query sequence during the training of the second initial model.
[0140] For example, using the second model, based on the historical query sequence in the first n+1 inference operations, to determine the second key-value data from multiple key-value cache data blocks, may include the following operations: using a gated recurrent unit to process the historical query sequence in the first n+1 inference operations to generate temporal dependency features; using a fully connected layer to process the temporal dependency features to generate the probability that any key-value cache data block is determined to be used to perform the n+2th inference operation; and based on the probability, determining the second key-value data from multiple key-value cache data blocks.
[0141] Figure 8 A schematic diagram of prefetch key-value cache data according to an embodiment of this application is shown.
[0142] like Figure 8 As shown, the gated recurrent unit 1041 of the second model 104 processes the historical query sequence of the previous n inference operations to generate temporal dependency features. Then, the fully connected layer of the second model 104 processes the temporal dependency features to generate the expected matching probability of each key-value cache data block in the target storage area 130 being determined for performing the (n+2)th inference operation.
[0143] Next, the key-value cache data blocks can be sorted from high to low based on the expected matching probability, and the first M key-value cache data blocks can be determined as the second key-value data. M is an integer greater than or equal to 1, and the value of M can be configured according to the actual application scenario requirements, without specific limitations here.
[0144] For example, the second key-value cache data can be the Page2 key-value cache data, and during the (n+1)th inference operation of the attention model, the Page2 key-value cache data is loaded into the third cache area 140.
[0145] Compared to passively waiting for data loading in related examples, this embodiment utilizes the local temporal correlation characteristics of attention distribution. It uses a trained second model to predict the second key-value data required for the n+2th inference operation based on the historical query sequence of the previous n+1 inference operations, and loads the second key-value data into the third cache area. This achieves the overlap of computation and data transmission processes, actively hides the data transmission latency inherent in the hierarchical storage architecture, reduces inference latency, improves inference speed, and increases the throughput of the attention model.
[0146] According to the embodiments of this application, the above method may further include the following operations: The method further includes:
[0147] After completing the (n+1)th inference operation, the state of the second key-value data is updated from the prefetch state to the read-ahead state, so that the attention model can read the second key-value data from the third cache area.
[0148] Since both the third and second cache regions reside in the GPU's high-bandwidth memory, after completing the (n+1)th inference operation, the state of the second key-value data is updated from the prefetch state to the read-ahead state. This means that when the attention model reads the second key-value data from the third cache region, the CPU scheduling module only needs to move the data read pointer from the management list of the second cache region to the management list of the third cache region. Therefore, no physical data transfer is required, further reducing the data transfer latency caused by the hierarchical storage architecture.
[0149] Based on the inference method of the above attention model, this application also provides an inference device for the attention model. The following will combine... Figure 9 The device is described in detail.
[0150] Figure 9 A structural block diagram of an inference apparatus for an attention model according to an embodiment of this application is shown.
[0151] like Figure 9 As shown, the inference device 900 of the attention model in this embodiment includes a generation module 910, a determination module 920 and an inference module 930.
[0152] The generation module 910 is used to process the first text block using the query encoder of the first model to generate first query information; the first query information has the same dimension as the key-value data summary; wherein, the first text block is generated by the attention model performing the nth inference operation; the key-value data summary is obtained by processing the key-value cache data block using the summary encoder of the first model; the key-value data summary indicates the importance of the key-value cache data block to the text block in the inference process of the attention model; the key-value cache data block is obtained by performing attention calculation on multiple text blocks in the prompt text using the attention model; the query encoder and the summary encoder of the first model are obtained by co-training using the importance of the sample key-value cache data block to the sample text block in the historical inference process generated by the attention model during the execution of the historical inference task as the first label. In one embodiment, the generation module 910 can be used to perform the operation S210 described above, which will not be repeated here.
[0153] The determining module 920 is used to determine, based on the first query information and multiple key-value data digests, the first key-value data that matches the first text block from multiple key-value cache data blocks. In one embodiment, the determining module 920 can be used to perform the operation S220 described above, which will not be repeated here.
[0154] The inference module 930 is used to perform the (n+1)th inference operation based on the first key-value data using the attention model, where n is an integer greater than or equal to 1. In one embodiment, the inference module 930 can be used to perform the operation S230 described above, which will not be repeated here.
[0155] According to an embodiment of this application, the inference device 900 of the attention model further includes a feature fusion module and a linear mapping module.
[0156] The feature fusion module is used to capture the correlation between the key data in any key-value cache data block and perform feature fusion to generate fused features between the key data in any key-value cache data block.
[0157] The linear mapping module is used to perform linear mapping on the fusion features between the key data in any key-value cache data block, and generate a key-value data digest for any key-value cache data block.
[0158] According to an embodiment of this application, the determining module 920 may include: a calculation submodule, a matching submodule, and a determining submodule.
[0159] The calculation submodule is used to obtain the matching degree between the first query information and any key-value data digest from multiple key-value data digests.
[0160] The matching submodule is used to determine a first key-value data digest that matches the first query information from multiple key-value data digests based on multiple matching degrees.
[0161] The determination submodule is used to determine the key-value data corresponding to the first key-value data digest from multiple key-value cache data blocks as the first key-value data.
[0162] According to an embodiment of this application, the inference device 900 of the attention model further includes: a multi-layer perception module and a first training module.
[0163] The feature fusion module is used to perform feature fusion by capturing the correlation between each key data in any sample key-value cache data block, and generate sample fusion features between each key data in any sample key-value cache data block.
[0164] The linear mapping module is used to perform linear mapping on the sample fusion features between the key data in any sample key-value cache data block, and generate a sample key-value data summary for any sample key-value cache data block.
[0165] The multi-layer perception module is used to process sample text blocks and generate sample query information; the sample query information has the same dimension as the sample key-value data summary.
[0166] The first training module is used to train the first initial model based on the first loss function, according to the sample key-value data summary, sample query information and first label of any key-value cache data block, to generate the first model.
[0167] According to an embodiment of this application, the inference device 900 for the attention model further includes: a first storage module, a second storage module, and a third storage module.
[0168] The first storage module is used to store key-value data digests in the first cache area.
[0169] The second storage module is used to store key-value cache data blocks in the second cache area.
[0170] The third storage module is used to store the target data block in the key-value cache data block in the target storage area based on an offloading strategy when the available storage space in the second cache area is less than or equal to a predetermined threshold.
[0171] According to an embodiment of this application, the inference apparatus 900 of the attention model further includes: a loading module, configured to, in response to determining that the first key-value data is located in the target storage area, load the first key-value data from the target storage area to the second cache area before performing the (n+1)th inference operation based on the first key-value data using the attention model, so that the attention model reads the first key-value data from the second cache area.
[0172] According to an embodiment of this application, the third storage module includes a quantization submodule and an unloading submodule.
[0173] The quantization submodule is used to quantize any target data block based on the quantization parameters of any target data block.
[0174] The unload submodule is used to perform an unload operation on any quantized target data block and store the quantized target data block in the target storage area.
[0175] According to an embodiment of this application, the inference device 900 of the attention model further includes: a dequantization module, used to dequantize the quantized first key-value data based on the quantization parameters used for the first key-value data, to generate the first key-value data.
[0176] According to an embodiment of this application, the inference device 900 of the attention model further includes a prediction module and a pre-loading module.
[0177] The prediction module is used to determine the second key-value data from the key-value cache data based on the historical query sequence in the previous n+1 inference operations using the second model.
[0178] The pre-loading module is used to pre-load the second key-value data into the third cache area during the (n+1)th inference operation of the attention model, so that after the (n+1)th inference operation is completed, the attention model can read the second key-value data from the third cache area and perform the (n+2)th inference operation based on the second key-value data.
[0179] According to an embodiment of this application, the inference device 900 of the attention model further includes: an update module, configured to update the state of the second key-value data from a prefetch state to a read-read state after the (n+1)th inference operation is completed, so that the attention model reads the second key-value data from the third cache area.
[0180] According to an embodiment of this application, the inference device 900 for the attention model further includes: a first processing module, a second processing module, and a key value determination module.
[0181] The first processing module is used to process the historical query sequence in the first n+1 inference operations using a gated loop unit to generate temporal dependency features.
[0182] The second processing module is used to process the temporal dependency features using a fully connected layer, and generate the probability that any key-value cache data in the key-value cache data is determined to be used to perform the (n+2)th inference operation.
[0183] The key-value determination module is used to determine the second key-value data from the key-value cache data based on probability.
[0184] According to an embodiment of this application, the inference device 900 of the attention model further includes: a feature extraction module, a probability prediction module, and a second training module.
[0185] The feature extraction module is used to process the sample query sequence using a gated loop unit to generate sample time-dependent features.
[0186] The probability prediction module uses a fully connected layer to process the temporal dependency features of samples, generating sample key-value data that is then used by the attention model to perform inference operations on the sample query sequence to predict the matching probability.
[0187] The second training module is used to train the second initial model based on the second loss function, the expected matching probability, and the second label to obtain the second model; the second label is the sample key-value data required by the attention model to perform inference operations on the sample query sequence.
[0188] According to embodiments of this application, any plurality of modules among the generation module 910, determination module 920, and inference module 930 can be merged into one module, or any one of these modules can be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules can be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the generation module 910, determination module 920, and inference module 930 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any appropriate combination of any of these three implementation methods. Alternatively, at least one of the generation module 910, determination module 920, and inference module 930 can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.
[0189] Figure 10 A block diagram of an electronic device suitable for implementing an attention model inference method according to an embodiment of this application is shown.
[0190] like Figure 10As shown, an electronic device 1000 according to an embodiment of this application includes a processor 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage portion 1008 into a random access memory (RAM) 1003. The processor 1001 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1001 may also include onboard memory for caching purposes. The processor 1001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.
[0191] RAM 1003 stores various programs and data required for the operation of electronic device 1000. Processor 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Processor 1001 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 1002 and / or RAM 1003. It should be noted that the programs may also be stored in one or more memories other than ROM 1002 and RAM 1003. Processor 1001 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.
[0192] According to embodiments of this application, the electronic device 1000 may further include an input / output (I / O) interface 1005, which is also connected to a bus 1004. The electronic device 1000 may also include one or more of the following components connected to the input / output (I / O) interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the input / output (I / O) interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1010 as needed so that computer programs read from it can be installed into the storage section 1008 as needed.
[0193] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0194] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 1002 and / or RAM 1003 and / or one or more memories other than ROM 1002 and RAM 1003 described above.
[0195] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this application.
[0196] When the computer program is executed by the processor 1001, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0197] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 1009, and / or installed from a removable medium 1011. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0198] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0199] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0200] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0201] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
[0202] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.
Claims
1. A reasoning method for an attention model, characterized in that, The method includes: The query encoder of the first model processes the first text block to generate first query information; the first query information has the same dimension as the key-value data summary; wherein, the first text block is generated by the attention model performing the nth inference operation; the key-value data summary is obtained by processing the key-value cache data block using the summary encoder of the first model; the query encoder and the summary encoder of the first model are obtained by co-training using the importance of the sample text block in the historical inference process of the sample key-value cache data block generated by the attention model during the historical inference task as the first label; Based on the first query information and the multiple key-value data digests, determine the first key-value data that matches the first text block from the multiple key-value cache data blocks; and The attention model is used to perform the (n+1)th inference operation based on the first key-value data, where n is an integer greater than or equal to 1.
2. The method according to claim 1, characterized in that, The first model's summary encoder includes: a feature fusion module and a linear mapping module; The method further includes: Using the feature fusion module, feature fusion is performed by capturing the correlation between key data in any key-value cache data block to generate fused features between key data in any key-value cache data block; and Using the linear mapping module, the fusion features between the key data in any key-value cache data block are linearly mapped to generate a key-value data digest for any key-value cache data block.
3. The method according to claim 1, characterized in that, The step of determining the first key-value data matching the first text block from multiple key-value cache data blocks based on the first query information and multiple key-value data digests includes: Based on the first query information and any one of the key-value data digests among the multiple key-value data digests, the matching degree between the first query information and any one of the key-value data digests is obtained; Based on multiple matching degrees, a first key-value data digest matching the first query information is determined from multiple key-value data digests; and The key-value data corresponding to the first key-value data digest is determined from the plurality of key-value cache data blocks as the first key-value data.
4. The method according to claim 1, characterized in that, The summary encoder includes a feature fusion module and a linear mapping module; the query encoder includes a multilayer perceptron module; the first model is trained using the following method: Using the feature fusion module, the correlation between each key data in any sample key-value cache data block is captured and fused to generate sample fusion features between each key data in any sample key-value cache data block; Using the linear mapping module, the sample fusion features between each key data in any sample key-value cache data block are linearly mapped to generate a sample key-value data summary for any sample key-value cache data block; The multilayer perceptron module is used to process the sample text block to generate sample query information; wherein the sample query information has the same dimension as the sample key-value data summary; and Based on the first loss function, the first initial model is trained according to the sample key-value data summary of any key-value cache data block, the sample query information, and the first label to generate the first model.
5. The method according to claim 1, characterized in that, The method further includes: The key-value data digest is stored in the first cache area; The key-value cache data block is stored in the second cache area; If the available storage space in the second cache area is less than or equal to a predetermined threshold, the target data block in the key-value cache data block is stored in the target storage area based on the unloading strategy.
6. The method according to claim 5, characterized in that, Before performing the (n+1)th inference operation based on the first key-value data using the attention model, the method further includes: In response to determining that the first key-value data is located in the target storage area, the first key-value data is loaded from the target storage area into the second cache area, so that the attention model reads the first key-value data from the second cache area.
7. The method according to claim 5, characterized in that, The target data block includes multiple blocks; storing the target data blocks in the key-value cache data block in the target storage area includes: Based on the quantization parameters of any target data block, quantize any target data block; An unload operation is performed on any quantized target data block, and the quantized target data block is stored in the target storage area.
8. The method according to claim 6 or 7, characterized in that, After loading the first key-value data into the second cache area, the method further includes: Based on the quantization parameters used for the first key-value data, the quantized first key-value data is dequantized to generate the first key-value data.
9. The method according to claim 1, characterized in that, The method further includes: Using the second model, based on the historical query sequence in the first n+1 inference operations, the second key-value data is determined from multiple key-value cache data blocks; and During the (n+1)th inference operation performed by the attention model, the second key-value data is preloaded into the third cache area so that after the (n+1)th inference operation is completed, the attention model reads the second key-value data from the third cache area and performs the (n+2)th inference operation based on the second key-value data.
10. The method according to claim 9, characterized in that, The method further includes: After the (n+1)th inference operation is completed, the state of the second key-value data is updated from the prefetch state to the read-ahead state, so that the attention model can read the second key-value data from the third cache area.
11. The method according to claim 9, characterized in that, The second model includes: a gated loop unit and a fully connected layer; The step of using the second model to determine the second key-value data from multiple key-value cache data blocks based on the historical query sequence in the previous n+1 inference operations includes: The gated loop unit is used to process the historical query sequence in the first n+1 inference operations to generate temporal dependency features; Using the fully connected layer, the temporal dependency features are processed to generate the probability that any key-value cache data block is determined to be used for the (n+2)th inference operation; Based on the probability, the second key-value data is determined from the plurality of key-value cache data blocks.
12. The method according to claim 11, characterized in that, The second initial model includes gated recurrent units and fully connected layers; the second model is trained using the following method: The gated loop unit is used to process the sample query sequence and generate sample temporal dependency features. The fully connected layer is used to process the temporal dependency features of the sample to generate the expected matching probability of the sample key-value data being inferred by the attention model for the sample query sequence. Based on the second loss function, the second initial model is trained according to the expected matching probability and the second label to obtain the second model; the second label is the sample key-value data required by the attention model to perform inference operations on the sample query sequence.
13. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 12.
14. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 12.
15. A program product comprising: A computer program or instruction, characterized in that, when executed by a processor, the computer program or instruction implements the steps of any one of claims 1 to 12.
Citation Information
Patent Citations
Model reasoning method and device based on key value matrix cache and medium
CN118036754A
Model reasoning optimization method and device, equipment, storage medium and program product
CN119201476A