A Method, Device and Medium for Optimizing the Compilation Times of a Chatbot

By compensating and masking the input and historical cache of the large model inference stage, the problem of difficult to statically large model input in the existing technology is solved, and the effect of reducing the number of compilation times and improving the inference efficiency is achieved.

CN118092927BActive Publication Date: 2025-06-24SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410100506.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-06-24
Estimated Expiration
2044-01-24

AI Technical Summary

Technical Problem

The prior art cannot make the input of the large model static shape while ensuring the correct calculation results, resulting in poor compiler optimization results, especially in dynamic shape input.

Method used

By compensating the input prompt tokens in the big model inference stage and the key-value cache of historical inference, and masking data processing before exponential normalization, the big model structure and code input by the compiler are modified to make it more suitable for static input shapes.

Benefits of technology

The number of compilation times is reduced, the inference efficiency of the compilation optimization acceleration solution is improved, the correctness of data calculation is ensured, and the number of compilation optimization times and model optimization effects are balanced through flexible slice length settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118092927B_ABST
    Figure CN118092927B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, device and medium for optimizing the compilation times of a chatbot. The method is implemented based on a deep learning compiler. By modifying some processes of large model inference, including padding the input prompt tokens and the key-value cache of historical inferences in the large model inference stage, and performing masked data processing before exponential normalization. Compared with the prior art, the present invention has the advantages of reducing the execution times of the deep learning compiler with static input shapes, thereby improving the inference efficiency of the compilation optimization acceleration scheme, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to large model chatbots, and more particularly to a method, device, and medium for optimizing the compilation times of chatbots. Background Art

[0002] With the development of large model technology, chatbots developed using large model technology have begun to become popular.

[0003] As Figure 1 shown, the simplified process of a chatbot's answer includes: the user submits a statement to the chatbot; the chatbot processes the user statement into a token sequence; the chatbot uses a large model and takes the token sequence as input for inference to obtain an output token sequence; the chatbot performs textification and post-processing on the token sequence output by the large model to obtain text that can be understood by humans.

[0004] Starting from the user's input statement as the starting point, the above process finally returns the next sentence of the chat conversation, and this process can be repeated sequentially to allow the user to have a conversation with the chatbot.

[0005] Since large models have high computing power requirements, it is necessary to optimize the compilation of large model inferences to accelerate the inference speed of large models; current deep learning compilers have good optimization effects on models with static input shapes, but have relatively average effects on models with dynamic input shapes.

[0006] Therefore, the disadvantages of the prior art include:

[0007] 1) Without ensuring the correctness of the calculation results, the input of the model optimized by the compiler is in a static shape; simply padding the input tokens and filling zeros in the key-value cache will result in incorrect intermediate calculation results;

[0008] 2) After the sequence dimension of the historical key-value cache is filled with zeros as a fixed value, even when the sequence length of the historical key-value cache is relatively small, the reduction dimension of the matrix multiplication of the score and value in the self-attention mechanism module is the "fixed value of the sequence dimension of the historical key-value cache filled with zeros", resulting in a large amount of matrix multiplication calculations. Summary of the Invention

[0009] The purpose of the present invention is to overcome the above-mentioned defects of the prior art and provide a method, device, and medium for optimizing the compilation times of chatbots.

[0010] The purpose of the present invention can be achieved through the following technical solutions:

[0011] According to the first aspect of the present invention, there is provided a method for optimizing the compilation times of a chatbot, which is implemented based on a deep learning compiler. By modifying part of the process of large model inference, including padding the input prompt tokens and the key-value cache of historical inferences during the large model inference stage, and performing masked data processing before exponential normalization.

[0012] As a preferred technical solution, the padding of the input prompt tokens and the key-value cache of historical inferences during the large model inference stage includes:

[0013] 1) Padding the "token sequence converted from the user input statement" in the large model inference at the front side in the sequence dimension;

[0014] 2) Padding the "historical key cache" and "historical value cache" in the attention layer at the front side in the sequence dimension.

[0015] As a preferred technical solution, only the first "token sequence converted from the user input statement" is padded at the front side in the sequence dimension, and in the case where a single token generated by the large model is used as the input subsequently, the padding length is regarded as 0.

[0016] As a preferred technical solution, the sum of the length of the padded input sequence and the actual token sequence length is a preset fixed value.

[0017] As a preferred technical solution, the "historical key cache" and "historical value cache" are first concatenated with the "current key" and "current value" respectively, and then padded at the front side in the sequence dimension, and a cutting operation is performed after concatenation; the length of the cut is the sequence length of the concatenated "current key" and "current value", and after cutting, the "historical key cache for matrix multiplication" and "historical value cache for matrix multiplication" are obtained respectively.

[0018] As a preferred technical solution, the cache length N of the "historical key cache" and "historical value cache" is pre-calculated. The remaining corresponding historical cache sequence length without considering the padding length of the historical cache is denoted as real_cache_len, and this real_cache_len is not greater than N. The method sets a slice length list, which is composed of a series of positive integers from small to large in sequence, the maximum value is N and N is in the list, and is used to further slice the historical cache after the front cut; for each real_cache_len, find the smallest value in the list that is greater than or equal to real_cache_len and denote it as slice_size; for the historical cache after the front cut, retain a tensor with a length of slice_size along the rear side in the sequence dimension.

[0019] As a preferred technical solution, the mask data processing before exponential normalization is specifically as follows:

[0020] The mask data is a two-dimensional matrix with dimensions: (input padding sequence length + actual token sequence length) * new historical cache sequence length, where the new historical cache sequence length is: cache padding residual sequence length + historical cache sequence length + input padding sequence length + current cache sequence length;

[0021] For the new historical cache sequence dimension, set the values of the corresponding dimensions of "cache padding residual sequence length" and "input padding sequence length" in the mask data to -inf, and the rest to 0, where -inf is negative infinity;

[0022] For the other dimension of the mask data matrix, its length is (input padding sequence length + actual token sequence length), set the corresponding "input padding sequence length" in the mask data to -inf, and the rest to 0;

[0023] If a value is to be set to 0 and -inf at the same time, it is preferentially set to -inf.

[0024] As a preferred technical solution, the processed mask data is summed to obtain the "masked score", which is exponentially normalized to obtain the "exponentially normalized score", and then matrix multiplication is performed with the "historical value cache for matrix multiplication" to obtain the attention score.

[0025] According to the second aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein a computer program is stored on the memory, and the method is implemented when the processor executes the program.

[0026] According to the third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and the method is implemented when the program is executed by a processor.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] 1) By modifying the structure and code of the large model input to the compiler, the present invention makes it more friendly to the deep learning compiler with static input shapes, enabling it to reduce the number of compilations while maintaining the original calculation results;

[0029] 1) The present invention solves the problem in the process of compilation optimization for the inference of large model chatbots, reduces the number of executions of the deep learning compiler with static input shapes, and thus improves the inference efficiency of the compilation optimization acceleration scheme;

[0030] 2) For the special mask processing in the input token and key-value cache padding scheme of the present invention, the correctness of data calculation is ensured;

[0031] 3) By using the setting of "slice length list", for the actual lengths of different sequences of historical key-value caches, the self-attention mechanism module of the large model can perform smaller matrix multiplication operations. Each reduction dimension length will increase the number of compilations by one, but the corresponding inference process will operate faster. This setting can flexibly balance the number of compilation optimizations and the optimization effect of the compiled output model. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a simplified process schematic diagram of the chatbot's answer;

[0033] Figure 2 It is a schematic diagram of the calculation process for generating a single token;

[0034] Figure 3 It is the specific flow chart after the improvement of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0036] The present invention focuses on the inference optimization in the large model process and uses a deep learning compiler for static input shapes for optimization. The large model involved in the present invention specifically refers to a decoder-only model based on the attention mechanism. In order to generate a whole chat reply, the chatbot will perform a token sequence converted from the user input statement. Using the large model to infer once will obtain the next token, and at the same time update the historical key and value caches. If a reply with a length of n tokens needs to be generated, then n - 1 more large model inferences are required. The input token used in each large model inference is the output token of the previous inference, and the historical key-value cache is also used and updated. In addition, the large model inference also uses the mask required for the current inference.

[0037] As Figure 2 shown, the calculation process for generating a single token is as follows:

[0038] Figure 2On the left is a schematic diagram of the large model inference process. The word embedding layer converts the input tokens into vectors corresponding to the word embeddings, which are then processed through several decoder modules. Finally, the output tokens are obtained after passing through the normalization layer, fully connected layer, and sampling post-processing in sequence.

[0039] Figure 2 On the right is a simplified schematic diagram of the attention mechanism in the decoder, which simplifies the descriptions of position encoding, matrix transpose, and some multi-head attention. This figure mainly describes the attention operation of a single head of a single batch sample under the input of multi-batch processing of multi-head attention. Among them, "the token sequence converted from the user input statement, the current value, the current query, the current mask, the current value" are the input information for the inference of the current attention layer, and "the historical key cache, the historical value cache" are the historical cache information.

[0040] As Figure 3 shown, the present invention has made the following improvements on the "computing process for generating a single token" described in Figure 2 :

[0041] (1) For the "token sequence converted from the user input statement" in the large model inference, padding is performed on the front side in the sequence dimension, and the value used for padding can be arbitrary.

[0042] (2) For the "historical key mask" and "historical value mask" in the attention layer, padding is performed on the front side in the sequence dimension, and the value used for padding can be arbitrary.

[0043] (3) Special processing is performed on the "current mask" in the attention layer, and the processing method will be described in detail in the following content.

[0044] (4) A new parameter "slice length" is added in the attention layer to reduce the amount of actual matrix multiplication operations.

[0045] The following is a detailed description of each improved part:

[0046] In the large model inference process, the present invention performs front-side padding on the token sequence converted from the first user input statement in the sequence dimension, and does not perform padding in the case where a single token generated by the large model subsequently serves as the input (regarded as a padding length of 0). For the single-head inference of a single batch sample, such an operation will change the dimension of the two-dimensional matrix of the token sequence from the token vector length * actual token sequence length to the token vector length * (input padding sequence length + actual token sequence length) (reference can be made to Figure 3The upper rectangle). We ensure that the input padding sequence length + the actual token sequence length is a preset fixed value. Then, for the deep learning compiler, the modified large model has a static input shape. The operation of input padding will cause the tensor data calculated based on the input in the subsequent inference process to have as much padding data as the input padding sequence length in the sequence dimension, and this data may affect the subsequent matrix operations.

[0047] In a more microscopic attention layer, its input tensor is composed of the output of the word embedding layer or the output of the previous decoder module, and both will be affected by the padding of the input sequence. The input tensor passes through 3 different linear mapping layers (simplified and not shown in the figure) to obtain "the query at this time", "the key at this time", and "the value at this time" for attention calculation, and they will all contain a part of the tensor of the padding product (along the sequence dimension) due to input padding, while the operation results of the remaining parts are the same as those without padding. Since the position encoding operation is an elementwise operation and does not affect the contribution of the padding sequence to matrix multiplication, it will be simplified and not considered here.

[0048] "The key at this time" and "the value at this time" then need to be concatenated with the "historical key cache" and "historical value cache" along the sequence dimension respectively; when the historical cache is empty, the concatenation behavior is to assign "the key at this time" and "the value at this time" to the "historical key cache" and "historical value cache" respectively. According to such a description, performing the inference of the "computation process of generating a single token" of the large model multiple times will cause the sequence dimension of the "historical key cache" and "historical value cache" to increase during each inference, and the change in the sequence dimension results in its non-static shape. The present invention also performs an operation of padding the front side of the "historical key cache" and "historical value cache" in the sequence dimension, and at the same time performs a cut-out operation after concatenation. The cut-out length is the sequence length of the concatenated "the key at this time" and "the value at this time" to obtain the "historical key cache for matrix multiplication" and "historical value cache for matrix multiplication". Specific example: If the current sequence length of the "historical key cache" is N, and a "the key at this time" cache with a sequence length of 1 is concatenated to its back along the sequence dimension, then the sequence length of the "historical key cache after front cut-out" after concatenation is N + 1, and then a tensor with a length of 1 is cut out from the front along the sequence dimension, and the remaining tensor with a sequence length of N is retained. In this way, the sequence length of the "historical key cache" faced by subsequent operations is always N, ensuring a static shape.

[0049] In addition, the maintained cache length N mentioned in the above example can be pre-computed. Without considering the padding length of the historical cache, the length of the corresponding historical cache sequence (denoted as real_cache_len, not greater than N, including the padding length of the token sequence converted from the first user input statement) can be known before each inference. The present invention sets a slice length list, which is composed of a series of positive integers in ascending order, with the maximum value being N and N must be in the list, for further slicing the historical cache after front cutting. For each real_cache_len, the smallest value greater than or equal to real_cache_len in the list can be found and denoted as slice_size. Specific example: The slice length list is set to [32, 64, 128], and the maximum value of the sequence dimension of the historical key-value cache is N = 128. Then when the historical key-value cache real_cache_len is from 0 to 32, slice_size is 32, and the reduction dimension length of the matrix multiplication of the score and the value is 32 (i.e., the inner product of two vectors of length 32 gives an element of the output matrix). When the historical key-value cache real_cache_len is from 32 to 64, slice_size is 64, and the reduction dimension length of the matrix multiplication of the score and the value is 64. When the historical key-value cache real_cache_len is from 64 to 128, slice_size is 128, and the reduction dimension length of the matrix multiplication of the score and the value is 128.

[0050] For the historical cache after front cutting, a tensor of length slice_size is retained along the rear side of the sequence dimension. Since slice_size is greater than or equal to real_cache_len, it can ensure that the final new historical cache contains the actual real historical cache. After such processing, for both matrix multiplications, the dimension length of part of the matrix will change from N to slice_size, which helps to reduce the computational complexity of the matrix multiplication. The meaning of "after front cutting" here is "after removing the data in the forward direction on the dimension", specifically: for the historical key-value cache, along the sequence dimension (this dimension is 1D), several index values with lower indexes are removed, and the larger index values are retained. For example, if the forward removal index length is 4 and the sequence dimension length is 10, then the data in the tensors corresponding to the 4 index values 0, 1, 2, 3 on the sequence dimension are removed, and the tensor data corresponding to the sequence dimension index values 4, 5, 6, 7, 8, 9 are retained.

[0051] Next, the present invention solves the contribution of padding and padding products in the sequence dimension to matrix multiplication through special processing of the "current mask". The main body of the "current mask" is a two-dimensional matrix with dimensions: (input padding sequence length + actual token sequence length) * new historical cache sequence length. After the slicing operation of slice_size, the length of the new historical cache (similar for keys and values) sequence is: cache padding residue sequence length + historical cache sequence length + input padding sequence length + current cache sequence length (the historical cache sequence length is 0 and the input padding sequence length is not 0 during the first inference, and the historical cache sequence length is not 0 and the input padding sequence length is 0 during subsequent inferences). Here, the property of exponential normalization is used: if the value of a certain component during normalization is negative infinity (denoted as -inf), then the corresponding value of this component becomes 0 after exponential normalization. For the new historical cache sequence dimension, set the values of the corresponding dimensions of "cache padding residue sequence length" and "input padding sequence length" in the "current mask" to -inf, and the rest to 0. For the other dimension of the mask matrix, whose length is (input padding sequence length + actual token sequence length), set the value corresponding to "input padding sequence length" in the "current mask" to -inf, and the rest to 0. There will be a conflict situation when setting values in these two directions. If a value needs to be set to 0 and -inf at the same time, it is preferentially set to -inf. The mask set in this way can eliminate the influence of input padding on the summation in matrix multiplication and also eliminate the influence of historical cache padding on the summation in matrix multiplication.

[0052] Finally, following the previous steps of concatenating to obtain the "historical key cache for matrix multiplication" and the "historical value cache for matrix multiplication", we can first perform matrix multiplication between the "query at this time" with padding products and the "historical key cache for matrix multiplication" to obtain a score, and then multiply it by a scaling factor to get the scaled score. Next, add it to the specially processed mask introduced in the previous subsection to obtain the "masked score", perform exponential normalization to obtain the "exponentially normalized score", and then perform matrix multiplication with the "historical value cache for matrix multiplication" to obtain the attention score. Some matrix transpositions that do not affect the calculation are omitted (the transposed cases of the matrices are shown in the figures), as well as the final transposition and linear transformation of the attention score. Thus, the process description of the improved attention layer is completed. In the overall large model inference process, next, normalization layer operations, feed-forward activation layer operations, and normalization operations will be performed on the output of the attention layer in sequence. These four operations are called a decoder module. After performing operations of several decoder modules on the input of the word embedding layer in sequence, an output is obtained, and then normalization, fully connected layer, and post-processing sampling are performed on this output in sequence to obtain the output token. Thus, the large model inference process for a single token is completed. After all the tokens that need to be inferred are inferred, since the token sequence obtained by converting the input statement is padded, the corresponding padding removal needs to be done for the sequence composed of the obtained output tokens.

[0053] Verification process of the present invention:

[0054] The feasibility and effectiveness of the solution were verified in the chatbot inference process based on the open-source Llama model (Llama: https: / / github.com / facebookresearch / llama). Under the condition of using a deep learning graph compiler with a static input shape on a certain hardware platform, without using the present invention, each token inference requires recompilation once. If the maximum support for the chatbot context is 128 tokens, then at least 128 recompilations are required; while using the present invention, the number of compilations can be fixed at 3 times under the same parameter settings, effectively reducing the number of compilations in chatbot inference.

[0055] The above is the introduction of the method embodiment. The following further illustrates the solution of the present invention through embodiments of electronic devices and storage media.

[0056] An embodiment of the present invention also provides an electronic device including a central processing unit (CPU), which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0057] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0058] The processing unit executes the various methods and processes described above, such as the method of the present invention. For example, in some embodiments, the method of the present invention can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of the method of the present invention described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute the method of the present invention by any other suitable means (e.g., by means of firmware).

[0059] The functions described above herein can be at least partially executed by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0060] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or a controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or the controller, the functions / operations specified in the flowchart and / or the block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed as an independent software package partially on the machine and partially on a remote machine, or executed entirely on a remote machine or a server.

[0061] In the context of the present invention, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0062] As described above, only the specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for optimizing the compilation times of a chat robot, characterized in that: This method is implemented based on a deep learning compiler by modifying some of the large model reasoning processes. The large model reasoning process includes: the word embedding layer converts the input token into a corresponding word vector sequence, which is then calculated through several decoder modules, and finally the output token is obtained through the normalization layer, the fully connected layer, and the sampling post-processing in sequence; Among them, before the word embedding layer, it includes: front-end padding of "token sequence converted from user input sentence" in the sequence dimension; Among them, the decoder module includes an attention layer, which generates a "historical key cache" and a "historical value cache" according to the word vector sequence, and front-end-fills the "historical key cache" and the "historical value cache" in the sequence dimension; the scores obtained based on the query and the filled historical key cache are masked to obtain masked scores, and the masked scores are exponentially normalized to obtain exponentially normalized scores.

2. A method for optimizing the compilation times of a chat robot according to claim 1, characterized in that: Only the first "token sequence converted from the user input sentence" is front-padded in the sequence dimension, and in the subsequent case where a single token generated by the large model is used as input, the padded length is regarded as 0.

3. A method for optimizing the compilation times of a chat robot according to claim 1, characterized in that: The input padding sequence length + the actual token sequence length is a preset fixed value.

4. The method for optimizing the compilation times of a chat robot according to claim 1, characterized in that: The "historical key cache" and "historical value cache" are first spliced ​​with the "current key" and "current value" respectively, and then the front side is padded in the sequence dimension, and the cut-out operation after splicing is performed at the same time; the cut-out length is the sequence length of the corresponding spliced ​​"current key" and "current value", and after cutting out, the "historical key cache for matrix multiplication" and "historical value cache for matrix multiplication" are obtained respectively.

5. A method for optimizing the compilation times of a chat robot according to claim 4, characterized in that: The cache length N of the "historical key cache" and "historical value cache" is pre-calculated, without considering the padded length of the historical cache, and the remaining corresponding historical cache sequence length is recorded as real_cache_len, which is not greater than N. The method sets a slice length list, which is composed of a series of positive integers from small to large, with a maximum value of N and N in the list, for further slicing the historical cache that was previously cut out; for each real_cache_len, find a value greater than or equal to real_cache_len and the smallest value in the list and record it as slice_size; for the historical cache that was previously cut out, retain a tensor of slice_size length along the back side of the sequence dimension.

6. A method for optimizing the compilation times of a chat robot according to claim 5, characterized in that: The mask data processing is performed before normalizing the index as follows: The mask data is a two-dimensional matrix with the dimensions of (input padding sequence length + actual token sequence length) * new history cache sequence length, where the new history cache sequence length is: cache padding residual sequence length + history cache sequence length + input padding sequence length + current cache sequence length; For the new historical cache sequence dimension, set the values ​​of the corresponding dimensions of "cache padding residual sequence length" and "input padding sequence length" in the mask data to -inf, and the rest to 0, where -inf is negative infinity; For the other dimension of the mask data matrix, its length is (input padding sequence length + actual token sequence length), the corresponding "input padding sequence length" in the mask data is -inf, and the rest are 0; If a value is set to either 0 or -inf, -inf takes precedence.

7. A method for optimizing the compilation times of a chat robot according to claim 6, characterized in that: The processed masked data is summed to obtain the "masked score", which is then normalized by exponential normalization to obtain the "exponential normalized score", which is then matrix multiplied with the "historical value cache for matrix multiplication" to obtain the attention score.

8. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Cache data processing method, electronic device and readable storage medium

    CN111400308A

  • Learned threshold token pruning for transformer neural networks

    US20220374766A1