Memory Management Method and Device for Large Language Model
By optimizing the memory management method of the large language model, including setting the target memory management block and dynamically adjusting the key-value cache, the memory fragmentation problem caused by frequent discarding is solved, and the stability and efficiency of memory management are improved.
Patent Information
- Application Number
- CN202510111389.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Large language models have low stability in memory management, mainly due to frequent discarding of key-value caches, resulting in excessive memory fragmentation.
By obtaining input vocabulary and performing inference processing, setting the target memory management block storage key-value cache, storing and canceling storage operations according to the comparison of key-value cache lengths, optimizing memory management strategies, including splitting input vocabulary for sequential reasoning and dynamically adjusting key-value caches.
It improves the memory management stability of large language models, reduces the frequency of key-value cache discarding, and improves the efficiency of memory management and the overall performance of the model.
Smart Images

Figure CN119576805B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computers, and in particular, to a method and device for memory management of large language models. Background Art
[0002] In the related art, during the process of large language processing and inference, in order to improve the memory management efficiency, it is necessary to continuously discard or selectively save some generated tokens, and it is also necessary to discard the key-value cache frequently. Frequent discarding of the key-value cache will generate a large amount of video memory fragmentation, which will in turn lead to the problem of low stability of the memory management of large language models. Therefore, there is a problem of low stability of the memory management of large language models.
[0003] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] The embodiments of the present application provide a method and device for memory management of large language models to at least solve the problem of low stability of the memory management of large language models in the related art.
[0005] According to an embodiment of the present application, a method for memory management of a large language model is provided, including: obtaining an input token, where the input token is a basic unit processed by the large language model; performing inference processing on the input token through the large language model to obtain a key-value cache of a first length, where the key-value cache of the first length is the key-value cache corresponding to the tokens to be deleted of the input token; in the case where the first length is less than a second length, setting a target memory management block to store the key-value cache of the first length, where the second length is the maximum length allowed for storing the key-value cache by the target memory management block; in the case where the first length is equal to the second length, setting the target memory management block to cancel storing the key-value cache of the first length.
[0006] According to another embodiment of the present application, a device for memory management of a large language model is provided, including: a first obtaining unit for obtaining an input token, where the input token is a basic unit processed by the large language model; an inference unit for performing inference processing on the input token through the large language model to obtain a key-value cache of a first length, where the key-value cache of the first length is the key-value cache corresponding to the tokens to be deleted of the input token; a storage unit for setting a target memory management block to store the key-value cache of the first length in the case where the first length is less than a second length, where the second length is the maximum length allowed for storing the key-value cache by the target memory management block; a setting unit for setting the target memory management block to cancel storing the key-value cache of the first length in the case where the first length is equal to the second length.
[0007] As an alternative solution, the above-mentioned inference unit includes: a first inference module, configured to perform inference processing on the above-mentioned input token during the pre-filling stage through the above-mentioned large language model.
[0008] As an alternative solution, the above-mentioned inference module includes: a splitting sub-module, configured to split the above-mentioned input token when the number of tokens of the above-mentioned input token is greater than the inference token upper limit, and perform sequential inference processing on the split input tokens, where the split input tokens include at least two segments of input tokens, and the number of tokens of each segment of the at least two segments of input tokens is less than or equal to the above-mentioned inference token upper limit.
[0009] As an alternative solution, the above-mentioned splitting sub-module includes: a setting sub-unit, configured to set the input word of the first segment among the above-mentioned n segments of input tokens to the first number of tokens, set the number of tokens of the input word of the i-th segment among the above-mentioned n segments of input tokens to the second number of tokens, and set the number of tokens of the input word of the n-th segment among the above-mentioned n segments of input tokens to the third number of tokens, where i is an integer greater than 1 and less than or equal to n, the first number of tokens is an integer multiple of the quantity corresponding to the second length, the first number of tokens is greater than the sum of the initial number of tokens and the rolling number of tokens, the initial number of tokens and the rolling number of tokens are used to determine the key-value cache to be deleted, the initial number of tokens is an integer multiple of the quantity corresponding to the second length, the second number of tokens, the initial number of tokens, and the rolling number of tokens are all integer multiples of the quantity corresponding to the second length, and the sum of the second number of tokens and the initial number of tokens and the rolling number of tokens is less than or equal to the above-mentioned inference token upper limit.
[0010] As an alternative solution, the above-mentioned setting sub-unit includes: a first inference component, configured to perform sequential inference processing on the above-mentioned split input tokens in a first inference manner when n is greater than 2; a second inference component, configured to perform sequential inference processing on the above-mentioned split input tokens in a second inference manner when n is equal to 2, where the second inference manner is different from the first inference manner.
[0011] As an alternative solution, the above-mentioned first inference component includes: a first inference sub-component, configured to perform inference processing on f input words of the first segment among the above-mentioned n segments of input tokens to obtain f key-value caches corresponding to the f input words; a second inference sub-component, configured to combine the above-mentioned f key-value caches and perform inference processing on s input words of the i-th segment among the above-mentioned n segments of input tokens to obtain s key-value caches corresponding to the s input words, where i is an integer greater than 1 and less than n; a third inference sub-component, configured to combine the above-mentioned s key-value caches and perform inference processing on m input words of the n-th segment among the above-mentioned n segments of input tokens to obtain m key-value caches.
[0012] As an alternative solution, the above-mentioned second inference sub-component includes: a first determination component, configured to determine, from the above-mentioned f key-value caches, the key-value caches with the first target quantity that is the number of the above-mentioned initial tokens and the last above-mentioned scrolling tokens; a first inference component, configured to perform inference processing on the above-mentioned s input words by combining the above-mentioned first target quantity of key-value caches to obtain the above-mentioned s key-value caches.
[0013] As an alternative solution, the above-mentioned third inference sub-component includes: a second determination component, configured to determine, from the above-mentioned first target quantity of key-value caches and the above-mentioned s key-value caches, the key-value caches with the second target quantity that is the number of the above-mentioned initial tokens and the last above-mentioned scrolling tokens; a second inference component, configured to perform inference processing on the above-mentioned m input words by combining the above-mentioned second target quantity of key-value caches to obtain the above-mentioned m key-value caches.
[0014] As an alternative solution, the above-mentioned second inference component includes: a fourth inference sub-component, configured to perform inference processing on the f input words in the first segment of the above-mentioned n segments of input tokens to obtain the f key-value caches corresponding to the above-mentioned f input words; a fifth inference sub-component, configured to perform inference processing on the m input words in the nth segment of the above-mentioned n segments of input tokens by combining the above-mentioned f key-value caches to obtain m key-value caches.
[0015] As an alternative solution, the above-mentioned fifth inference sub-component includes: a third determination component, configured to determine, from the above-mentioned f key-value caches, the key-value caches with the second target quantity, where the key-value caches with the second target quantity are the key-value caches after the above-mentioned f key-value caches discard the key-value caches with the third target quantity, and the above-mentioned target quantity is the smallest integer obtained by rounding up the quotient of m and the quantity corresponding to the above-mentioned second length; a third inference component, configured to perform inference processing on the above-mentioned m input words by combining the above-mentioned second target quantity of key-value caches to obtain the above-mentioned m key-value caches.
[0016] As an alternative solution, the above-mentioned inference module includes: a decoding sub-module, configured to perform inference processing in the decoding stage on the key-value caches obtained in the above-mentioned pre-filling stage through the above-mentioned large language model.
[0017] As an alternative solution, the above-mentioned inference sub-module includes: a deletion sub-unit, configured to, when the quantity of the key-value caches obtained in the above-mentioned pre-filling stage is greater than the key-value cache upper limit, perform deletion processing on the key-value caches obtained in the above-mentioned pre-filling stage to obtain the deleted key-value caches, where the quantity of the above-mentioned deleted key-value caches is less than or equal to the above-mentioned key-value cache upper limit.
[0018] As an alternative solution, the above-mentioned deletion subunit includes: a deletion component, configured to delete the key-value cache from the initial token number to the rolling token number in the key-value cache obtained in the above-mentioned pre-filling stage, so as to obtain the above-mentioned deleted key-value cache, where the above-mentioned initial token number and the above-mentioned rolling token number are used to determine the key-value cache to be deleted, and both the above-mentioned initial token number and the above-mentioned rolling token number are integer multiples of the quantity corresponding to the above-mentioned second length.
[0019] As an alternative solution, the above-mentioned deletion subunit includes: a first output component, configured to output a current new token and a current new key-value cache through the above-mentioned large language model, in combination with the above-mentioned deleted key-value cache and the tokens obtained in the above-mentioned pre-filling stage; an acquisition component, configured to acquire the above-mentioned output tokens when the above-mentioned current new token meets the termination condition, where the above-mentioned output tokens include the tokens obtained in the above-mentioned pre-filling stage and the new tokens obtained in the above-mentioned pre-filling stage; a second output component, configured to, when the above-mentioned current new token does not meet the above-mentioned termination condition, output a next new token and a next new key-value cache through the above-mentioned large language model, in combination with the above-mentioned current new token and the above-mentioned current new key-value cache, and use the above-mentioned next new token as the above-mentioned current new token and the next new key-value cache as the above-mentioned current new key-value cache.
[0020] As an alternative solution, the above-mentioned first acquisition unit includes: a first acquisition module, configured to acquire an input token corresponding to a single token processing request; a second acquisition module, configured to acquire a first input token corresponding to a first token processing request and a second input token corresponding to a second token processing request; a combination module, configured to, when the token quantity of the above-mentioned first input token is less than the token processing upper limit, perform a combination process on the above-mentioned first input token and some tokens of the above-mentioned second input token through the above-mentioned large language model, where the sum of the token quantity of the above-mentioned first input token and the token quantity of some tokens of the above-mentioned second input token is less than the above-mentioned token processing upper limit.
[0021] As an alternative solution, the above-mentioned combination module includes: a combination sub-module, configured to, when the token quantity of the above-mentioned first input token is less than the above-mentioned second token processing upper limit, perform a combination process on the above-mentioned first input token and some tokens of the above-mentioned second input token through the above-mentioned large language model.
[0022] As an alternative solution, the above-mentioned inference unit includes: a pause module, configured to pause the inference process on the above-mentioned key-value cache when the remaining space corresponding to the memory management space of the above-mentioned large language model is less than a preset threshold; a continue module, configured to continue the inference process on the key-value cache whose inference process has been paused when the remaining space corresponding to the memory management space of the above-mentioned large language model is greater than or equal to the preset threshold.
[0023] As an alternative solution, the above device further includes: a second acquisition unit configured to acquire the model parameters and data precision of the above large language model; a third acquisition unit configured to acquire the memory space required for each of the above key-value caches according to the above model parameters and the above data precision; a fourth acquisition unit configured to utilize the memory space required for each of the above key-value caches and the memory occupied by the above large language model according to a preset weight to acquire the maximum length allowed for storing key-value caches in the above target memory management block.
[0024] According to another embodiment of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0025] According to another embodiment of the present application, there is also provided an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0026] Through the present application, an input token is acquired, where the input token is a basic unit processed by a large language model; through the large language model, the input token is inferentially processed to obtain a key-value cache of a first length, where the key-value cache of the first length is the key-value cache corresponding to the tokens to be deleted from the input token; in the case where the first length is less than a second length, the target memory management block is set to store the key-value cache of the first length, where the second length is the maximum length allowed for storing key-value caches in the target memory management block; in the case where the first length is equal to the second length, the target memory management block is set to cancel storing the key-value cache of the first length.
[0027] Specifically, after the large language model inferentially processes the input token to obtain a key-value cache of a first length, in the case where the first length is less than the second length, the key-value cache of the first length is stored, and in the case where the number of key-value caches to be deleted in the key-value cache of the first length is equal to the quantity corresponding to the second length, storing of the first quantity of key-value caches is cancelled, so that only when the key-value cache of the first length is equal to the second length, the first quantity of key-value caches is cancelled, reducing the frequency of discarding key-value caches and improving the stability of the memory management of the large language model. Therefore, the problem of low stability of the memory management of the large language model can be solved, and thus the technical effect of improving the stability of the memory management of the large language model is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a schematic diagram of the application environment of the memory management method of the large language model according to the embodiment of the present application;
[0029] Figure 2 is a flowchart of a memory management method for a large language model according to an embodiment of the present application;
[0030] Figure 3 is a schematic diagram of a memory management method for a large language model according to an embodiment of the present application;
[0031] Figure 4 is a schematic diagram of a memory management method for a large language model according to an embodiment of the present application;
[0032] Figure 5 is a schematic diagram of a memory management method for a large language model according to an embodiment of the present application;
[0033] Figure 6 is a schematic diagram of a memory management method for a large language model according to an embodiment of the present application;
[0034] Figure 7 is a structural block diagram of a memory management device for a large language model according to an embodiment of the present application. Detailed implementation manners
[0035] In the following, embodiments of the present application will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.
[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0037] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 is a hardware structural block diagram of a server device for a memory management method of a large language model according to an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in the figure is only schematic, and it does not limit the structure of the above-mentioned server device. For example, the server device may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0038] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the memory management method of the large language model in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0039] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0040] In this embodiment, a memory management method for a large language model is provided. Figure 2 It is a flowchart of the memory management method for the large language model according to the embodiments of the present application, as Figure 2 shown, and the process includes the following steps:
[0041] Step S202, obtain input tokens, where the input tokens are the basic units processed by the large language model;
[0042] In an exemplary embodiment, the input tokens can be but are not limited to the smallest units extracted from the input text, can be but are not limited to one or more characters, words, sub-word units, or symbols, used for the input of the model, can be but are not limited to obtained from the input sequence input into the large language model, and the input sequence can be but are not limited to the sequence input by the user into the large language model for inference.
[0043] In an exemplary embodiment, the large language model can be but is not limited to a deep learning model that can understand and generate complex natural language.
[0044] Step S204, use a large language model to perform inference processing on the input token to obtain a key-value cache of a first length, where the key-value cache of the first length is the key-value cache corresponding to the tokens to be deleted from the input token.
[0045] In an exemplary embodiment, the inference processing can, but is not limited to, be understood as a process of allowing the large language model to generate an output based on the input token, and can, but is not limited to, include the understanding of the input and the prediction of the output.
[0046] In an exemplary embodiment, the key-value cache can, but is not limited to, be a cache used to store the generated keys and values during the inference of the large language model to avoid repeated calculations when generating subsequent tokens. The first quantity of key-value caches can, but is not limited to, refer to the number of key-value caches generated and required to be stored in the inference step.
[0047] It should be noted that the key-value cache of the first length can, but is not limited to, be understood as a certain number of key-value caches. For further illustration, assuming the length of the key-value cache is 8, it can be understood as 8 key-value caches.
[0048] In addition, when processing ultra-long sequences, to avoid excessive memory consumption, it is necessary to periodically discard some key-value caches that are no longer needed. The number of key-value caches to be deleted can, but is not limited to, refer to the number of key-value caches that need to be removed to maintain the efficient use of memory in the current inference stage or memory management strategy.
[0049] In an exemplary embodiment, the memory management block can, but is not limited to, be a spatial unit used to store a specific number of key-value caches, and can, but is not limited to, allow storing up to the number corresponding to the second length of key-value caches.
[0050] Step S206, in the case where the first length is less than the second length, set the target memory management block to store the key-value cache of the first length, where the second length is the maximum length that the target memory management block allows to store key-value caches.
[0051] In an exemplary embodiment, the target memory management block can, but is not limited to, be understood as a single memory management block designated to store the key-value cache of the first length.
[0052] Step S208, in the case where the first length is equal to the second length, set the target memory management block to cancel storing the key-value cache of the first length.
[0053] It should be noted that when the total amount of key-value caches is equal to the capacity of a single memory management block, i.e., the second length, it indicates that a single memory management block is full. At this time, all key-value caches need to be discarded to empty the memory management block. The operation when the storage space reaches the upper limit and space needs to be released is specified, ensuring the continuity of the process and the efficiency of memory management.
[0054] Through the above steps, input tokens are obtained. Among them, the input tokens are the basic units processed by the large language model. Through the large language model, the input tokens are inferentially processed to obtain key-value caches of the first length. Among them, the key-value caches of the first length are the key-value caches corresponding to the tokens to be deleted from the input tokens. When the first length is less than the second length, the target memory management block is set to store the key-value caches of the first length. Among them, the second length is the maximum length that the target memory management block allows to store key-value caches. When the first length is equal to the second length, the target memory management block is set to cancel storing the key-value caches of the first length.
[0055] Specifically, after the large language model inferentially processes the input tokens to obtain key-value caches of the first length, when the first length is less than the second length, the key-value caches of the first length are stored. When the number of key-value caches to be deleted in the key-value caches of the first length is equal to the quantity corresponding to the second length, the storage of the first quantity of key-value caches is cancelled. Thus, only when the key-value caches of the first length are equal to the second length, the first quantity of key-value caches is cancelled, reducing the frequency of discarding key-value caches and improving the stability of the memory management of the large language model. Therefore, the problem of low stability of the memory management of the large language model can be solved, and the technical effect of improving the stability of the memory management of the large language model is achieved.
[0056] Among them, the execution subject of the above steps can be a server, a terminal, etc., but is not limited thereto.
[0057] As an optional solution, inferentially processing the input tokens through the large language model includes:
[0058] Through the large language model, inferentially process the input tokens in the prefill stage.
[0059] In an optional embodiment, the prefill stage can but is not limited to be understood as an initial stage of the large language model inference process, and can but is not limited to include the process of the large language model receiving and processing input tokens and inferentially generating a series of tokens.
[0060] Through the embodiments of the present application, inferentially process the input tokens in the prefill stage through the large language model. The technical effect of inferentially processing the input tokens in the prefill stage and generating tokens is achieved.
[0061] As an alternative solution, during the process of performing inference processing on input tokens in the pre-filling stage through a large language model, the method further includes:
[0062] In the case where the number of tokens of the input tokens is greater than the inference token limit, the input tokens are split, and the split input tokens are sequentially processed for inference, where the split input tokens include at least two segments of input tokens, and the number of tokens of each segment of the at least two segments of input tokens is less than or equal to the inference token limit.
[0063] In an alternative embodiment, the inference token limit may but is not limited to refer to the maximum number of tokens that the model or hardware can process at one time during the inference process of the large language model.
[0064] In an alternative embodiment, the split input tokens can be but are not limited to being understood as when the number of tokens in the input sequence exceeds the inference token limit, the original input sequence is segmented to generate multiple subsequences, and the number of tokens in each subsequence does not exceed the inference token limit.
[0065] It should be noted that the inference token limit is usually determined by limitations in computing resources, memory size, or model architecture. If this limit is exceeded, the inference performance of the large language model will be affected, or it may cause a memory overflow. When the total number of input tokens exceeds the upper limit that the model can process at one time, the input token sequence must be split into smaller parts, and then inference processing is performed one by one in the order of these smaller parts. The splitting and sequential inference strategies can significantly improve the practicality and efficiency of the large language model in processing long texts or high-concurrency scenarios, while ensuring the accuracy and integrity of model inference.
[0066] Through the embodiments of the present application, in the case where the number of tokens of the input tokens is greater than the inference token limit, the input tokens are split, and the split input tokens are sequentially processed for inference, where the split input tokens include at least two segments of input tokens, and the number of tokens of each segment of the at least two segments of input tokens is less than or equal to the inference token limit. The technical purpose of splitting the input token sequence into smaller parts and then performing inference processing one by one in the order of these smaller parts is achieved, and thus the technical effect of ensuring the accuracy and integrity of model inference is realized.
[0067] As an alternative solution, the split input tokens include n segments of input tokens, where n is an integer greater than 1. Splitting the input tokens includes:
[0068] Set the input word of the first segment among the n segments of input tokens as the first token number, set the number of input tokens of the i-th segment among the n segments of input tokens as the second token number, and set the number of input tokens of the n-th segment among the n segments of input tokens as the third token number, where i is an integer greater than 1 and less than or equal to n. The first token number is an integer multiple of the quantity corresponding to the second length, and the first token number is greater than the sum of the initial token number and the rolling token number. The initial token number and the rolling token number are used to determine the key-value cache to be deleted. The initial token number is an integer multiple of the quantity corresponding to the second length. The second token number, the initial token number, and the rolling token number are all integer multiples of the quantity corresponding to the second length. The sum of the second token number and the initial token number and the rolling token number is less than or equal to the inference token upper limit.
[0069] In an alternative embodiment, the first token number may, but is not limited to, refer to the number of tokens of the first segment of input tokens among the n segments of input tokens, and the first token number may, but is not limited to, be an integer multiple of the quantity corresponding to the second length.
[0070] In an alternative embodiment, the second token number may, but is not limited to, refer to the number of tokens of each segment of input tokens except the first segment and the n-th segment after splitting the input token sequence, and may, but is not limited to, be an integer multiple of the quantity corresponding to the second length.
[0071] In an alternative embodiment, the third token number may, but is not limited to, refer to the number of tokens of the last segment of input tokens after splitting the input token sequence.
[0072] In an alternative embodiment, the initial token number may, but is not limited to, refer to a part of the number of tokens retained from the beginning of the sequence to ensure the inference accuracy of the model when processing the input sequence. These tokens may, but are not limited to, be used as assistance information in subsequent inference processes to assist the large language model in understanding the text information of the input sequence.
[0073] In an alternative embodiment, the rolling token number may, but is not limited to, refer to the number of the most recent part of tokens retained after each inference to maintain the coherence of model inference when processing the input sequence. These tokens may, but are not limited to, be used as the latest input in subsequent inference to ensure the coherence of the model's understanding and generation of the input sequence.
[0074] It should be noted that reasonably splitting the input token sequence into multiple segments and setting the number of tokens for each segment to adapt to the processing capacity and memory limit of the large language model can ensure the accuracy of large language model inference and reduce processing latency when processing long texts or high-concurrency requests.
[0075] Through the embodiments of the present application, the input word of the first segment among the n segments of input tokens is set as the first number of tokens, the number of input tokens of the i-th segment among the n segments of input tokens is set as the second number of tokens, and the number of input tokens of the n-th segment among the n segments of input tokens is set as the third number of tokens, where i is an integer greater than 1 and less than or equal to n. The first number of tokens is an integer multiple of the quantity corresponding to the second length, and the first number of tokens is greater than the sum of the initial number of tokens and the rolling number of tokens. The initial number of tokens and the rolling number of tokens are used to determine the key-value cache to be deleted. The initial number of tokens is an integer multiple of the quantity corresponding to the second length. The second number of tokens, the initial number of tokens, and the rolling number of tokens are all integer multiples of the quantity corresponding to the second length. The sum of the second number of tokens and the initial number of tokens and the rolling number of tokens is less than or equal to the inference token upper limit. The technical purpose of reasonably splitting the input token sequence into multiple segments and setting the number of tokens in each segment is achieved, and further the technical effect of ensuring the accuracy of the large language model inference is realized.
[0076] The split input tokens include n segments of input tokens, where n is an integer greater than 1. The split input tokens are sequentially inference-processed, including:
[0077] S1-1, when n is greater than 2, the first inference method is adopted to sequentially inference-process the split input tokens;
[0078] S1-2, when n is equal to 2, the second inference method is adopted to sequentially inference-process the split input tokens, where the second inference method is different from the first inference method.
[0079] In an alternative embodiment, the first inference method may but is not limited to refer to the inference processing strategy adopted by the large language model when the input token sequence is split into more than two segments.
[0080] In an alternative embodiment, the second inference method may but is not limited to refer to the inference processing strategy adopted by the large language model when the input token sequence is only split into two segments.
[0081] It should be noted that by selecting different inference methods according to the split number of input tokens, it is possible to flexibly adapt to the inference processing schemes in different scenarios, ensure that the large language model can efficiently and accurately perform inference processing regardless of how many segments the sequence is split into when processing an ultra-long input sequence, optimize memory usage, avoid resource waste, and improve the inference performance and response speed of the large language model.
[0082] Through the embodiments of the present application, when n is greater than 2, the first reasoning method is adopted to sequentially reason about the split input tokens; when n is equal to 2, the second reasoning method is adopted to sequentially reason about the split input tokens, where the second reasoning method is different from the first reasoning method. The purpose of being able to perform reasoning processing efficiently and accurately regardless of how many segments the sequence is split into is achieved, and thus the technical effect of improving the reasoning performance and response speed of the large language model is realized.
[0083] As an alternative solution, adopting the first reasoning method to sequentially reason about the split input tokens includes:
[0084] S2-1, reason about the f input tokens in the first segment of the n segments of input tokens to obtain f key-value caches corresponding to the f input tokens;
[0085] S2-2, in combination with the f key-value caches, reason about the s input tokens in the i-th segment of the n segments of input tokens to obtain s key-value caches corresponding to the s input tokens, where i is an integer greater than 1 and less than n;
[0086] S2-3, in combination with the s key-value caches, reason about the m input tokens in the n-th segment of the n segments of input tokens to obtain m key-value caches.
[0087] It should be noted that by using the key-value caches generated by the reasoning of the first segment to reason about the input tokens of the i-th segment and obtain new S key-value caches, the continuous understanding and reasoning of the input sequence by the large language model are ensured, so that the reasoning results of the intermediate segments not only depend on the input tokens of this segment, but also depend on the key-value caches generated by the previous reasoning, thus maintaining the coherence of the reasoning process.
[0088] Furthermore, by further using the s key-value caches generated by the reasoning to perform reasoning processing to generate m key-value caches, the reasoning of the input tokens is completed, providing complete information for the subsequent decoding stage.
[0089] Through the embodiments of the present application, reason about the f input tokens in the first segment of the n segments of input tokens to obtain f key-value caches corresponding to the f input tokens; in combination with the f key-value caches, reason about the s input tokens in the i-th segment of the n segments of input tokens to obtain s key-value caches corresponding to the s input tokens, where i is an integer greater than 1 and less than n; in combination with the s key-value caches, reason about the m input tokens in the n-th segment of the n segments of input tokens to obtain m key-value caches. The technical purpose of using the key-value caches generated by the reasoning of the first segment to reason about the input tokens of the i-th segment and obtain new S key-value caches is achieved, and thus the coherence of the reasoning of the input sequence by the large language model is realized.
[0090] As an alternative solution, in combination with f key-value caches, s input words in the i-th segment of the n segments of input tokens are processed for inference to obtain s key-value caches, including:
[0091] S3-1: Determine, from the f key-value caches, key-value caches with a first target quantity of the number of initial tokens and the number of rolling tokens at the end;
[0092] S3-2: In combination with the key-value caches with the first target quantity, process the s input words for inference to obtain s key-value caches.
[0093] It should be noted that after processing the f input words in the first segment and generating f key-value caches, screening out a specific quantity of key-value caches from these key-value caches, that is, the key-value caches corresponding to the number of initial tokens and the number of rolling tokens, as the basis for subsequent inference, can ensure the coherence and inference accuracy in the subsequent inference process, while avoiding unnecessary memory occupation and improving resource utilization.
[0094] Furthermore, during the inference process of the i-th segment, by using the previously determined key-value caches with the first target quantity, the s input words in the current segment are processed for inference to generate new s key-value caches for subsequent inference, ensuring the coherence and efficiency of the model during the inference process. By using the key cache information in the previous inference results, the large language model can more accurately understand and generate tokens, avoiding repeated calculations and resource waste.
[0095] Through the embodiments of the present application, key-value caches with a first target quantity of the number of initial tokens and the number of rolling tokens at the end are determined from the f key-value caches; in combination with the key-value caches with the first target quantity, the s input words are processed for inference to obtain s key-value caches. It achieves the technical purpose of using the key-value caches corresponding to the number of initial tokens and the number of rolling tokens as the basis for subsequent inference, and realizes the technical effect of avoiding resource waste.
[0096] As an alternative solution, in combination with s key-value caches, m input words in the n-th segment of the n segments of input tokens are processed for inference to obtain m key-value caches, including:
[0097] S4-1: Determine, from the key-value caches with the first target quantity and the s key-value caches, key-value caches with a second target quantity of the number of initial tokens and the number of rolling tokens at the end;
[0098] S4-2: In combination with the key-value caches with the second target quantity, process the m input words for inference to obtain m key-value caches.
[0099] It should be noted that after processing the s input words in the middle segment and generating s key-value caches, a specific number of caches need to be selected from these s caches and the reserved first target number of key-value caches. This part of the caches contains key-value pairs corresponding to the initial token number and the rolling token number, which are used for the inference processing of the last segment, ensuring that the model can utilize the most relevant information during the inference processing of the last segment. At the same time, it reduces unnecessary memory occupation, improves resource utilization rate and inference efficiency.
[0100] Furthermore, in the inference processing of the last segment, the second target number of key-value caches selected is combined with the m input words of the current segment for inference processing to generate m key-value caches, completing the inference of the entire input sequence, ensuring the coherence and accuracy of the inference process of the entire input sequence. Especially when processing ultra-long sequence inputs, by effectively managing and utilizing key-value caches, the model can more accurately understand and generate tokens, and is more efficient in memory management, avoiding repeated calculations and resource waste, and improving the inference efficiency of the large language model.
[0101] Through the embodiments of the present application, the second target number of key-value caches of the first initial token number and the last rolling token number are determined from the first target number of key-value caches and s key-value caches; combined with the second target number of key-value caches, the m input words are subjected to inference processing to obtain m key-value caches. It achieves the technical purpose of being able to more accurately understand and generate tokens by effectively managing and utilizing key-value caches when processing ultra-long sequence inputs, and realizes the technical effect of improving the inference efficiency of the large language model.
[0102] As an alternative solution, the second inference method is adopted to perform sequential inference processing on the split input tokens, including:
[0103] S5-1, perform inference processing on the f input words in the first segment of the n segments of input tokens to obtain f key-value caches corresponding to the f input words;
[0104] S5-2, combine the f key-value caches and perform inference processing on the m input words in the nth segment of the n segments of input tokens to obtain m key-value caches.
[0105] It should be noted that in the pre-fill stage, inference processing is performed on the first segment of the input sequence to generate key-value caches equal in number to the input words for subsequent inference use, which can generate a part of the key-value caches in advance, reduce the computational overhead in the subsequent decoding stage, and improve the response speed and processing efficiency of the model.
[0106] Furthermore, when processing the last segment of input tokens, the existing f key-value caches are used as context information to perform inference processing on this part of the input tokens, generating corresponding m key-value caches, completing the inference process of the entire input sequence, ensuring the coherence and context integrity of the last inference process, avoiding repeated calculations, improving the utilization rate of memory resources, and thus enhancing the performance and response speed of the large language model when processing long sequence inputs.
[0107] Through the embodiments of this application, inference processing is performed on the f input tokens in the first segment of the n segments of input tokens to obtain f key-value caches corresponding to the f input tokens; in combination with the f key-value caches, inference processing is performed on the m input tokens in the nth segment of the n segments of input tokens to obtain m key-value caches. The technical purpose of using the existing f key-value caches as context information to perform inference processing on this part of the input tokens and generating corresponding m key-value caches when processing the last segment of input tokens is achieved, and thus the technical effect of improving the utilization rate of memory resources is realized.
[0108] As an optional solution, performing inference processing on the m input tokens in the nth segment of the n segments of input tokens in combination with the f key-value caches to obtain m key-value caches includes:
[0109] S6-1, determining a second target number of key-value caches from the f key-value caches, where the second target number of key-value caches is the key-value caches after discarding a third target number of key-value caches from the f key-value caches, and the target number is the smallest integer obtained by rounding up the quotient of m and the number corresponding to the second length;
[0110] S6-2, performing inference processing on the m input tokens in combination with the second target number of key-value caches to obtain m key-value caches.
[0111] In an optional embodiment, the second target number of key-value caches is not limited to being understood as a set of key-value caches screened from the f key-value caches for accelerating inference when processing the last segment of input tokens.
[0112] In an optional embodiment, the third target number of key-value caches: The third target number of key-value caches is the number of key-value caches that need to be discarded from the f key-value caches when determining the second target number of key-value caches.
[0113] It should be noted that when processing the last input token, among the f key-value caches generated during the pre-filling stage, according to a specific calculation rule, that is, the smallest integer obtained by rounding up the quotient of m and the quantity corresponding to the second length, the third target quantity of key-value caches to be excluded is determined. The remaining key-value caches, that is, the key-value caches of the second target quantity, will be used for the inference processing of the last segment, ensuring that when processing the last input token, the model can accurately utilize the most relevant and necessary context information for inference, while avoiding excessive memory usage overhead, improving resource utilization efficiency and inference speed.
[0114] Furthermore, using the key-value caches of the second target quantity selected from the f key-value caches, the m input words in the nth input sequence are subjected to inference processing to generate m key-value caches, completing the processing of the entire input sequence. This ensures the accuracy and coherence of the model inference. Even when processing the last input token, it can make full use of the key-value cache information, reduce repeated calculations, accelerate the inference process, and improve the inference performance of the large language model when processing long sequence inputs.
[0115] Through the embodiments of the present application, the key-value caches of the second target quantity are determined from the f key-value caches, where the key-value caches of the second target quantity are the key-value caches after discarding the third target quantity of key-value caches from the f key-value caches, and the target quantity is the smallest integer obtained by rounding up the quotient of m and the quantity corresponding to the second length; combining the key-value caches of the second target quantity, the m input words are subjected to inference processing to obtain m key-value caches. It achieves the technical purpose of being able to make full use of the key-value cache information, reduce repeated calculations, and accelerate the inference process when processing the last input token, thereby realizing the improvement of the inference performance of the large language model when processing long sequence inputs.
[0116] As an optional solution, after the inference processing of the input tokens in the pre-filling stage by the large language model, the method further includes:
[0117] Through the large language model, the key-value caches obtained in the pre-filling stage are subjected to inference processing in the decoding stage.
[0118] In an optional embodiment, the decoding stage can be but is not limited to understood as the stage where after generating the initial tokens in the pre-filling stage, new tokens are gradually generated based on the existing key-value caches until a specific end condition is met.
[0119] It should be noted that the goal of the decoding stage is to generate new tokens or complete the generation of the sequence based on the existing key-value cache. Through the inference process in the decoding stage, the model can quickly generate new tokens based on the key-value cache generated in the prefill stage. This not only accelerates the generation process but also improves the model's response speed and processing efficiency, significantly enhancing the user experience and the performance of the large language model when dealing with real-time and streaming data processing requirements. Additionally, during the decoding process, by leveraging the key-value cache generated in the prefill stage, it is also possible to avoid redundant calculations, reduce memory occupancy, improve real-time response capabilities and overall processing speed, and enhance the inference performance of the large language model.
[0120] Through the embodiments of the present application, the large language model performs inference processing in the decoding stage on the key-value cache obtained in the prefill stage, achieving the technical goal of quickly generating new tokens based on the key-value cache generated in the prefill stage, and thus realizing the technical effect of enhancing the inference performance of the large language model.
[0121] As an alternative solution, performing inference processing in the decoding stage on the key-value cache obtained in the prefill stage through the large language model includes:
[0122] In the case where the number of key-value caches obtained in the prefill stage is greater than the key-value cache upper limit, the key-value caches obtained in the prefill stage are deleted to obtain the deleted key-value caches, where the number of the deleted key-value caches is less than or equal to the key-value cache upper limit.
[0123] It should be noted that when the number of key-value caches generated in the prefill stage exceeds the key-value cache upper limit set by the large language model and a process of deleting the cache is required, the most relevant key-value cache information should be retained at this time, while discarding the non-critical parts to ensure that the memory resource limit is not exceeded, effectively avoiding memory overflow, reducing resource waste, and ensuring that when processing ultra-long sequences or high-concurrency requests, the model can automatically adjust the storage of the key-value cache to avoid a decline in the inference performance of the large language model due to excessive consumption of memory resources.
[0124] Through the embodiments of the present application, in the case where the number of key-value caches obtained in the prefill stage is greater than the key-value cache upper limit, the key-value caches obtained in the prefill stage are deleted to obtain the deleted key-value caches, where the number of the deleted key-value caches is less than or equal to the key-value cache upper limit, achieving the technical effect of automatically adjusting the storage of the key-value cache and realizing the avoidance of a decline in the inference performance of the large language model due to excessive consumption of memory resources.
[0125] As an alternative solution, deleting the key-value caches obtained in the prefill stage to obtain the deleted key-value caches includes:
[0126] Delete the key-value cache from the initial token number to the rolling token number in the key-value cache obtained in the pre-filling stage to obtain the deleted key-value cache, where the initial token number and the rolling token number are used to determine the key-value cache to be deleted, and both the initial token number and the rolling token number are integer multiples of the quantity corresponding to the second length.
[0127] It should be noted that after the pre-filling stage, delete the key-value cache between the initial token number and the rolling token number from the key-value cache. Both the initial token number and the rolling token number are integer multiples of the quantity corresponding to the second length, avoiding retaining unnecessary intermediate key-value caches, reducing memory occupancy, ensuring the coherence of context information, effectively managing memory resources, avoiding memory overflow and fragmentation, and ensuring that the large language model retains key context information during the inference process. Thus, when processing ultra-long sequence inputs or streaming data, the inference efficiency and memory resource utilization rate can be improved.
[0128] Through the embodiments of the present application, delete the key-value cache from the initial token number to the rolling token number in the key-value cache obtained in the pre-filling stage to obtain the deleted key-value cache, where the initial token number and the rolling token number are used to determine the key-value cache to be deleted, and both the initial token number and the rolling token number are integer multiples of the quantity corresponding to the second length. The technical purpose of discarding unnecessary intermediate key-value caches and reducing memory occupancy is achieved, and further the technical effect of improving the inference efficiency and memory resource utilization rate of the large language model is realized.
[0129] As an optional solution, after deleting the key-value cache obtained in the pre-filling stage to obtain the deleted key-value cache, the method further includes:
[0130] S7-1, perform the following steps until the output token is obtained:
[0131] S7-2, through the large language model, combine the deleted key-value cache and the tokens obtained in the pre-filling stage to output the current new token and the current new key-value cache;
[0132] S7-3, when the current new token meets the termination condition, obtain the output token, where the output token includes the tokens obtained in the pre-filling stage and the new tokens obtained in the pre-filling stage;
[0133] S7-4, when the current new token does not meet the termination condition, through the large language model, combine the current new token and the current new key-value cache to output the next new token and the next new key-value cache, and use the next new token as the current new token and the next new key-value cache as the current new key-value cache.
[0134] In an alternative embodiment, the termination condition can be, but is not limited to, understood as a condition used to determine when to stop generating new tokens during the decoding process.
[0135] In an alternative embodiment, at the start of the decoding phase, the large language model can, but is not limited to, generate new tokens and corresponding key-value caches based on the key-value cache and token input generated during the prefill phase.
[0136] In an alternative embodiment, when the newly generated tokens reach a predefined termination condition, which can, but is not limited to, include generating a termination token or reaching the output length limit, the tokens generated during the prefill phase and the decoding phase are used as the final output tokens.
[0137] In an alternative embodiment, when the new tokens do not reach the termination condition, the large language model can, but is not limited to, utilize the existing tokens and key-value cache data to generate the next token and corresponding key-value cache, and then use the generated tokens and key-value cache for the next token generation process until the output requirements are met.
[0138] It should be noted that continuously generating new tokens based on the optimized key-value cache and token input until reaching the predefined termination condition ensures that the text generated by the model is not only coherent but also accurately reflects the context of the input sequence. By dynamically managing the key-value cache, unnecessary repeated calculations and memory consumption are avoided, improving the efficiency of model inference and the utilization rate of computing resources.
[0139] Through the embodiments of the present application, the large language model combines the deleted key-value cache and the tokens obtained in the prefill phase to output the current new token and the current new key-value cache; in the case where the current new token meets the termination condition, the output tokens are obtained, where the output tokens include the tokens obtained in the prefill phase and the new tokens obtained in the prefill phase; in the case where the current new token does not meet the termination condition, the large language model combines the current new token and the current new key-value cache to output the next new token and the next new key-value cache, and uses the next new token as the current new token and the next new key-value cache as the current new key-value cache. The technical purpose of continuously generating new tokens based on the optimized key-value cache and token input until reaching the termination condition is achieved, and the technical effect of improving the efficiency of large language model inference and the utilization rate of computing resources is realized.
[0140] As an alternative solution, obtaining the input tokens includes:
[0141] S8-1, obtaining the input tokens corresponding to a single token processing request;
[0142] The method further includes:
[0143] S8-2, obtain the first input token corresponding to the first token processing request and the second input token corresponding to the second token processing request;
[0144] S8-3, when the number of tokens in the first input token is less than the token processing limit, use a large language model to perform combined processing on the first input token and some tokens in the second input token, where the sum of the number of tokens in the first input token and the number of tokens in some tokens in the second input token is less than the token processing limit.
[0145] In an alternative embodiment, when the number of tokens in the first token request is less than the processing limit, the request can be combined with the tokens of the second token request through, but not limited to, a large language model and input as a whole into the model for processing, which can more effectively utilize computing resources, avoid resource waste caused by allocating independent resources for each request, and can also improve the inference speed and response time of the model.
[0146] It should be noted that when the number of tokens in a single token processing request is less than the processing limit, the tokens in multiple requests can be combined and processed as a whole. By combining and processing the requests, resource consumption can be significantly reduced, the processing speed and response ability of the model can be improved, while maintaining the inference quality, which can play an important role in applications such as dialogue systems, text generation services, and multi-task inference scenarios, and can optimize the resource allocation of the model to ensure that the large language model can run with relatively high efficiency when processing multiple requests.
[0147] Through the embodiments of the present application, obtain the input token corresponding to a single token processing request; obtain the first input token corresponding to the first token processing request and the second input token corresponding to the second token processing request; when the number of tokens in the first input token is less than the token processing limit, use a large language model to perform combined processing on the first input token and some tokens in the second input token, where the sum of the number of tokens in the first input token and the number of tokens in some tokens in the second input token is less than the token processing limit. Achieve the technical purpose of combining and processing the request with the tokens of the second token request through a large language model when the number of tokens in the first token request is less than the processing limit, and realize the technical effect of improving the processing efficiency of the large language model.
[0148] As an alternative solution, a single token processing request corresponds to a first token processing limit, and multiple token processing requests correspond to a second token processing limit, where the second token processing limit is less than or equal to the first token processing limit. When the number of tokens in the first input token is less than the token processing limit, use a large language model to perform combined processing on the first input token and some tokens in the second input token, including:
[0149] When the number of tokens of the first input token is less than the upper limit of the second token processing, through the large language model, the first input token and some tokens in the second input token are combined for processing.
[0150] In an alternative embodiment, when the number of tokens in the first token request is less than the upper limit of tokens that the large language model is set to process two or more requests simultaneously, the tokens in the first token and the second token request can be, but are not limited to, combined and then processed through the large language model. This can improve the utilization rate of computing resources, reduce the latency in the model inference process, process multiple token requests simultaneously, thereby improving the overall processing efficiency and throughput of the large language model, and ensuring that the model can handle more concurrent requests while running efficiently.
[0151] It should be noted that in a multi-task or high-concurrency scenario, to more effectively utilize computing resources, reduce resource waste, improve the inference speed and response time of the model, through combined processing, the data of multiple requests can be, but are not limited to, packed, making the processing of each batch more efficient. Combined processing can also alleviate the bottleneck problem in a resource-constrained environment to a certain extent, providing a more flexible solution for the deployment and operation of the large language model.
[0152] Through the embodiments of the present application, when the number of tokens of the first input token is less than the upper limit of the second token processing, through the large language model, the first input token and some tokens in the second input token are combined for processing. The technical purpose of processing multiple token requests simultaneously is achieved, and thus the technical effect of improving the processing efficiency of the large language model is realized.
[0153] As an alternative solution, in the process of inferring and processing input tokens through the large language model, it includes:
[0154] S9-1, when the remaining space in the memory management space corresponding to the large language model is less than the preset threshold, suspend the inference processing of the key-value cache;
[0155] S9-2, when the remaining space in the memory management space corresponding to the large language model is greater than or equal to the preset threshold, continue the inference processing of the key-value cache for which the inference processing has been suspended.
[0156] In an alternative embodiment, when the remaining memory in the memory management space used by the large language model is lower than the preset threshold, the further inference processing of the key-value cache can be, but is not limited to, suspended to avoid memory overflow or overuse of resources, effectively preventing the computing device from crashing due to insufficient memory, ensuring the stability and reliability of the large language model, and at the same time avoiding excessive consumption of resources and maintaining the efficient operation of the large language model.
[0157] In an alternative embodiment, once the remaining memory in the memory management space returns to or exceeds a preset threshold, inference processing on the key-value cache that was previously paused can be restarted, but is not limited to this, to continue generating output tokens. By dynamically monitoring memory usage and intelligently adjusting the processing flow based on the remaining space, the large language model can flexibly handle different loads and memory requirements, ensuring continuous service under resource-constrained conditions and avoiding inference interruptions due to insufficient memory.
[0158] It should be noted that by real-time monitoring of memory usage and dynamically adjusting the inference strategy, while ensuring the inference quality of the model, it is possible to effectively avoid performance degradation or service interruption caused by insufficient memory. For application scenarios that require real-time response and large-scale deployment, it can significantly improve the stability and user experience of the large language model. Additionally, by real-time monitoring of memory usage and dynamically adjusting the inference strategy, it is also possible to enable the large language model to handle more requests and tasks under limited computing resources.
[0159] Through the embodiments of the present application, when the remaining space corresponding to the memory management space of the large language model is less than the preset threshold, the inference processing of the key-value cache is paused; when the remaining space corresponding to the memory management space of the large language model is greater than or equal to the preset threshold, the inference processing of the key-value cache that has been paused is continued. It achieves the technical purpose of enabling the large language model to handle more requests and tasks under limited computing resources, and thus realizes the technical effect of improving the efficiency of the large language model in processing requests and tasks.
[0160] As an alternative solution, the above memory management method for the large language model further includes:
[0161] S10-1, obtaining the model parameters and data precision of the large language model;
[0162] S10-2, obtaining the memory space required for each key-value cache according to the model parameters and the data precision;
[0163] S10-3, using the memory space required for each key-value cache and the memory occupied by the large language model according to a preset weight, to obtain the maximum length of the key-value cache that the target memory management block allows to store.
[0164] In an alternative embodiment, the model parameters can include, but are not limited to, the number of layers, the number of key-value pair attention heads, and the dimension of the heads.
[0165] In an alternative embodiment, the data precision can include, but is not limited to, the representation precision during the calculation process, and can be, but is not limited to, 32-bit floating point or 16-bit floating point.
[0166] It should be noted that the capacity of the memory management block affects the number of disk I / O operations, which in turn affects the continuous data reading speed. The larger the capacity of the memory management block, the fewer the disk I / O operations and the faster the continuous data reading speed. However, if the capacity of the memory management block is too large, it may lead to waste of memory resources. In addition, if the capacity of the memory management block is too small, it will result in too many memory fragments, thus affecting the normal performance of the system. Therefore, for the overall performance of the system, it is necessary to set an appropriate capacity for the memory management block.
[0167] Furthermore, to set an appropriate capacity for the memory management block, it is necessary to obtain the size of the key-value cache corresponding to the token, which can be understood as obtaining the unit memory. In this embodiment, the unit memory is obtained through the following formula: unit memory = number of layers × number of key-value pair attention heads × dimension of the head × number of bytes occupied by data precision × 2. Here, since each key-value pair includes a key and a value, the number of bytes occupied by data precision should be multiplied by 2.
[0168] Through the embodiments of the present application, the model parameters and data precision of the large language model are obtained; according to the model parameters and the data precision, the memory space required for each key-value cache is obtained; using the memory space required for each key-value cache and the memory occupied by the large language model according to the preset weight, the maximum length of the key-value cache allowed to be stored in the target memory management block is obtained. Thus, the technical purpose of setting an appropriate capacity for the memory management block is achieved, and further the technical effect of improving the overall performance of the system is realized.
[0169] As an alternative solution, the above memory management method of the large language model is applied to the scenario of large language model inference.
[0170] In an alternative embodiment, in the prefill stage, if the number of input tokens exceeds the upper limit of the number of tokens that can be inferred and calculated by the board in one go or exceeds the set threshold of the number of tokens inferred in one go, the input is split for inference and the first token is generated; otherwise, the first token is directly inferred.
[0171] In an alternative embodiment, in the decode stage, if the length of the kvcache of the request exceeds a certain limit, a part of the kvcache corresponding to the intermediate blocks is discarded and the inference continues to generate new tokens.
[0172] For further illustration by way of example, optionally as Figure 3 shown, the specific steps are as follows:
[0173] S302, in the prefill stage: If the number of input tokens requested exceeds the upper limit, split the input for inference and generate the first token; otherwise, directly infer the first token stage;
[0174] S304, in the decode stage: If the length of the key-value cache (kvcache) for this request exceeds a certain limit, discard the kvcache corresponding to some intermediate memory management blocks (blocks), and continue the inference to generate a new token;
[0175] S306, terminate the inference when the generated token is a terminator or the number of generated tokens reaches the set threshold.
[0176] In an alternative embodiment, assume that when using the paged-attention technique to schedule multi-task inference of a large language model on a certain board, the size of the set block is bn, that is, each block can store bn kvcaches. In approximate inference, assume that the number of initial tokens reserved is in (where in % bn = 0), and the corresponding number of rolling tokens is rn (where rn % bn = 0). Here, the paged-attention technique can be understood, but not limited to, as an attention algorithm inspired by the virtual memory and paging concepts of traditional operating systems.
[0177] In an alternative embodiment, the steps of multi-task inference of a large language model are as follows:
[0178] Step 1, in the prefill stage, if the number of input request tokens qn exceeds the upper limit of the number of tokens that the board can infer and calculate at one time or exceeds the upper limit of the number of input tokens for one-time inference set (this upper limit is represented by tn, and in the paged-attention mode, tn % bn = 0), then enter Step 1.3 to split the input for inference and generate the first token; otherwise, enter Step 1.2 to directly infer the first token.
[0179] Step 1.2, pass the input request and the corresponding mask and position_id into the large language model to obtain the kvcache of all input requests and the first token for inference in the decode stage, where the mask is a lower triangular matrix of qn × qn, and the position_id is a vector from 0 to qn - 1.
[0180] Step 1.3, split the input request, perform inference separately, and generate the first token, specifically as follows;
[0181] Step 1.3.1, divide the request with qn tokens into n segments (n > 1), where the length of the first segment is fn, the length of the second to the (n - 1)-th segments is sn, and the length of the n-th segment is the remaining length nn, i.e., nn = qn - fn - (n - 2) * sn; where fn % bn = 0, fn > in + rn, fn ≤ tn, sn % bn = 0; sn + in + rn ≤ tn;
[0182] Step 1.3.2, pass the first segment of the request and the corresponding mask and position_id into the large language model to obtain the kvcache corresponding to fn requests for subsequent segmented request calculations, where the mask is a lower triangular matrix of fn×fn, and the position_id is a vector from 0 to fn - 1. If n > 2, go to Step 1.3.3, otherwise skip Step 1.3.3 and directly go to Step 1.3.4.
[0183] Step 1.3.3, for each of the second to the (n - 1)-th segments, called the i-th segment (2 ≤ i ≤ n - 1), perform the following operations in sequence: for all the previous kvcache (actually the kvcache corresponding to fn or sn + in + rn tokens), only keep the kvcache corresponding to the first in and the last rn tokens, discard the kvcache corresponding to the middle tokens, obtain the kvcache corresponding to in + rn tokens, and pass it and the i-th segment of the request and the corresponding mask and position_id into the large language model to get sn new kvcache for subsequent segmented request calculations, where the mask is a special matrix such as Figure 4 shown as a matrix of sn×(in + rn + sn) formed by concatenating a matrix of all 1s of sn×(in + rn) and a lower triangular matrix of sn×sn in the row direction, and the position_id is a vector from fn + (i - 2) * sn to fn + (i - 1) * sn - 1.
[0184] Step 1.3.4. For the nth segmented input request with a length of nn, let dnn = ⌈nn / bn⌉ × bn, where the ceiling function ⌈x⌉ represents the smallest integer greater than or equal to x. For all previous kvcaches (actually the kvcaches corresponding to fn or sn + in + rn tokens), discard the kvcaches corresponding to dnn tokens after in are discarded, obtaining the kvcaches corresponding to lnn (lnn is actually fn - dnn or sn + in + rn - dnn) tokens, and pass it, along with the nth segment request and the corresponding mask and position_id, into the large language model. The newly obtained nn kvcaches and the corresponding tokens are used for the calculation in the subsequent decode stage. Here, the mask is a special matrix, which is a nn×(nn + lnn) matrix formed by concatenating a nn×lnn all-ones matrix and a nn×nn lower triangular matrix in the row direction, and the position_id is a vector from fn+(n - 2)*sn to qn - 1.
[0185] It should be noted that the discard operations in Steps 1.3.3 and 1.3.4 are only for convenience of expression. In the actual inference process, the relevant kvcaches can be directly not saved during the calculation according to the above logic.
[0186] For further illustration, the calculation logic diagram of the entire prefill stage is as Figure 5 shown:
[0187] S502, Determine whether the number of requested tokens exceeds the threshold. If yes, execute S504; otherwise, execute S506.
[0188] S504, Divide the request with qn tokens into n segments (n > 1), where the length of the first segment is fn, the length of the second to the (n - 1)th segments is sn, and the length of the nth segment is the remaining length nn, that is, qn - fn - (n - 2)*sn.
[0189] S506, Pass the input request and the corresponding mask and position_id into the large language model to obtain the kvcaches of all input requests and the first token for inference in the decode stage. Here, the mask is a qn×qn lower triangular matrix, and the position_id is a vector from 0 to qn - 1.
[0190] S508, Calculate the kvcaches with a length of fn for the first segment.
[0191] S510, Determine whether n is greater than 2. If yes, execute S512; otherwise, directly execute S514.
[0192] S512. For each segment from the second segment to the (n - 1)-th segment, called the i-th segment (2 ≤ i ≤ n - 1), perform the following operations in sequence: For all previous kvcaches, only retain the kvcaches corresponding to the first in and the last m tokens, and pass them, along with the i-th segment request and the corresponding mask and position_id, into the large language model to obtain sn new kvcaches;
[0193] S514. For the segmented input request of the n-th segment with a length of nn, let dnn = [nn / bn] × bn. For all previous kvcaches, discard the kvcaches corresponding to the tokens after in in dnn, and pass them, along with the n-th segment request and the corresponding mask and position_id, into the large language model to obtain nn new kvcaches and the first token for inference in the decode stage.
[0194] Step 2. This step describes the computational inference process in the decode stage. When the exit condition mentioned in Step 3 is not reached, multiple inferences are performed in the decode stage. The details of each inference are shown in Steps 2.1 - 2.2.
[0195] Step 2.1. In the decode stage, if the number of existing kvcaches reaches the set support upper limit tn, delete the kvcaches corresponding to the tokens from the in-th to the (in + bn)-th token (counting from 1), and proceed to Step 2.2; if the number of existing kvcaches does not reach the set support upper limit tn, directly proceed to Step 2.2;
[0196] Step 2.2. Suppose there are currently din kvcaches corresponding to tokens. Then, use the retained kvcaches, the new token obtained from the previous inference, and the corresponding mask and position_id as the input to the large language model. Here, the mask is a two-dimensional all-ones matrix of 1 × (1 + din), and the position_id is a one-dimensional vector containing only the data din. The output is a new token and the corresponding kvcaches, which can be used for the next inference.
[0197] Step 3. Describes the exit judgment condition for the end of the large language model inference. That is, before each inference in Step 2, make the following judgments: Judge whether the newly generated token is the termination symbol of the model. If it is the termination symbol, the inference ends; Judge whether the total number of generated tokens reaches the set threshold. If it reaches the threshold, the inference ends.
[0198] For further illustration by example, the inference processes of Step 2 and Step 3 are as Figure 6 shown. The specific steps are as follows:
[0199] S602. Check whether the token is a terminator or whether the total number of tokens has reached the threshold. If so, execute S604; otherwise, execute S610.
[0200] S604. Check whether the current number of kvcaches has reached the set support upper limit tn. If so, execute S606; otherwise, execute S608.
[0201] S606. Delete the kvcaches corresponding to the tokens from the in - th to the in+bn - th.
[0202] S608. Assume that there are currently din tokens' corresponding kvcaches. Then, use the reserved kvcaches, the new tokens obtained from the previous inference, and the corresponding mask and position_id as the input to the large - language model, and the output is a new token and the corresponding kvcache.
[0203] S610. End the inference.
[0204] In an alternative embodiment, if the upper limit of the number of input tokens for a single inference is tn, and in the case of multiple batches, the upper limit of the number of input tokens for a single - batch single inference is stn (stn≤tn), then in the prefill stage, when the length of a certain request is less than tn, it can be concatenated with other requests or a part of other requests for continue - batch splicing. Therefore, the space occupied by the spliced requests, block_num*bn, is not greater than tn, where block_num is the number of blocks occupying the kvcache.
[0205] Furthermore, in the decode stage, multiple batches can perform decode inference simultaneously. If there is still some remaining kvcache space, new requests can be added, and the decode stage and the prefill stage can be combined for inference. The mask and position_id of each request are the same as those described in the scheme of Section 2.2.1. In the decode inference stage, when the kvcache space is insufficient, the kvcaches of some requests are transferred from the acceleration computing card to the CPU, and the corresponding requests pause the inference; when there is still some kvcache space, the kvcaches on the CPU are transferred back to the computing acceleration card, and the relevant requests continue to be inferred.
[0206] It should be noted that the embodiments of the present application can process multiple requests simultaneously, improving the processing speed of the requests. Additionally, even if the number of input tokens for a single request is less than tn but greater than stn, using the approximate inference scheme proposed in this embodiment will have a certain improvement in computing performance compared to the existing computing scheme.
[0207] It should be noted that each block can store the kvcache information corresponding to a certain number of tokens. The corresponding number can be called the block_size, and the occupied memory space is called the block_mem. Each block can be assigned a unique block ID, corresponding to a section of address; the number of blocks that can be allocated in memory is the block_num. Each block only stores the kvcache of the same request and does not store the kvcache of other requests.
[0208] Furthermore, for different computing acceleration cards and different models, the value of the block_size is generally 32 / 64 / 128 or 256, generally a power of 2, and the reading speed of such a size on the board is generally relatively fast;
[0209] For different large language models, the space size required for the kvcache corresponding to each token is different, and its value is unit_mem = num_hidden_layers×num_key_value_heads×head_dim×bytes×2, where num_hidden_layers is the number of layers of the transformers structure in the large language model, num_key_value_heads and head_dim are the number of kv heads and the head dimension in the multi-head self-attention mechanism in the large language model, bytes is the number of bytes occupied by the computing data precision, and 2 means one copy of k and v each; among them, the values of num_hidden layers, num_key_value_heads, and head_dim may be different for each model.
[0210] For different models, since the space size required for the kvcache corresponding to each token is different, even if the block_size in the block is the same, the size of the block_mem corresponding to this block is also different. The available space size Memb for allocating blocks in memory is "total board memory - weight occupied memory - temporary calculation and system occupied memory", then block_num = Memb divided by block_mem and rounded down.
[0211] For a specific model, the block_size determines the block_mem, and the available space for allocating blocks in the board affects the size of block_num. The block_size should not be too small, as this will affect the data reading and storage speed; the block_size should not be too large either, as this will result in too small a block_num, fewer blocks available for different requests, which is not conducive to memory scheduling, and will also lead to too few blocks available for discarding, and it is only possible to discard once after storing a relatively large amount of kvcache each time.
[0212] In addition, the method in this embodiment is applicable not only to single-card inference, but also to multi-card inference (including multi-card pipelining parallelism, multi-card tensor parallelism, multi-card hybrid parallelism) or multi-machine inference.
[0213] Through the embodiments of the present application, the inference in the prefill stage and the decode stage of the large language model is optimized respectively, which can be directly docked with the existing large language model inference service system based on the paged-attention technology, achieving the technical purpose of only modifying the kvcache space management strategy of the request without modifying the calculation logic and the underlying operators, and further realizing the technical effect of improving the calculation performance of the large language model.
[0214] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), including several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present application.
[0215] In this embodiment, a memory management device for a large language model is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0216] Figure 7 is a structural block diagram of the memory management device for a large language model according to an embodiment of the present application. As Figure 7 shown, the device includes:
[0217] The first acquisition unit 702 is configured to acquire input tokens, where the input tokens are basic units processed by a large language model;
[0218] The inference unit 704 is configured to perform inference processing on the input tokens through the large language model to obtain a key-value cache of a first length, where the key-value cache of the first length is the key-value cache corresponding to the tokens to be deleted from the input tokens;
[0219] The storage unit 706 is configured to, when the first length is less than a second length, set a target memory management block to store the key-value cache of the first length, where the second length is the maximum length that the target memory management block allows to store the key-value cache;
[0220] The setting unit 708 is configured to, when the first length is equal to the second length, set the target memory management block to cancel storing the key-value cache of the first length.
[0221] For specific embodiments, reference can be made to the examples shown in the memory management method of the above large language model, and details are not described herein again.
[0222] As an optional solution, the inference unit 704 includes: a first inference module configured to perform inference processing on the input tokens in the pre-fill stage through the large language model.
[0223] For specific embodiments, reference can be made to the examples shown in the memory management method of the above large language model, and details are not described herein again.
[0224] As an optional solution, the inference module includes: a splitting sub-module configured to, when the number of tokens in the input tokens is greater than the inference token upper limit, split the input tokens and perform sequential inference processing on the split input tokens, where the split input tokens include at least two segments of input tokens, and the number of tokens in each segment of the at least two segments of input tokens is less than or equal to the inference token upper limit.
[0225] For specific embodiments, reference can be made to the examples shown in the memory management method of the above large language model, and details are not described herein again.
[0226] As an alternative solution, the above-mentioned splitting sub-module includes: a setting subunit configured to set the input word of the first segment among the above n segments of input word tokens as the first number of word tokens, set the number of input word tokens of the i-th segment among the above n segments of input word tokens as the second number of word tokens, and set the number of input word tokens of the n-th segment among the above n segments of input word tokens as the third number of word tokens, where i is an integer greater than 1 and less than or equal to n. The above first number of word tokens is an integer multiple of the quantity corresponding to the above second length. The above first number of word tokens is greater than the sum of the initial number of word tokens and the rolling number of word tokens. The above initial number of word tokens and the above rolling number of word tokens are used to determine the key-value cache to be deleted. The above initial number of word tokens is an integer multiple of the quantity corresponding to the above second length. The above second number of word tokens, the above initial number of word tokens, and the above rolling number of word tokens are all integer multiples of the quantity corresponding to the above second length. The sum of the above second number of word tokens and the above initial number of word tokens and the above rolling number of word tokens is less than or equal to the above inference word token upper limit.
[0227] For specific embodiments, reference can be made to the examples shown in the above-mentioned memory management method of the large language model, and details will not be elaborated in this example.
[0228] As an alternative solution, the above-mentioned setting subunit includes: a first inference component configured to, when n is greater than 2, perform sequential inference processing on the split input word tokens by using a first inference method; a second inference component configured to, when n is equal to 2, perform sequential inference processing on the split input word tokens by using a second inference method, where the above second inference method is different from the above first inference method.
[0229] For specific embodiments, reference can be made to the examples shown in the above-mentioned memory management method of the large language model, and details will not be elaborated in this example.
[0230] As an alternative solution, the above-mentioned first inference component includes: a first inference sub-component configured to perform inference processing on f input words of the first segment among the above n segments of input word tokens to obtain f key-value caches corresponding to the f input words; a second inference sub-component configured to, in combination with the above f key-value caches, perform inference processing on s input words of the i-th segment among the above n segments of input word tokens to obtain s key-value caches corresponding to the s input words, where i is an integer greater than 1 and less than n; a third inference sub-component configured to, in combination with the above s key-value caches, perform inference processing on m input words of the n-th segment among the above n segments of input word tokens to obtain m key-value caches.
[0231] For specific embodiments, reference can be made to the examples shown in the above-mentioned memory management method of the large language model, and details will not be elaborated in this example.
[0232] As an alternative solution, the above-mentioned second inference sub-component includes: a first determination component, configured to determine, from the above-mentioned f key-value caches, the first target number of key-value caches of the above-mentioned initial number of word tokens and the last above-mentioned rolling number of word tokens; a first inference component, configured to perform an inference process on the above-mentioned s input words by combining the above-mentioned first target number of key-value caches to obtain the above-mentioned s key-value caches.
[0233] For specific embodiments, reference may be made to the examples shown in the above-mentioned memory management method of the large language model, which will not be elaborated herein.
[0234] As an alternative solution, the above-mentioned third inference sub-component includes: a second determination component, configured to determine, from the above-mentioned first target number of key-value caches and the above-mentioned s key-value caches, the second target number of key-value caches of the above-mentioned initial number of word tokens and the last above-mentioned rolling number of word tokens; a second inference component, configured to perform an inference process on the above-mentioned m input words by combining the above-mentioned second target number of key-value caches to obtain the above-mentioned m key-value caches.
[0235] For specific embodiments, reference may be made to the examples shown in the above-mentioned memory management method of the large language model, which will not be elaborated herein.
[0236] As an alternative solution, the above-mentioned second inference component includes: a fourth inference sub-component, configured to perform an inference process on the f input words in the first segment of the above-mentioned n segments of input word tokens to obtain the f key-value caches corresponding to the above-mentioned f input words; a fifth inference sub-component, configured to perform an inference process on the m input words in the nth segment of the above-mentioned n segments of input word tokens by combining the above-mentioned f key-value caches to obtain m key-value caches.
[0237] For specific embodiments, reference may be made to the examples shown in the above-mentioned memory management method of the large language model, which will not be elaborated herein.
[0238] As an alternative solution, the above-mentioned fifth inference sub-component includes: a third determination component, configured to determine, from the above-mentioned f key-value caches, the second target number of key-value caches, where the above-mentioned second target number of key-value caches is the key-value caches after the above-mentioned f key-value caches discard the third target number of key-value caches, and the above-mentioned target number is the smallest integer obtained by rounding up the quotient of m and the number corresponding to the above-mentioned second length; a third inference component, configured to perform an inference process on the above-mentioned m input words by combining the above-mentioned second target number of key-value caches to obtain the above-mentioned m key-value caches.
[0239] For specific embodiments, reference may be made to the examples shown in the above-mentioned memory management method of the large language model, which will not be elaborated herein.
[0240] As an alternative solution, the above-mentioned inference module includes: a decoding sub-module, which is used to perform inference processing in the decoding stage on the key-value cache obtained in the above-mentioned pre-filling stage through the above-mentioned large language model.
[0241] For specific embodiments, reference can be made to the examples shown in the above-mentioned memory management method of the large language model, and these examples will not be elaborated here.
[0242] As an alternative solution, the above-mentioned inference sub-module includes: a deletion sub-unit, which is used to perform deletion processing on the key-value cache obtained in the above-mentioned pre-filling stage when the number of the key-value caches obtained in the above-mentioned pre-filling stage is greater than the upper limit of the key-value cache, so as to obtain a deleted key-value cache, where the number of the above-mentioned deleted key-value caches is less than or equal to the upper limit of the key-value cache.
[0243] For specific embodiments, reference can be made to the examples shown in the above-mentioned memory management method of the large language model, and these examples will not be elaborated here.
[0244] As an alternative solution, the above-mentioned deletion sub-unit includes: a deletion component, which is used to perform deletion processing on the key-value cache from the initial token number to the rolling token number in the key-value cache obtained in the above-mentioned pre-filling stage, so as to obtain the above-mentioned deleted key-value cache, where the above-mentioned initial token number and the above-mentioned rolling token number are used to determine the key-value cache to be deleted, and both the above-mentioned initial token number and the above-mentioned rolling token number are integer multiples of the quantity corresponding to the above-mentioned second length.
[0245] For specific embodiments, reference can be made to the examples shown in the above-mentioned memory management method of the large language model, and these examples will not be elaborated here.
[0246] As an alternative solution, the above-mentioned deletion sub-unit includes: a first output component, which is used to output a current new token and a current new key-value cache through the above-mentioned large language model by combining the above-mentioned deleted key-value cache and the token obtained in the above-mentioned pre-filling stage; an acquisition component, which is used to acquire the above-mentioned output token when the above-mentioned current new token meets the termination condition, where the above-mentioned output token includes the token obtained in the above-mentioned pre-filling stage and the new token obtained in the above-mentioned pre-filling stage; a second output component, which is used to output the next new token and the next new key-value cache through the above-mentioned large language model by combining the above-mentioned current new token and the above-mentioned current new key-value cache when the above-mentioned current new token does not meet the above-mentioned termination condition, and use the above-mentioned next new token as the above-mentioned current new token and use the next new key-value cache as the above-mentioned current new key-value cache.
[0247] For specific embodiments, reference can be made to the examples shown in the above-mentioned memory management method of the large language model, and these examples will not be elaborated here.
[0248] As an alternative, the above-mentioned first acquisition unit 702 includes: a first acquisition module for acquiring an input token corresponding to a single-token processing request; a second acquisition module for acquiring a first input token corresponding to a first token processing request and a second input token corresponding to a second token processing request; and a combination module for, when the token amount of the above-mentioned first input token is less than the token processing upper limit, performing combination processing on the above-mentioned first input token and some tokens of the above-mentioned second input token through the above-mentioned large language model, wherein the sum of the token amount of the above-mentioned first input token and the token amount of some tokens of the above-mentioned second input token is less than the above-mentioned token processing upper limit.
[0249] For specific embodiments, reference may be made to the examples shown in the memory management method of the above-mentioned large language model, and details are not described herein again in this example.
[0250] As an alternative, the above-mentioned combination module includes: a combination sub-module for, when the token amount of the above-mentioned first input token is less than the above-mentioned second token processing upper limit, performing combination processing on the above-mentioned first input token and some tokens of the above-mentioned second input token through the above-mentioned large language model.
[0251] For specific embodiments, reference may be made to the examples shown in the memory management method of the above-mentioned large language model, and details are not described herein again in this example.
[0252] As an alternative, the above-mentioned inference unit 704 includes: a pause module for pausing the inference processing of the above-mentioned key-value cache when the remaining amount of the space corresponding to the memory management space of the above-mentioned large language model is less than a preset threshold; and a resume module for resuming the inference processing of the key-value cache whose inference processing has been paused when the remaining amount of the space corresponding to the memory management space of the above-mentioned large language model is greater than or equal to the preset threshold.
[0253] For specific embodiments, reference may be made to the examples shown in the memory management method of the above-mentioned large language model, and details are not described herein again in this example.
[0254] As an alternative, the above-mentioned device further includes: a second acquisition unit for acquiring the model parameters and data precision of the above-mentioned large language model; a third acquisition unit for acquiring the memory space required for each of the above-mentioned key-value caches according to the above-mentioned model parameters and the above-mentioned data precision; and a fourth acquisition unit for using the memory space required for each of the above-mentioned key-value caches and the memory occupied by the above-mentioned large language model according to a preset weight to acquire the maximum length of the key-value cache that the above-mentioned target memory management block allows to store.
[0255] For specific embodiments, reference may be made to the examples shown in the memory management method of the above-mentioned large language model, and details are not described herein again in this example.
[0256] It should be noted that the above virtual devices (modules, units, sub - modules, sub - units, components, etc.) can be implemented by software or hardware. For the latter, it can be achieved in the following ways, but not limited to: the above virtual devices are all located in the same processor; or, the above virtual devices are respectively located in different processors in any combination form.
[0257] An embodiment of the present application also provides a computer - readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any one of the above - mentioned method embodiments when running.
[0258] In an exemplary embodiment, the above - mentioned computer - readable storage medium may include, but is not limited to: USB flash drive, read - only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disk and other various media that can store computer programs.
[0259] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above - mentioned method embodiments.
[0260] In an exemplary embodiment, the above - mentioned electronic device may further include a transmission device and an input / output device. Wherein, the transmission device is connected to the above - mentioned processor, and the input / output device is connected to the above - mentioned processor.
[0261] Specific examples in this embodiment can refer to the examples described in the above - mentioned embodiments and exemplary embodiments, and will not be elaborated here.
[0262] Obviously, those skilled in the art should understand that the above - mentioned virtual devices or steps of the present application can be implemented by a general - purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program code executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0263] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the protection scope of the present application.
Claims
1. A memory management method for a large language model, characterized in that, Including: Obtain input tokens, where the input tokens are the basic units processed by the large language model; Through the large language model, perform inference processing on the input tokens in the prefill stage to obtain a key-value cache of a first length, where the key-value cache of the first length is the key-value cache corresponding to the tokens to be deleted from the input tokens; When the number of tokens in the input tokens is greater than the inference token limit, split the input tokens and perform sequential inference processing on the split input tokens, where the split input tokens include at least two segments of input tokens, and the number of tokens in each segment of the at least two segments of input tokens is less than or equal to the inference token limit; When the split input tokens include n segments of input tokens, the splitting of the input tokens includes: setting the number of tokens in the first segment of the n segments of input tokens to a first number of tokens, setting the number of tokens in the i-th segment of the n segments of input tokens to a second number of tokens, and setting the number of tokens in the n-th segment of the n segments of input tokens to a third number of tokens, where i is an integer greater than 1 and less than or equal to n, the first number of tokens is an integer multiple of the quantity corresponding to the second length, the first number of tokens is greater than the sum of the initial number of tokens and the rolling number of tokens, the initial number of tokens and the rolling number of tokens are used to determine the key-value cache to be deleted, the initial number of tokens is an integer multiple of the quantity corresponding to the second length, the second number of tokens, the initial number of tokens, and the rolling number of tokens are all integer multiples of the quantity corresponding to the second length, and the sum of the second number of tokens and the initial number of tokens and the rolling number of tokens is less than or equal to the inference token limit; When the first length is less than the second length, set the target memory management block to store the key-value cache of the first length, where the second length is the maximum length allowed for the target memory management block to store the key-value cache; When the first length is equal to the second length, set the target memory management block to cancel storing the key-value cache of the first length.
2. The method according to claim 1, characterized in that, The split input tokens include n segments of input tokens, where n is an integer greater than 1, and the sequential inference processing on the split input tokens includes: When n is greater than 2, use the first inference method to perform sequential inference processing on the split input tokens; When n is equal to 2, use the second inference method to perform sequential inference processing on the split input tokens, where the second inference method is different from the first inference method.
3. The method according to claim 2, wherein The using the first inference method to perform sequential inference processing on the split input tokens includes: Perform inference processing on the f input tokens in the first segment of the n segments of input tokens to obtain f key-value caches corresponding to the f input tokens; Combined with the f key-value caches, perform inference processing on the s input tokens in the i-th segment of the n segments of input tokens to obtain s key-value caches corresponding to the s input tokens, where i is an integer greater than 1 and less than n; In combination with the s key-value caches, perform inference processing on the m input words in the nth segment of the n segments of input tokens to obtain m key-value caches.
4. The method according to claim 3, characterized in that, The step of performing inference processing on the s input words in the ith segment of the n segments of input tokens in combination with the f key-value caches to obtain s key-value caches includes: Determine, from the f key-value caches, the key-value caches of the first target quantity that is the sum of the number of the initial tokens and the number of the rolling tokens; In combination with the key-value caches of the first target quantity, perform inference processing on the s input words to obtain the s key-value caches.
5. The method according to claim 3, characterized in that The step of performing inference processing on the m input words in the nth segment of the n segments of input tokens in combination with the s key-value caches to obtain m key-value caches includes: Determine, from the key-value caches of the first target quantity and the s key-value caches, the key-value caches of the second target quantity that is the sum of the number of the initial tokens and the number of the rolling tokens; In combination with the key-value caches of the second target quantity, perform inference processing on the m input words to obtain the m key-value caches.
6. The method according to claim 2, wherein The step of sequentially performing inference processing on the split input tokens by using the second inference method includes: Perform inference processing on the f input words in the first segment of the n segments of input tokens to obtain f key-value caches corresponding to the f input words; In combination with the f key-value caches, perform inference processing on the m input words in the nth segment of the n segments of input tokens to obtain m key-value caches.
7. The method according to claim 6, wherein The step of performing inference processing on the m input words in the nth segment of the n segments of input tokens in combination with the f key-value caches to obtain m key-value caches includes: Determine, from the f key-value caches, the key-value caches of the second target quantity, where the key-value caches of the second target quantity are the key-value caches obtained by discarding the key-value caches of the third target quantity from the f key-value caches, and the target quantity is the smallest integer obtained by rounding up the quotient of m and the quantity corresponding to the second length; In combination with the key-value caches of the second target quantity, perform inference processing on the m input words to obtain the m key-value caches.
8. The method according to claim 1, characterized in that, After performing inference processing in the pre-filling stage on the input tokens by using the large language model, the method further includes: Performing inference processing in the decoding stage on the key-value caches obtained in the pre-filling stage by using the large language model.
9. The method according to claim 8, wherein The step of performing inference processing in the decoding stage on the key-value caches obtained in the pre-filling stage by using the large language model includes: In the case where the length of the key-value caches obtained in the pre-filling stage is greater than the key-value cache upper limit, perform deletion processing on the key-value caches obtained in the pre-filling stage to obtain the deleted key-value caches, where the length of the deleted key-value caches is less than or equal to the key-value cache upper limit.
10. The method according to claim 9, characterized in that, The step of performing deletion processing on the key-value caches obtained in the pre-filling stage to obtain the deleted key-value caches includes: Delete the key-value cache from the initial token number to the rolling token number in the key-value cache obtained in the pre-filling stage to obtain the deleted key-value cache, where the initial token number and the rolling token number are used to determine the key-value cache to be deleted, and both the initial token number and the rolling token number are integer multiples of the quantity corresponding to the second length.
11. The method according to claim 9, wherein After deleting the key-value cache obtained in the pre-filling stage to obtain the deleted key-value cache, the method further includes: Execute the following steps until the output token is obtained: Through the large language model, combine the deleted key-value cache and the tokens obtained in the pre-filling stage to output the current new token and the current new key-value cache; When the current new token meets the termination condition, obtain the output token, where the output token includes the tokens obtained in the pre-filling stage and the new tokens obtained in the pre-filling stage; When the current new token does not meet the termination condition, through the large language model, combine the current new token and the current new key-value cache to output the next new token and the next new key-value cache, and use the next new token as the current new token and the next new key-value cache as the current new key-value cache.
12. The method according to claim 1, wherein obtaining the input token includes: obtaining the input token corresponding to a single token processing request; The method further includes: obtaining the first input token corresponding to the first token processing request and the second input token corresponding to the second token processing request; When the token quantity of the first input token is less than the token processing upper limit, through the large language model, perform a combination process on the first input token and some tokens in the second input token, where the sum of the token quantity of the first input token and the token quantity of some tokens in the second input token is less than the token processing upper limit.
13. The method according to claim 12, wherein The single token processing request corresponds to a first token processing upper limit, and multiple token processing requests correspond to a second token processing upper limit, where the second token processing upper limit is less than or equal to the first token processing upper limit. When the token quantity of the first input token is less than the token processing upper limit, performing a combination process on the first input token and some tokens in the second input token through the large language model includes: When the token quantity of the first input token is less than the second token processing upper limit, through the large language model, perform a combination process on the first input token and some tokens in the second input token.
14. The method according to any one of claims 1 to 13, characterized in that, During the process of performing inference processing on the input token through the large language model, it includes: When the remaining space corresponding to the memory management space of the large language model is less than the preset threshold, suspend the inference processing of the key-value cache; When the remaining space corresponding to the memory management space of the large language model is greater than or equal to the preset threshold, continue the inference processing on the key-value cache whose inference processing has been suspended.
15. The method according to any one of claims 1 to 13, characterized in that, The method further includes: Obtain the model parameters and data precision of the large language model; According to the model parameters and the data precision, obtain the memory space required for each of the key-value caches; Utilize the memory space required for each of the key-value caches and the memory occupied by the large language model according to a preset weight to obtain the maximum length that the target memory management block allows to store key-value caches.
16. A computer-readable storage medium, characterized in that A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 15 are implemented.
17. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the computer program, the steps of the method described in any one of claims 1 to 15 are implemented.
Citation Information
Patent Citations
Local caching method and related equipment
CN117632793A
Data transmission method and device, electronic equipment and storage medium
CN118550857A