Memory allocation method capable of optimizing memory usage of large language model

By dynamically allocating memory in chunk units based on token generation and probability of sentence completion, the method optimizes key-value cache usage in large language models, addressing memory efficiency challenges and improving performance.

WO2025127195A1PCT designated stage expired Publication Date: 2025-06-19KOREA ELECTRONICS TECH INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2023/020533
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2023-12-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The increasing memory requirements of large language models, particularly due to large key-value caches, make it difficult to operate these models efficiently, as they consume excessive memory resources and hinder batching improvements for execution performance.

Method used

A dynamic memory allocation method that divides the memory used by the language model in the key-value cache into chunk units and allocates memory in stages, optimizing the key-value cache by allocating only the necessary memory for each operation and adjusting based on token generation and probability of sentence completion.

Benefits of technology

This approach allows for the operation of large language models with reduced memory usage, enhancing energy efficiency and improving service performance by dynamically managing memory allocation across multiple services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2023020533_19062025_PF_FP_ABST
    Figure KR2023020533_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a memory allocation method capable of optimizing the memory usage of a large language model. A memory allocation method according to an embodiment of the present invention divides memory that a language model uses in a KV cache into chunks, and dynamically allocates the memory to the language model step by step in chunks. Accordingly, large language models can be driven using less memory than in the prior art and operate with greater energy efficiency compared to the operation modes of existing language models, and the service performance of large language models can be increased when driving multiple services.
Need to check novelty before this filing date? Find Prior Art

Description

Memory allocation methods that can optimize memory usage for large language models.

[0001] The present invention relates to memory management / allocation, and more particularly, to a memory allocation method capable of optimizing memory usage when executing a large language model.

[0002] When executing a conventional large language model, a key-value cache is pre-allocated using the maximum sequence length supported by the language model, and large language model operations are performed using the pre-allocated memory area.

[0003] However, as the memory size required to run large language models continues to increase, there is a need to optimize the memory area used by large language models, and there is a problem that running large language models becomes increasingly difficult due to excessively large key-value caches.

[0004] Additionally, there is a problem that batching to improve the execution performance of large language models becomes difficult due to the memory area of ​​the key-value cache.

[0005] The present invention has been devised to solve the above problems, and the purpose of the present invention is to provide a method for dynamically allocating memory of a key-value cache in stages as a means for optimizing a key-value cache of a large language model.

[0006] A memory allocation method according to one embodiment of the present invention for achieving the above object includes a step of dividing memory used by a language model for a KV (Key, Value) cache into chunk units; and a step of dynamically allocating memory to the language model in stages into chunk units.

[0007] The allocation step includes a step of allocating memory of 1 Chunk size for the KV cache when the language model is first operated and initializing the counter to 0; a step of operating the language model once and increasing the counter by 1; and a step of terminating the language model operation when the output token of the language model is a token indicating the end of sentence generation.

[0008] 1 Chunk size of memory may be the size of memory required to store the KV cache for a certain number of tokens.

[0009] The counter can be the number of tokens generated from the current chunk.

[0010] The assignment step may further include a step of comparing a counter value with a counter threshold value when the output token of the language model is not a token indicating the end of sentence generation; and a step of operating the language model once and increasing the counter by 1 when the value of the counter is less than the counter threshold value.

[0011] The threshold value may be a threshold value of the number of tokens to prepare for the next memory allocation.

[0012] The allocation step may further include: comparing a probability value that a sentence will be terminated within the currently allocated KV cache without additionally using the KV cache with a threshold probability value when the value of the counter is greater than the threshold probability value; allocating additional memory of the size of 1 Chunk for the KV cache when the probability value is less than the threshold probability value; and initializing the counter and recalculating the counter threshold value.

[0013] The allocation step may further include: a step of checking whether the KV cache capacity will be insufficient in the next execution if the probability value is greater than the threshold probability value; and a step of additionally allocating memory of the size of 1 Chunk for the KV cache if the KV cache capacity is determined to be insufficient.

[0014] The allocation step may further include a step of running the language model once and increasing the counter by 1, if it is confirmed that the KV cache capacity will not be insufficient.

[0015] According to another aspect of the present invention, a language model execution system is provided, characterized in that it includes a memory in which a storage space used by a language model for a KV (Key, Value) cache is divided into chunk units; and a processor that dynamically allocates memory to the language model in stages in chunk units.

[0016] According to another aspect of the present invention, a memory allocation method is provided, comprising: a step of operating a language model once and increasing a counter by 1; a step of comparing a counter value with a counter threshold value when an output token of the language model is not a token indicating the end of sentence generation; a step of comparing a probability value that a sentence will end within a currently allocated KV (Key, Value) cache without additionally using the KV (Key, Value) cache with a threshold probability value when the value of the counter is greater than the threshold probability value; and a step of additionally allocating memory of a size of 1 Chunk for the KV cache when the probability value is less than the threshold probability value.

[0017] According to another aspect of the present invention, a language model execution system is provided, characterized in that it includes a memory in which the storage space used by the language model for the KV (Key, Value) cache is divided into chunk units; a processor which operates the language model once and increases a counter by 1, compares the counter value with a counter threshold value when the output token of the language model is not a token indicating the end of sentence generation, and compares a probability value that the sentence will end within the currently allocated KV cache without additionally using the KV (Key, Value) cache with a threshold probability value when the value of the counter is greater than the threshold probability value, and additionally allocates memory of the size of 1 chunk for the KV cache when the probability value is less than the threshold probability value.

[0018] As described above, according to embodiments of the present invention, it is possible to operate a large language model with less memory than before, thereby enabling energy-efficient operation compared to the operation method of existing language models.

[0019] In particular, according to embodiments of the present invention, by minimizing the memory usage of the Key-Value cache, memory allocation can be dynamically managed when multiple services are operated, thereby improving the service performance of a large language model.

[0020] Figure 1. Large language model structure to which the present invention can be applied.

[0021] Figure 2. Example of the operation of the language model attention layer.

[0022] Figure 3. Memory allocation method for KV cache memory optimization according to one embodiment of the present invention.

[0023] Figure 4. Language model execution system according to another embodiment of the present invention.

[0024] Hereinafter, the present invention will be described in more detail with reference to the drawings.

[0025] An embodiment of the present invention presents a memory allocation method capable of optimizing the memory usage of a large language model. This technology relates to an allocation strategy and memory management algorithm for running a large language model while minimizing the usage of the Attention Layer's key-value cache.

[0026] Figure 1 illustrates the structure of a large-scale language model to which the present invention can be applied. Initial input data is input to a decoder block after undergoing input embedding and positional encoding.

[0027] The data input at this time is determined by a parameter called Sequence length, which means the maximum number of tokens that the language model can understand and process.

[0028] The decoder block is largely divided into a multi-head self-attention layer and a linear layer. The multi-head self-attention layer calculates query, key, and value vectors for the input data. The linear layer uses the calculated query, key, and value vectors to perform matrix multiplication with the weights to calculate the final output value of the decoder block.

[0029] These decoder blocks are performed repeatedly depending on the size of the entire language model. For example, GPT-3, one of the large language models, performs these decoder block operations 96 times.

[0030] When the decoder block's operation is finished, the output value is converted into a probability value for a specific token through the execution of the Linear layer and Softmax layer, one of the tokens with a high output probability is selected and output, the output token is added to the existing input data and inputted back into the language model, and the language model is re-executed.

[0031] The language model's way of working is to repeat these actions until sentence generation is complete.

[0032] Figure 2 is a detailed diagram illustrating the operation of the language model's Attention Layer. Figure 2 assumes that the input sentence is a 4-token sentence. During the initial execution (Step 1), Query, Key, and Value vectors are generated for each input token, and these are multiplied to derive the result. The resulting Key and Value vectors are then stored in memory.

[0033] From the second execution onwards, after generating a query vector for a newly added token (the token generated as the final result of the previous Step 1), since the key and value vectors for the previous tokens are identical, the values ​​stored in memory are retrieved and reused, and a key and value vector is generated for only one newly input token. Once all vectors are generated, multiplication is performed again and the attention result for the newly added token is derived.

[0034] Once generated, these Key and Value vector values ​​are continuously reused until sentence generation is complete, and this is called the KV Cache. The memory space used for this KV Cache is typically reserved in advance when the model is first executed, based on the maximum Sequence Length the model can process. This reserved space is then continuously utilized during model execution.

[0035] However, the sequence lengths of recently announced language models are steadily increasing to improve performance. Consequently, the memory space consumed by the KV cache is also significantly hindering the smooth execution of the model. For example, in the case of ChatGPT, which has a sequence length of 8,192 and a model dimension of 12,288, the KV cache accounts for approximately 10% of the total memory usage, and this figure continues to increase as the sequence length increases.

[0036] The simplest way to optimize the capacity of this KV cache is to dynamically allocate new memory space each time a new token is computed. However, large language models typically run on separate accelerators, such as GPUs. Dynamically allocating and managing the accelerator's memory space incurs significant performance penalties, making it difficult to utilize. Therefore, a separate memory management strategy and algorithm are needed to optimize the memory space for the KV cache.

[0037] FIG. 3 is a flowchart illustrating a memory allocation method for optimizing the memory of a KV cache according to one embodiment of the present invention. The method according to the embodiment of the present invention divides the memory used by the language model in the KV cache into specific chunk units and allocates only one chunk at a time, thereby maximizing memory space efficiency and minimizing the overhead of dynamically allocating memory for each execution.

[0038] To this end, when the language model first runs, memory sized 1 Chunk is allocated for the KV cache, and the counter (Cntchunk) is initialized to 0 (S110). The 1 Chunk memory is defined as the size of the memory required to store the KV cache for a specific number of tokens, and this value can be changed according to user convenience. Additionally, the counter (Cntchunk) indicates the number of tokens generated in the current Chunk.

[0039] After the next language model is run once and the counter (Cntchunk) is increased by 1 (S120), if the output token, which is the result of the language model execution, is the [End] token indicating the end of sentence generation (S130-Y), the language model operation is terminated.

[0040] If the language model operation is not terminated, i.e., if the output token of the language model is not an [End] token (S130-N), the counter value and the NThres value are compared, and if the counter value is smaller (S140-N), the operation returns to executing the language model once again without any additional operations (S120).

[0041] At this time, the initial NThres value represents the threshold for the number of tokens to prepare for the next memory allocation, and can be changed according to user convenience. For example, if the chunk size is set to store 128 tokens and the NThres value is 100, a decision is made on whether to allocate an additional memory chunk after generating 100 tokens. Before generating 100 tokens, the model is continuously executed repeatedly without making a decision on memory allocation.

[0042] Meanwhile, if the value of the counter is NThres or greater (S140-Y), the probability that the sentence will be terminated within the currently allocated KV cache without using additional KV cache (Pend) is compared with the threshold probability value (PThreshold) (S150).

[0043] The Pend value can be estimated in several ways. For example, if the memory size capable of storing 128 tokens is set to 1 Chunk and 100 tokens have been generated, the probability that sentence generation will end within 28 tokens (128-100) is lower than the critical probability. In this case, the probability that model operation will end within the remaining memory capacity is low.

[0044] If the probability of termination (Pend) within the currently allocated KV cache is less than the threshold probability value (PThreshold) as a result of the comparison at step S150 (S150-N), memory of 1 Chunk size is additionally allocated for the KV cache (S160), the counter (Cntchunk) is initialized, and NThres is recalculated (S170).

[0045] In step S170, NThres may recalculate the remaining unused memory capacity by considering the empty capacity rather than using the initial value as is if a new chunk is allocated without fully utilizing the previous chunk. After the additional memory allocation operation is completed, the model is repeatedly executed continuously again (S120).

[0046] If the probability of the sentence terminating is high, that is, if the probability (Pend) is greater than the threshold probability value (PThreshold) (S150-Y), all currently allocated memory is used to check whether the KV cache capacity will be insufficient for the next execution (S180).

[0047] If the allocated memory is almost completely used, additional memory allocation is performed (S160). On the other hand, if the allocated memory is almost completely used, there is a high probability that the KV cache capacity is sufficient (S180-N), so the language model operation is repeated without additional memory allocation (S120).

[0048] FIG. 4 is a diagram illustrating the configuration of a language model execution system according to another embodiment of the present invention. As illustrated, the language model execution system according to an embodiment of the present invention can be implemented as a computing system comprising a communication unit (210), an output unit (220), a processor (230), an input unit (240), and a storage unit (250).

[0049] The communication unit (210) is a communication interface for connection with an external network or external device, the output unit (220) is an output means for displaying the results of calculations performed by the processor (230), and the input unit (240) is a user interface for receiving user commands and transmitting them to the processor (230).

[0050] The processor (230) executes the language model according to the procedure illustrated in FIG. 1 described above, and in this process, allocates memory according to the procedure illustrated in FIG. 3. The storage unit (250) provides the storage space required for the language model to be executed by the processor (230).

[0051] So far, we have described in detail a preferred embodiment of a memory allocation method that can optimize the memory usage of a large language model.

[0052] The above example dynamically allocates memory for the key-value cache to optimize the key-value cache of a large language model. This allows for the operation of large language models with less memory than before, resulting in more energy-efficient operation compared to existing language models and improved multi-execution performance.

[0053] Meanwhile, it goes without saying that the technical idea of ​​the present invention can also be applied to a computer-readable recording medium containing a computer program that performs the functions of the device and method according to the present embodiment. In addition, the technical idea according to various embodiments of the present invention can be implemented in the form of computer-readable code recorded on a computer-readable recording medium. The computer-readable recording medium can be any data storage device that can be read by a computer and store data. For example, the computer-readable recording medium can be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical disk, a hard disk drive, etc. In addition, the computer-readable code or program stored on the computer-readable recording medium can be transmitted through a network connected between computers.

[0054] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.

Claims

1. A step for dividing the memory used by the language model in the KV (Key, Value) cache into chunk units; and A memory allocation method, characterized by including a step of dynamically allocating memory to a language model in stages in chunk units.

2. In claim 1, The allocation step is, Step to allocate 1 Chunk of memory for KV cache and initialize the counter to 0 when the language model is first run; A step of running the language model once and increasing the counter by 1; and A memory allocation method characterized by including a step of terminating a language model operation when an output token of a language model is a token indicating the end of sentence generation.

3. In claim 1, Memory of 1 Chunk size, A memory allocation method characterized by the size of memory required to store a KV cache for a specific number of Tokens.

4. In claim 3, The counter is, A memory allocation method characterized by the number of tokens generated in the current Chunk.

5. In claim 2, The allocation step is, If the output token of the language model is not a token indicating the end of sentence generation, a step of comparing the counter value with the counter threshold value; and A memory allocation method, characterized in that it further includes a step of operating the language model once and increasing the counter by 1 when the value of the counter is less than the counter threshold value.

6. In claim 5, The threshold is, A memory allocation method characterized by having a threshold value for the number of tokens for preparing the next memory allocation.

7. In claim 5, The allocation step is, If the value of the counter is greater than the threshold value, a step of comparing the threshold probability value with the probability value that the sentence will be terminated within the currently allocated KV cache without using additional KV cache; If the probability value is less than the critical probability value, a step of additionally allocating memory of 1 Chunk size for the KV cache; and A memory allocation method, characterized in that it further comprises the steps of initializing a counter and recalculating a counter threshold value.

8. In claim 7, The allocation step is, If the probability value is greater than the threshold probability value, a step for checking whether the KV cache capacity will be insufficient in the next execution; A memory allocation method, characterized in that it further includes a step of additionally allocating memory of the size of 1 Chunk for the KV cache if it is determined that the KV cache capacity is insufficient.

9. In claim 8, The allocation step is, A memory allocation method, characterized in that it further includes a step of operating a language model once and increasing a counter by 1 if it is confirmed that the KV cache capacity is not insufficient.

10. Memory where the storage space used by the language model in the KV (Key, Value) cache is divided into chunk units; A language model execution system, characterized by including a processor that dynamically allocates memory to a language model in stages in chunk units.

11. Step of running the language model once and increasing the counter by 1; A step of comparing the counter value and the counter threshold value when the output token of the language model is not a token indicating the end of sentence generation; A step of comparing the threshold probability value with the probability value that the sentence will be terminated within the currently allocated KV cache without additionally using the KV (Key, Value) cache if the value of the counter is greater than the threshold probability value; A memory allocation method, characterized by including a step of additionally allocating memory of the size of 1 Chunk for a KV cache when the probability value is less than a critical probability value.

12. Memory where the storage space used by the language model in the KV (Key, Value) cache is divided into chunk units; A language model execution system, characterized in that it includes a processor which operates a language model once and increases a counter by 1, compares the counter value with a counter threshold value if an output token of the language model is not a token indicating the end of sentence generation, and if the value of the counter is greater than the threshold value, compares a probability value that a sentence will end within the currently allocated KV cache without additionally using the KV (Key, Value) cache with a threshold probability value, and additionally allocates memory of the size of 1 Chunk for the KV cache if the probability value is less than the threshold probability value.

Citation Information

Patent Citations

  • KR20210123236A