Large model memory management method and device, electronic equipment and readable storage medium

By building a text length estimate model and pre-allocating memory blocks, the problem of low memory access efficiency in large model memory management is solved, the inference speed and throughput are improved, and the user experience is optimized.

CN120353603AActive Publication Date: 2025-07-22ZHEJIANG LAB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510821422.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-22
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The memory management method of large models has problems such as low memory access efficiency, large data handling overhead, and slow overall inference speed of large models, especially the bandwidth limitation between the computing unit and the storage unit causes serious memory wall problems.

Method used

Build a text length estimate model, and predict the text length of the large model output through the input layer, the embedding layer, the MQA-based decoder structure, the root mean square normalization layer, the linear projection layer and the output layer, and calculate the number of cache blocks based on the memory page size and the kv cache dimension, allocate the kv cache memory blocks in advance to avoid memory re-allocation and memory access discontinuity.

Benefits of technology

It improves the inference speed of large models, reduces the delay caused by dynamic adjustment, improves throughput in batch inference scenarios, and optimizes the user experience in streaming output scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353603A_ABST
    Figure CN120353603A_ABST
Patent Text Reader

Abstract

The invention discloses a memory management method and device for a large model, electronic equipment and a readable storage medium, and the method comprises the steps: inputting data into a trained text length prediction model, estimating the length of an output text of the large model, upwards adjusting the length into an integer, calculating the number of cache blocks according to the size of a memory page and the kv cache dimension, and storing the cache blocks in the memory page. The number of the cache blocks is adjusted upwards to be an integer; and finally, kv cache memory blocks are allocated for large model decoding. By allocating enough video memory or internal memory in advance, delay caused by dynamic adjustment is effectively avoided; in a batch reasoning scene, computing resources can be reasonably planned, and the throughput is improved; in a streaming output scene, in a word-by-word generation scene, the user experience can be optimized by estimating the output length, such as progress bar display or advanced truncation processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of large model data generation and inference, and particularly to a memory management method and device for a large model, an electronic device, and a readable storage medium. Background Art

[0002] With the rapid development of technology, large model inference technology, as one of the core technologies in the field of artificial intelligence, is showing broad application prospects. In actual application scenarios, optimizing the inference performance of large models faces various challenges, among which the memory access efficiency of computing units has become a key bottleneck restricting the overall system performance. Research shows that the inference speed of large models is not only limited by the theoretical computing power of computing units, but is also closely related to the memory access characteristics of the computing architecture. This correlation is mainly reflected in the following aspects: First, the large number of parameters in large models leads to complex memory access patterns; second, the data reuse characteristics in the neural network computing process pose special requirements on the storage hierarchy; third, the bandwidth limitation between computing units and storage units may cause serious "memory wall" problems. Therefore, existing large model memory management methods have problems such as low memory access efficiency, large data transfer overhead, and slow overall inference speed of large models. Summary of the Invention

[0003] Aiming at the deficiencies of the prior art, the present invention provides a memory management method and device for a large model, an electronic device, and a readable storage medium, which can avoid problems such as repeated memory reallocation and discontinuous memory access during large model decoding, and improve the inference speed of large models.

[0004] The object of the present invention is achieved by the following technical solutions:

[0005] A memory management method for a large model includes the following steps:

[0006] Step 1: Construct and train a text length prediction model;

[0007] The input of the text length prediction model is text data and relevant information of a large model, and the output is the length of the text output by the large model; the text length prediction model includes an input layer, an embedding layer, three cascaded MQA-based decoder structures, a third root mean square normalization layer, a linear projection layer, and an output layer connected in sequence; the MQA-based decoder structure includes a first root mean square normalization layer, a multi-query attention layer, a first splicing layer, a second root mean square normalization layer, a feed-forward neural network, and a second splicing layer; the multi-query attention layer shares some attention heads, that is, it shares Key and Value, which is used to capture context relationships and learn information from different subspaces; at the same time, different weights are assigned to different information, and the learned information and the corresponding weights are calculated, and the calculation result is used as a feature output; the first splicing layer is used to splice the input of the first root mean square normalization layer and the output of the multi-query attention layer; the feed-forward neural network is used to perform a non-linear mapping operation on the result normalized by the second root mean square normalization layer; the second splicing layer is used to perform the same splicing operation as the first splicing layer on the output of the feed-forward neural network and the output of the first splicing layer;

[0008] Step 2: Round up the predicted text length output by the text length prediction model in Step 1 to an integer;

[0009] Step 3: Calculate the cache block number according to the memory page size and the kv cache dimension;

[0010] Step 4: Round up the cache block number to an integer;

[0011] Step 5: Allocate kv cache memory blocks for the large model decoding.

[0012] Further, the predicted text length output by the output layer is an exact length or a quantization interval;

[0013] If the text length prediction model outputs an exact length prediction, the rounding up in Step 2 is specifically: first round up the length value, and then add an additional buffer value;

[0014] If the text length prediction model outputs a quantization interval, the rounding up in Step 2 is specifically: first round up the upper limit value of the quantization interval, and then add an additional buffer value.

[0015] Further, the rounding up is rounding up to the nearest unit, or rounding up to the nearest hundred, or rounding up to the nearest thousand, which is selected according to business requirements; the buffer value is a fixed value, or a percentage of the predicted value.

[0016] Further, in Step 4, one of the following methods is adopted to round up the cache block number to an integer:

[0017] Method 1: Round up to the nearest integer for the units digit;

[0018] Method 2: After rounding up to the nearest integer for the units digit, an additional safety block quantity is added;

[0019] Method 3: Memory alignment optimization, specifically:

[0020] N' = ceil(N / A) × A

[0021] where N' is the adjusted cache block quantity, N is the cache block quantity calculated in step three, A is the alignment coefficient, and ceil() is the rounding-up function;

[0022] Method 4: For the ultra-long text generation task, adopt a dynamic chunking strategy:

[0023] N' = min(ceil(N) + S, N_max)

[0024] where S represents the safety margin and N_max represents the maximum available block number of the device.

[0025] Furthermore, the specific content of step five is as follows: Use cudaMalloc(GPU) or posix_memalign(CPU) for aligned memory allocation, with the continuous space corresponding to the memory page size of each block as the allocation unit, and a total of N × D bytes of continuous memory space is allocated; The N' cache block quantities are arranged in sequence in the N × D bytes of continuous memory space, and each memory block contains:

[0026] Key buffer: Occupies 50% of the single block space, that is, D / 2 bytes;

[0027] Value buffer: Occupies 50% of the single block space, that is, D / 2 bytes;

[0028] Meta information header: An 8-byte block status identifier.

[0029] Furthermore, after completing the aligned memory allocation in step five, immediately execute the following memory guarantee strategy:

[0030] (1) Pre-clearing processing, immediately execute memset clearing after allocation to avoid numerical anomalies caused by uninitialized memory;

[0031] (2) Backup mechanism: Automatically switch to the fragmented allocation mode when continuous memory is insufficient.

[0032] Furthermore, the memory block allocation in step five is an asynchronous allocation and is executed in parallel with the calculation pipeline.

[0033] A memory management device for a large model, comprising:

[0034] An input data acquisition module, configured to acquire and combine input data;

[0035] A text length estimation and adjustment module, which estimates the text length of the output of the large model through a text length estimation model, and upward-adjusts the estimated text length value to an integer;

[0036] A memory block calculation and adjustment module, which calculates the number of memory blocks to be allocated by using the upward-adjusted text length and the memory page size, and at the same time upward-adjusts the number of memory blocks to an integer;

[0037] A kv cache memory block allocation module, configured to allocate kv cache memory blocks for the decoding of the large model.

[0038] An electronic device, comprising:

[0039] One or more processors;

[0040] A storage device, configured to store one or more programs, and when the one or more programs are executed by the electronic device, the electronic device implements the memory management method of the large model.

[0041] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the memory management method of the large model is implemented.

[0042] The beneficial effects of the present invention are as follows:

[0043] The present invention uses a text length estimation model to estimate the output text length of the large model, and then calculates the number of memory blocks to be allocated according to the estimated output text length. By allocating the memory required by the kv cache in the large model decoding at one time in this way, problems such as repeated allocation of memory and discontinuous memory access during the inference process are avoided, and the inference speed of the large model is improved. The method of the present invention can play different effects in different application scenarios. By allocating sufficient video memory or memory in advance, the delay caused by dynamic adjustment is effectively avoided; in the batch inference scenario, the computing resources are reasonably planned to improve the throughput; in the streaming output scenario, in the token-by-token scenario, predicting the output length can optimize the user experience, such as progress bar display or early truncation processing. Description of the Drawings

[0044] Figure 1 It is a schematic flowchart of the memory management method of the large model according to the embodiment of the present invention.

[0045] Figure 2 It is a schematic structural diagram of the text length prediction model according to the embodiment of the present invention.

[0046] Figure 3 Schematic diagram of the MQA-based decoder structure according to an embodiment of the present invention.

[0047] Figure 4 Schematic diagram of the memory management device of the large model according to an embodiment of the present invention.

[0048] Figure 5 Schematic diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners

[0049] The present invention will be described in detail below according to the accompanying drawings and preferred embodiments. The objectives and effects of the present invention will become more apparent. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0050] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0051] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0052] The present application provides a memory management method for a large model. Please refer to Figure 1 , as Figure 1 shown, the method includes the following steps:

[0053] Step 1: Construct and train a text length prediction model (Text Length Prediction Model).

[0054] The input of the text length prediction model is text data and relevant information of the large model, and the output is the length of the text output by the large model. Among them, the text data is the original text input by the user into the large model, such as questions, instructions, or natural language content to be processed, etc. The relevant information of the large model includes the model name (such as GPT-4, LLaMA-3, etc.), quantization bits (such as FP16, INT8, etc.), model size (the number of parameters, such as 7B, 70B, etc.), and adjustable parameters (such as temperature coefficient, top-p sampling value, etc.).

[0055] As Figure 2 shown, the text length prediction model includes an input layer (input), an embedding layer (embedding), three cascaded MQA-based decoder structures, a third Root Mean Square (RMS) normalization layer, a linear projection layer (Linear), and an output layer (Output) connected in sequence. Figure 3 The MQA-based decoder structure includes a first Root Mean Square normalization layer, a Multi Query Attention (MQA) layer, a first concatenation layer, a second Root Mean Square normalization layer, a Feed-Forward Network (FFN), and a second concatenation layer.

[0056] Among them, the input layer is used to receive input data, including the text data to be processed and relevant information of the large model (model name, quantization bits, model size, adjustable parameters), and concatenate the text data to be processed and the relevant information of the large model into a sequence.

[0057] The embedding layer is used to map the sequence concatenated by the input layer into a high-dimensional vector representation. In this embodiment, the embedding dimension of the embedding layer is 4096, and the vocabulary size is 32000.

[0058] In the three cascaded MQA-based decoder structures, the input of the first Root Mean Square normalization layer of the first MQA-based decoder structure is the high-dimensional vector representation mapped by the embedding layer; the input of the first Root Mean Square normalization layer of the latter two MQA-based decoder structures is the output of the second concatenation layer of the previous MQA-based decoder structure. In this embodiment, the number of attention heads in each MQA-based decoder structure is 32, the dimension of each head = 4096 / 32 = 128, the dimensions of the three matrices of Q, K, and V are all (4096, 128), and the output dimension of each MQA-based decoder structure is (4096, 4096).

[0059] The first Root Mean Square normalization layer is used to normalize the input to enhance the training stability of the model. The calculation formula for normalization is as follows:

[0060]

[0061] where d represents the dimension of the layer, which is 4096 here. x i represents the i-th input data, and ε is a parameter to prevent the RMS calculation result from being 0, ε = 1e-7.

[0062] The multi-query attention layer is used to capture context relationships and learn information from different subspaces. At the same time, different weights are assigned to different information, and the learned information and the corresponding weights are calculated, and the calculation result is used as the feature output. The specific calculation formula is as follows:

[0063]

[0064]

[0065] Where Q represents Query, K represents key, V represents Value, and d k represents the dimension of Query and Key, and W O is the output projection matrix. Each query head Q i in Attention(Q i , K, V) is an independently calculated attention weight, but all query heads share K and V. MQA(Q, K, V) represents the merging of the output results of all query heads, and Concat() is a string concatenation function.

[0066] The multi-query attention layer, as a variant of Self-Attention, reduces the computational overhead by sharing some attention heads (Key and Value), while maintaining good expressive power.

[0067] The first concatenation layer is used to concatenate the input of the first root mean square normalization layer and the output of the multi-query attention layer. The specific concatenation operation is to add the input of the first root mean square normalization and the output of the multi-query attention layer element-wise.

[0068] The second root mean square normalization layer is responsible for normalizing the output of the first concatenation layer.

[0069] The feed-forward neural network consists of fully connected layers and activation functions, and is used to perform non-linear mapping operations on the results normalized by the second root mean square normalization layer, so as to enhance the non-linear modeling ability of the model. In this embodiment, the feed-forward neural network contains three weight matrices, where the dimensions of two upsampling matrices are (4096, 16384), and the dimension of the downsampling matrix is (16384, 4096).

[0070] The second concatenation layer is used to perform the same concatenation operation as the first concatenation layer on the output of the feed-forward neural network and the output of the first concatenation layer. The output of the second concatenation layer is used as the output of the entire MQA-based decoder structure.

[0071] The third root mean square normalization layer is used to normalize the output of the third MQA-based decoder structure. The specific normalization operation is as follows:

[0072]

[0073]

[0074] Among them, RMSNorm is the final calculation result, where x represents the input vector, and γ represents the scaling factor, which is a trainable parameter.

[0075] The linear projection layer is used to map the output of the third root mean square normalization layer to the target dimension to prepare for the final prediction.

[0076] The output layer is used to output the predicted text length, which can be used as a basis for subsequent inference optimization (such as memory pre-allocation, batch scheduling, etc.). In this embodiment, the dimension of the output layer is 32000, which also corresponds to the size of the vocabulary. Here, the predicted text length can be represented in different ways, including precise length prediction (accurate to the unit digit, such as outputting "256"), and quantization interval prediction (such as in the order of 1000, outputting "≈2k" indicating that the predicted length is around 2000). Which specific method to use depends on the requirements of the application scenario. For example, precise prediction is suitable for scenarios that require strict control of memory occupancy (such as embedded devices), and quantization prediction is suitable for resource scheduling optimization (such as batch processing tasks in a cloud computing environment).

[0077] The MQA-based decoder structure in the text length prediction model of the present invention is a variant of the traditional transformer structure. Compared with the Self-Attention mechanism of the standard Transformer, this structure has more advantages in computational efficiency and is particularly suitable for lightweight prediction tasks. Three MQA-based decoder structures are used as the core data processing units in the text length prediction model of the present invention to gradually extract the deep features of the input data and improve the prediction accuracy.

[0078] When training the text length prediction model of the present invention, the actual output text length (such as the number of tokens or characters) is used as the supervision signal during training, that is, the expected output of the text length prediction model during training. The input data during training includes the input text of the large model (such as user queries, Prompts, etc.), and the relevant information of the large model (model name, quantization bits, model size, adjustable parameters). The input data during online prediction needs to have the same format and structure as the input data during training to ensure the same effect during training and online inference. During the training process, the text length prediction model learns the mapping relationship between the input data and the output length to optimize the prediction accuracy. After training, this model can be used in the inference stage to dynamically predict the text output length of the large model, and its prediction result can be used as a key input for subsequent optimization strategies (such as batch scheduling or video memory allocation), thereby improving the overall efficiency of large model inference.

[0079] Step 2: Round up the predicted text length output by the text length prediction model in Step 1 to an integer.

[0080] The predicted value output by the text length prediction model may have certain fluctuations or deviations. To ensure the stability and reliability of subsequent memory allocation, it is necessary to round up the predicted text length.

[0081] The specific method of rounding up is as follows:

[0082] If the text length prediction model outputs a length value accurate to the units digit (e.g., the predicted value is 1256), first round up this value (Ceiling Adjustment, e.g., round up to 1300), and then add an additional buffer value (e.g., +500, finally adjusted to 1,800); the buffer value can be adjusted according to the actual application scenario, and its purpose is to reserve additional memory space to prevent memory shortages caused by prediction deviations; adjustment for quantization interval prediction;

[0083] If the predicted text length output by the text length prediction model is a quantization interval, such as outputting "≈2k", indicating around 2000, that is, 1500–2500, then also round up, take the upper limit of 2500, and add a buffer value of 500 to avoid memory shortages in critical situations; after adjustment, it is 3000.

[0084] The adjustment strategy can effectively avoid the following problems:

[0085] (1) Memory Reallocation: During large model inference, if the initially allocated memory is insufficient, the system may be forced to dynamically expand the memory, resulting in additional computational overhead and latency; by rounding up the predicted value, sufficient memory can be allocated at once, reducing the performance loss caused by runtime adjustments.

[0086] (2) Preventing Out-of-Memory (OOM): If the predicted value is too low, it may lead to video memory / memory overflow, which in turn causes inference failure; increasing the buffer value can improve fault tolerance and ensure the stable operation of the large model when generating longer texts.

[0087] The buffer value (Buffer Size) can be flexibly selected according to the actual situation: A fixed value (e.g., +500) is suitable for most general scenarios and is computationally simple; dynamic adjustment (e.g., increasing by 20% of the predicted value) is suitable for tasks with large variations in output length (such as dialogue generation vs. summary generation).

[0088] Meanwhile, the rounding method can also be combined with business requirements to select different strategies such as rounding to the nearest hundred (1234 → 1300), rounding to the nearest thousand (1234 → 2000), etc. Through this optimization step, the system can manage memory resources more reliably and improve the stability and efficiency of large model inference. The adjusted length value will be used as a key parameter for subsequent memory pre-allocation to optimize the overall inference process.

[0089] Step 3: Calculate the number of cache chunks based on the memory page size and the dimension of the kv cache.

[0090] During the autoregressive large model inference process, the efficient management of the KV Cache (Key-Value cache) is crucial for the inference speed. To optimize the video memory utilization and reduce the overhead of dynamic memory allocation, it is necessary to reasonably calculate the number of chunks of the KV Cache according to the estimated text length. The specific calculation method is as follows:

[0091] N = W * L / D

[0092] Where, N represents the number of KV Cache blocks to be allocated; W represents the dimension of the KV Cache, which is determined by the model structure, such as each token occupying 128 bytes; L represents the adjusted estimated output text length; D represents the system memory page size, such as 4KB or 16KB per page, depending on the hardware architecture. W × L represents the total memory size required to store the KV Cache of the entire sequence, and dividing by D represents chunking according to the memory page size to calculate the number of physical storage units required.

[0093] Example calculation: Assume the adjusted output length L = 1,800 tokens, the KV Cache dimension W = 128 bytes / token, and the memory page size D = 4KB (4,096 bytes), then: N = 128 × 1800 / 4096 ≈ 56.25.

[0094] Step 4: Round up the number of cache chunks calculated in Step 3 to an integer.

[0095] After completing the initial calculation N = W × L / D, the obtained value is usually a floating point number (such as 56.25). Since memory allocation must be in integer blocks, the number of chunks needs to be adjusted. The quantity adjustment can be carried out in various ways and can be flexibly selected according to different application scenarios:

[0096] (1) Force rounding up, using the ceil(N) operation to ensure sufficient memory space is allocated. Example: When N = 56.25, 57 blocks are taken. This operation can completely avoid the risk of insufficient memory, but the cost is that there may be about 1.5% memory waste (taking 56.25 as an example).

[0097] (2) Minimum safety margin strategy, adding safety blocks additionally on the basis of rounding, adjustment formula: N' = ceil(N) + S (S is the number of safety blocks set according to experience, usually 1 - 3 blocks). Adding such safety blocks can provide memory safety guarantee for scenarios with zero tolerance for memory overflow.

[0098] (3) Memory alignment optimization, making alignment adjustment considering the characteristics of hardware memory access: calculating the number of aligned blocks: N' = ceil(N / A) × A (A is the architecture - specific alignment coefficient, such as a multiple of 4). This can effectively improve the memory access efficiency.

[0099] (4) For the ultra - long text generation task, adopt the dynamic chunking strategy: adopt the segmented allocation mechanism, set the initial allocated basic number of chunks (such as 32 chunks), and at the same time set the monitoring threshold to trigger dynamic expansion, and cooperate with cudaMallocAsync of CUDA to achieve non - stop expansion. The complete adjusted calculation formula:

[0100] N' = min(ceil(N) + S, N_max)

[0101] Among them, S represents the safety margin, with the default value of 1; N_max represents the maximum number of available blocks of the device.

[0102] Step Five: Allocate kv cache memory blocks for the large - model decoding.

[0103] According to the integerized number of KV Cache blocks N calculated in Step Four, the system performs video memory / memory allocation operations. The specific allocation method is as follows:

[0104] Use cudaMalloc(GPU) or posix_memalign(CPU) for aligned memory allocation; take the continuous space of each block strictly corresponding to D bytes (memory page size) as the allocation unit, and the total allocated continuous memory space is N' × D bytes.

[0105] Memory layout optimization, adopt the optimal memory arrangement scheme: [Block 0][Block 1]...[Block N - 1]. Each Block contains:

[0106] Key buffer: occupies 50% of the single - block space (D / 2 bytes);

[0107] Value buffer: occupies 50% of the single - block space (D / 2 bytes);

[0108] Meta - information header: 8 - byte block status identifier (including valid bits, etc.).

[0109] Meanwhile, some memory guarantee strategies are adopted according to different scenarios, such as:

[0110] Pre-clearing processing, immediately execute memset to clear after allocation, avoiding numerical anomalies caused by uninitialized memory;

[0111] Backup mechanism: automatically switch to the fragmented allocation mode when continuous memory is insufficient.

[0112] The following performance optimization measures should also be adopted for memory allocation:

[0113] Asynchronous allocation, executed in parallel with the computing pipeline to improve efficiency;

[0114] Unified memory management, supporting the Unified Virtual Memory (UVM) access mode;

[0115] Intelligent cache, automatically recycling idle blocks according to the LRU (Least Recently Used) strategy.

[0116] The memory area after allocation will be dedicated to storing the Key-Value matrix generated by each token during the decoding process.

[0117] This embodiment uses a text length prediction model to estimate the length of the text output by the large model, and then calculates the number of cache blocks to be allocated using the memory page size and the kv cache dimension, avoiding problems such as repeated memory reallocation and discontinuous memory access during the decoding of the large model. This embodiment can achieve different effects in different application scenarios. By pre-allocating sufficient video memory or memory in advance, it effectively avoids the latency caused by dynamic adjustment; in the batch inference scenario, it reasonably plans the computing resources and improves the throughput; in the streaming output scenario, in the token-by-token generation scenario, predicting the output length can optimize the user experience, such as progress bar display or early truncation processing.

[0118] On the other hand, as Figure 4 shown, this application also provides a memory management device for a large model, including:

[0119] An input data acquisition module, used to acquire and combine input data;

[0120] A text length estimation and adjustment module, estimating the length of the text output by the large model through a text length estimation model and rounding up the estimated text length value to an integer;

[0121] A memory block calculation and adjustment module, calculating the number of memory blocks to be allocated using the rounded-up text length and the memory page size, and at the same time rounding up the number of memory blocks;

[0122] The kv cache memory block allocation module is used to allocate kv cache memory blocks for the decoding of large models.

[0123] This application also provides an electronic device. Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an embodiment of the electronic device of this application. Except for Figure 5 the processors, memories, network interfaces, and non-volatile memories shown, any device with data processing capabilities where the device in the embodiment is located usually may further include other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated here.

[0124] The implementation processes of the functions and roles of each unit in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.

[0125] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0126] This embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the memory management method of the large model in the above embodiment.

[0127] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a SmartMedia card (SMC), an SD card, a Flash card, etc. equipped on the device. Further, the computer-readable storage medium may also include both the internal storage unit of any device with data processing capabilities and the external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store the data that has been output or will be output.

[0128] Other embodiments of the present application will be readily contemplated by those skilled in the art upon consideration of the specification and practice of the disclosure herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the claims.

[0129] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A memory management method for a large model, characterized in that It includes the following steps: Step 1: Construct and train a text length prediction model; The input of the text length prediction model is text data and relevant information of a large model, and the output is the length of the text output by the large model; the text length prediction model includes an input layer, an embedding layer, three cascaded MQA-based decoder structures, a third root mean square normalization layer, a linear projection layer, and an output layer connected in sequence; the MQA-based decoder structure includes a first root mean square normalization layer, a multi-query attention layer, a first splicing layer, a second root mean square normalization layer, a feed-forward neural network, and a second splicing layer; the multi-query attention layer shares some attention heads, that is, shares Key and Value, to capture context relationships and learn information from different subspaces; at the same time, different weights are assigned to different information, and the learned information and the corresponding weights are calculated, and the calculation result is used as a feature output; the first splicing layer is used to splice the input of the first root mean square normalization layer and the output of the multi-query attention layer; The feed-forward neural network is used to perform a non-linear mapping operation on the result normalized by the second root mean square normalization layer; the second splicing layer is used to perform the same splicing operation as the first splicing layer on the output of the feed-forward neural network and the output of the first splicing layer; Step 2: Round up the predicted text length output by the text length prediction model in Step 1 to an integer; Step 3: Calculate the cache block number according to the memory page size and the kv cache dimension; Step 4: Round up the cache block number to an integer; Step 5: Allocate kv cache memory blocks for the large model decoding.

2. The memory management method of the large model according to claim 1, wherein The predicted text length output by the output layer is an exact length or a quantization interval; If the text length prediction model outputs an exact length prediction, the rounding up in Step 2 is specifically: first round up the length value, and then add an additional buffer value; If the text length prediction model outputs a quantization interval, the rounding up in Step 2 is specifically: first round up the upper limit value of the quantization interval, and then add an additional buffer value.

3. The memory management method of the large model according to claim 2, wherein The rounding up is rounding up to the nearest unit or the nearest hundred or the nearest thousand, which is selected according to business requirements; the buffer value is a fixed value or a percentage of the predicted value.

4. The memory management method of the large model according to claim 3, wherein In Step 4, one of the following methods is used to round up the cache block number to an integer: Method 1: Round up to the nearest unit; Method 2: After rounding up to the nearest unit, add an additional safety block number; Method 3: Memory alignment optimization, specifically: N' = ceil(N / A)×A where N' is the adjusted cache block number, N is the cache block number calculated in Step 3, A is the alignment coefficient, and ceil() is the rounding-up function; Method 4: For ultra-long text generation tasks, adopt a dynamic block strategy: N' = min( ceil(N) + S, N_max ) where S represents the safety margin and N_max represents the maximum available blocks of the device.

5. The memory management method of the large model according to claim 4, characterized in that, Step five is specifically as follows: Use cudaMalloc or posix_memalign for aligned memory allocation, with the continuous space corresponding to the memory page size of each block as the allocation unit, and a total of N×D bytes of continuous memory space is allocated; The N' cache block numbers are arranged in sequence in the N×D bytes of continuous memory space, and each memory block contains: Key buffer: Occupies 50% of a single block, that is, D / 2 bytes; Value buffer: Occupies 50% of a single block, that is, D / 2 bytes; Meta-information header: An 8-byte block status flag.

6. The memory management method of the large model according to claim 1 or 5, characterized in that After completing the aligned memory allocation in step five, immediately execute the following memory guarantee policy: (1) Pre-clearing processing: Immediately execute memset for clearing after allocation to avoid numerical anomalies caused by uninitialized memory; (2) Backup mechanism: Automatically switch to the sharded allocation mode when continuous memory is insufficient.

7. The memory management method of the large model according to claim 1, characterized in that, The memory block allocation in step five is an asynchronous allocation and is executed in parallel with the computing pipeline.

8. A memory management device for a large model, characterized in that, It includes: An input data acquisition module for acquiring and combining input data; A text length estimation and adjustment module that estimates the text length output by the large model through a text length estimation model and rounds up the estimated text length value to an integer; A memory block calculation and adjustment module that calculates the number of memory blocks to be allocated using the rounded-up text length and the memory page size, and at the same time rounds up the number of memory blocks to an integer; A kv cache memory block allocation module for allocating kv cache memory blocks for the large model decoding.

9. An electronic device, characterized in that, It includes: One or more processors; A storage device for storing one or more programs, which, when executed by the electronic device, cause the electronic device to implement the memory management method of the large model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A program is stored thereon, which, when executed by the processor, implements the memory management method of the large model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Large language model reasoning optimization method and device, computer equipment and storage medium

    CN117194056A

  • Optimization method and system for dynamic reasoning memory allocation based on predictor

    CN118227336A

  • Large language model reasoning throughput testing method and device and program product

    CN120045896A

  • Lease cache memory devices and methods

    US20200310985A1

  • Efficiently serving machine-learned model computations with high throughput and low latency

    WO2025090955A1