Memory management method and device for reasoning system
By maintaining the scheduling queue and dynamically managing GPU memory in the inference engine, the problem of inefficient resource management of inference systems in the existing technology is solved, and more efficient GPU memory usage and more flexible resource management are achieved.
Patent Information
- Application Number
- CN202411784836.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-05
AI Technical Summary
Existing inference systems have problems with inefficiency in resource management, especially in the management of GPU memory by inference engines, resulting in waste of resources and degradation of performance.
By maintaining the Prefill scheduling queue and the Decode scheduling queue in the inference engine, the memory management time window is determined based on the data processing time, and the GPU memory requirements are calculated, and the GPU memory is dynamically allocated to achieve more flexible and efficient resource management.
It realizes that during the inference engine batch execution of inference requests, dynamically adjusts the GPU memory allocation to avoid resource waste, ensures that the inference engine has sufficient GPU memory usage, and improves the accuracy and adaptability of GPU memory management.
Smart Images

Figure CN119248525B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a memory management method and device for an inference system. Background Art
[0002] An inference system is a computer program that uses logical rules and known facts to draw new conclusions or decisions. An inference system is an important part of the field of artificial intelligence, mainly used to simulate the human decision-making process. It derives conclusions based on a set of defined knowledge bases and inference engines. An inference system can execute the inference request it obtains and output the corresponding inference results.
[0003] A typical reasoning system usually consists of the following parts: Knowledge Base, Inference Engine, User Interface, and Explanation Facility. The knowledge base includes all the facts and rules known to the storage system. These facts can be about the state of the world, object properties, etc., while the rules are logical expressions that describe how to draw new conclusions from known facts. The inference engine is the core component of the reasoning system. It is responsible for performing logical operations in the reasoning process, that is, drawing new conclusions or decisions from a given knowledge base; the inference engine uses a series of rules and known facts to derive new knowledge to help the system solve problems or make decisions. The user interface allows users to interact with the system, input queries, or observe the results of the reasoning process. The explanation mechanism is used to explain how the system draws specific conclusions, which is very important for transparency and trust.
[0004] Inference engines usually use computing resources, storage resources, network resources, etc. (e.g., GPU, GPU memory, storage, network interface, etc.) to perform inference tasks. The efficient use of these resources directly affects the performance of the inference engine. Therefore, it is expected to manage the resources used by the inference engine better and more flexibly. Summary of the invention
[0005] One or more embodiments of the present application provide the following technical solutions:
[0006] The present application provides a memory management method for an inference system, which is applied to an inference engine in the inference system; the computing resources of the inference engine include a GPU mounted on a computing device for deploying the inference engine; the inference engine maintains a Prefill scheduling queue for scheduling an inference request set in a Prefill stage, and a Decode scheduling queue for scheduling an inference request set in a Decode stage; the method includes:
[0007] Determine a Prefill memory management time window according to a data processing duration associated with a set of Prefill inference requests being executed in the Prefill scheduling queue, calculate a GPU memory requirement corresponding to the set of Prefill inference requests within the Prefill memory management time window, and allocate GPU memory to the set of Prefill inference requests according to the GPU memory requirement;
[0008] At the end of the Prefill memory management time window, re-determine a subsequent Prefill memory management time window corresponding to the Prefill memory management time window based on the data processing duration associated with the set of Prefill inference requests being executed in the Prefill scheduling queue;
[0009] Determine a Decode memory management time window according to a data processing duration associated with a Decode inference request set being executed in the Decode scheduling queue, calculate a GPU memory requirement corresponding to the Decode inference request set within the Decode memory management time window, and allocate GPU memory to the Decode inference request set according to the GPU memory requirement;
[0010] When the Decode memory management time window ends, a next Decode memory management time window corresponding to the Decode memory management time window is determined again based on the data processing duration associated with the Decode inference request set being executed in the Decode scheduling queue.
[0011] The present application also provides a memory management device for an inference system, which is applied to an inference engine in the inference system; the computing resources of the inference engine include a GPU mounted on a computing device for deploying the inference engine; the inference engine maintains a Prefill scheduling queue for scheduling an inference request set in a Prefill stage, and a Decode scheduling queue for scheduling an inference request set in a Decode stage; the device includes:
[0012] A first memory management module determines a Prefill memory management time window according to a data processing duration associated with a set of Prefill inference requests being executed in the Prefill scheduling queue, calculates a GPU memory requirement corresponding to the set of Prefill inference requests within the Prefill memory management time window, and allocates GPU memory to the set of Prefill inference requests according to the GPU memory requirement; at the end of the Prefill memory management time window, re-determines a next Prefill memory management time window corresponding to the Prefill memory management time window according to a data processing duration associated with a set of Prefill inference requests being executed in the Prefill scheduling queue;
[0013] a second memory management module, which determines a Decode memory management time window according to a data processing duration associated with a Decode inference request set being executed in the Decode scheduling queue, calculates a GPU memory requirement corresponding to the Decode inference request set within the Decode memory management time window, and allocates GPU memory to the Decode inference request set according to the GPU memory requirement; and when the Decode memory management time window ends, re-determines a next Decode memory management time window corresponding to the Decode memory management time window according to the data processing duration associated with the Decode inference request set being executed in the Decode scheduling queue.
[0014] The present application also provides an electronic device, comprising:
[0015] processor;
[0016] a memory for storing processor-executable instructions;
[0017] The processor implements the steps of any of the above methods by running the executable instructions.
[0018] The present application also provides a computer-readable storage medium having computer instructions stored thereon, which implement the steps of any of the above methods when executed by a processor.
[0019] In the above technical solution, the inference engine in the inference system can use the GPU mounted on the computing device where it is located as a computing resource to execute inference requests and manage GPU memory. Specifically, the inference engine can maintain a scheduling queue for scheduling inference request sets, and determine the memory management time window according to the data processing duration associated with the inference request set being executed in the scheduling queue, and subsequently calculate the GPU memory demand corresponding to the inference request set within the memory management time window, and allocate GPU memory to the inference request set according to the GPU memory demand, and after the end of the memory management time window, the next memory management time window corresponding to the memory management time window can be determined again according to the data processing duration associated with the inference request set being executed in the scheduling queue, so as to perform GPU memory management again within the next memory management time window.
[0020] By adopting the above method, on the one hand, there is no need to reserve a large amount of GPU memory for the inference engine when it is started. Instead, in the process of the inference engine executing inference requests in batches, a memory management time window can be continuously set, and the GPU memory demand corresponding to the inference request set being executed within the memory management time window can be predicted, so as to allocate GPU memory to the inference request set according to the GPU memory demand. This not only ensures that the inference engine has sufficient GPU memory available during the process of batch executing inference requests, but also avoids waste of GPU memory. On the other hand, the memory management time window can be set according to different stages of the inference process, and the GPU demand within the set memory management time window can be predicted, so as to realize staged GPU memory management in the inference process, thereby improving the accuracy and adaptability of GPU memory management. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The following is a description of the accompanying drawings required for describing the exemplary embodiments, wherein:
[0022] Figure 1 It is a schematic diagram of an inference system shown in an exemplary embodiment of the present application.
[0023] Figure 2 It is a schematic diagram of another reasoning system shown in an exemplary embodiment of the present application.
[0024] Figure 3 It is a flowchart of a memory management method of an inference system shown in an exemplary embodiment of the present application.
[0025] Figure 4 It is a structural schematic diagram of a device shown in an exemplary embodiment of the present application.
[0026] Figure 5It is a block diagram of a memory management device of an inference system shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0027] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description relates to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of the present application. On the contrary, they are only examples consistent with some aspects of one or more embodiments of the present application.
[0028] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this application. In some other embodiments, the steps included in the method may be more or less than those described in this application. In addition, a single step described in this application may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this application may also be combined into a single step for description in other embodiments.
[0029] In this application, the inference engine in the inference system can be deployed on a computing device, and the computing resources, storage resources, network resources, etc. (e.g., GPU, GPU memory, storage, network interface, etc.) of the computing device are used to perform calculations in the inference process and execute inference tasks. Alternatively, the inference engine in the inference system can be deployed on a computing cluster to meet the needs of processing large-scale data, high-concurrency requests, or high-performance computing, thereby achieving higher computing power, better fault tolerance, and more flexible resource management.
[0030] In general, the inference engine may use a CPU as a computing resource, where the CPU may be a CPU mounted on a computing device for deploying the inference engine.
[0031] With the development of artificial intelligence technology, reasoning systems based on large models (such as Large Language Model) are becoming more and more widely used.
[0032] Large models refer to machine learning models with a large number of parameters, such as various variants of the Transformer architecture, including but not limited to natural language processing models such as GPT and BERT. These models achieve powerful representation learning capabilities through large amounts of training data and complex architectures.
[0033] In a large model-based reasoning system, the large model can be viewed as a large knowledge base that contains information learned from a large amount of data. The large model learns a large number of patterns and features during training, which represent complex relationships in the data. Therefore, to some extent, the information stored inside the large model can be viewed as a form of knowledge.
[0034] The inference engine designed for large models is responsible for calling the large model to perform specific inference work, that is, it can manage the loading of large models, execute inference operations of large models, and manage the interaction with hardware (for example: CPU, GPU or other accelerators). The inference engine can also include optimization algorithms to improve the speed and efficiency of inference.
[0035] In practical applications, the large model can be deployed separately from the reasoning system, or it can be integrated into the reasoning system and the reasoning engine in the system can be used to call the large model to efficiently perform reasoning tasks.
[0036] Since large models are artificial intelligence applications that require a lot of computing resources, the inference engine usually uses GPU as a computing resource when calling large models to perform inference tasks.
[0037] GPUs are designed for parallel processing and can process multiple data points at the same time. This is particularly useful for deep learning models, as they often need to perform the same operation on large amounts of data. Modern AI models, especially neural network models, require a large number of floating-point operations. GPUs generally perform better than CPUs on floating-point operations, especially when processing large-scale matrix operations, which are common operations in deep learning. GPUs generally have higher memory bandwidth than CPUs, which means they are able to read and write data from memory faster. This is very important for applications that need to frequently access large amounts of data. Many GPU vendors (such as NVIDIA) have optimized their hardware and software stacks for machine learning tasks. For example, they provide hardware specifically designed to accelerate tensor operations, as well as programming models like CUDA to take full advantage of these hardware features. GPUs can increase inference speed. For applications deployed at large scale, using GPUs can reduce latency and increase system responsiveness.
[0038] During the operation of the inference engine, whether the computing resources, storage resources, network resources and other resources of the computing device can be efficiently utilized usually directly affects the performance of the inference engine.
[0039] For example, during the reasoning process of a large model, KV Cache (Key-Value Cache) is usually used to store the generated tokens.
[0040] It should be noted that the reasoning process of a large model refers to the process of using a trained large machine learning model to generate corresponding output results (i.e., reasoning results) based on given input data (i.e., data contained in the reasoning request). In the field of natural language processing, especially in large language models, the reasoning process usually involves generating a continuous text sequence based on a given input prompt, and generating a token for each iteration. The token here usually refers to a series of small units into which text is segmented in natural language processing, such as words or characters, depending on the word segmentation method of the model.
[0041] Specifically, a piece of text can be provided to the model as input, which is called a Prompt; the model receives the Prompt and converts it into a form that can be processed internally, usually by converting each Token in the Prompt into a corresponding digital representation, such as an embedding vector. An embedding vector is a numerical vector in a high-dimensional space that is used to represent the semantic information of each Token; based on the current Prompt, the model predicts the next most likely Token through complex calculations (such as multi-layer neural network operations); the model adds the predicted Token to the existing Prompt to form a new Prompt; the above steps can be repeated multiple times, and each repetition is called an iteration. In each iteration, the model predicts the next Token based on the latest Prompt until the set stop condition is reached, such as a certain number of Tokens are generated or a specific end mark is encountered.
[0042] KV Cache is a data storage method that is usually used to speed up data access. It is a form of data storage as key-value pairs, where the key is a unique identifier used to find data, and the value is the data associated with it or the information pointing to the location of the data. When an application requests specific data, the system can quickly retrieve the corresponding value through the key, thereby improving the system's response speed.
[0043] In large models, especially those used to perform natural language processing tasks, KV Cache is used to store previously calculated results for rapid reuse during decoding, thereby accelerating the inference process.
[0044] In actual applications, using KV Cache to store generated tokens consumes a lot of GPU memory. For an inference engine based on a large model, when the inference engine is started, a certain amount of GPU memory needs to be reserved for the inference engine based on past experience and actual conditions, for subsequent model inference in the inference engine. The allocation ratio is usually 80% to 90% of the total GPU memory. In addition, due to the uncertainty of the intermediate data generated during the inference process of the large model, a certain amount of GPU memory needs to be reserved for these intermediate data.
[0045] However, the traffic of the inference engine is not fixed, that is, the number of inference tasks that need to be performed by the inference engine may change over time. For example, more inference requests are sent to the inference engine in the current time period, while fewer inference requests are sent to the inference engine in the next time period. In this case, the GPU memory reserved for the inference engine will be idle when the traffic is small and cannot be used by other services; in addition, the remaining 10% to 20% of the GPU memory that is not reserved for the inference engine cannot be used by the inference engine. Therefore, the GPU memory of the computing device used to deploy the inference engine cannot be used efficiently, which affects the performance of the inference engine.
[0046] In the technical solution provided by one or more embodiments of the present application, the inference engine in the inference system can use the GPU mounted on the computing device where it is located as a computing resource to execute inference requests and manage GPU memory. Specifically, the inference engine can maintain a scheduling queue for scheduling inference request sets, and determine the memory management time window based on the data processing duration associated with the inference request set being executed in the scheduling queue, and subsequently calculate the GPU memory demand corresponding to the inference request set within the memory management time window, and allocate GPU memory to the inference request set based on the GPU memory demand, and after the memory management time window ends, the next memory management time window corresponding to the memory management time window can be determined based on the data processing duration associated with the inference request set being executed in the scheduling queue, so as to perform GPU memory management again within the next memory management time window.
[0047] By adopting the above method, on the one hand, there is no need to reserve a large amount of GPU memory for the inference engine when it is started. Instead, in the process of the inference engine executing inference requests in batches, a memory management time window can be continuously set, and the GPU memory demand corresponding to the inference request set being executed within the memory management time window can be predicted, so as to allocate GPU memory to the inference request set according to the GPU memory demand. This not only ensures that the inference engine has sufficient GPU memory available during the process of batch executing inference requests, but also avoids waste of GPU memory. On the other hand, the memory management time window can be set according to different stages of the inference process, and the GPU demand within the set memory management time window can be predicted, so as to realize staged GPU memory management in the inference process, thereby improving the accuracy and adaptability of GPU memory management.
[0048] Please refer to Figure 1 and Figure 2 , Figure 1 and Figure 2 They are respectively schematic diagrams of an inference system shown by an exemplary embodiment of the present application.
[0049] like Figure 1 As shown, the above reasoning system may include an API server (API Server) and a reasoning engine.
[0050] An API server is a server specifically designed to process client requests and return responses. It is usually an indispensable component of an application. It is responsible for processing client requests, executing business logic, generating responses, and ensuring the security, scalability, and performance of the system.
[0051] Specifically, the API server usually serves as the unified entrance of the reasoning system, handles all requests from clients, and can distribute requests to different backend services to achieve decoupling between services. The API server can be responsible for authentication and authorization to ensure that only legitimate requests can access system resources, and implement fine-grained permission control to protect sensitive data and operations. The API server can also work with the load balancer to evenly distribute requests to multiple backend instances to improve the availability and performance of the reasoning system.
[0052] The above-mentioned inference engine can be deployed on an independent computing device or on a computing cluster composed of at least one computing node. The computing device can be a physical or virtual local computing device or a physical or virtual cloud computing device.
[0053] A computing cluster is a system that is connected by a network and works together. These computing devices work together to complete computing tasks, providing stronger computing power and higher availability than a single computing device. Computing nodes are one of the core components of a computing cluster, usually referring to computing devices used to perform computing tasks, and the GPUs installed on them can be used as computing resources.
[0054] Correspondingly, the above-mentioned API server can be deployed on the computing device where the above-mentioned inference engine is located; or, the above-mentioned API server can be deployed on the head node in the above-mentioned computing cluster, where the head node usually refers to the computing node responsible for the management and scheduling of the computing cluster, and can also be deployed on other computing devices outside the computing cluster.
[0055] It should be noted that the reasoning process may include multiple stages in some cases. For example, for a reasoning system based on a large model, the reasoning process of the large model (especially the generative model in natural language processing) may include the Prefill stage and the Decode stage.
[0056] In practical applications, the two stages of Prefill and Decode are closely linked. The Prefill stage provides the initial text basis, while the Decode stage is responsible for gradually expanding on this basis to generate a complete text output. These two stages together determine the quality and coherence of the final generated text.
[0057] Based on the above, if Figure 2 As shown, for an independent computing device used to deploy the above-mentioned inference engine, the resources of the computing device (for example, computing resources, storage resources, network resources, etc.) can be allocated as needed, so that a part of these resources are used to perform inference calculations in the Prefill stage, and the inference engine using this part of resources can be called a Prefill engine; another part of these resources are used to perform inference calculations in the Decode stage, and the inference engine using this part of resources can be called a Decode engine.
[0058] For the computing cluster used to deploy the above-mentioned inference engine, the computing nodes in the computing cluster can be divided into two categories, the first category is the computing nodes used to perform the inference calculation of the Prefill stage, and the second category is the computing nodes used to perform the inference calculation of the Decode stage. In this case, the above-mentioned inference engine can be divided into a Prefill engine and a Decode engine according to the inference calculation performed; wherein the Prefill engine is the inference engine deployed on the first category of computing nodes, and the Decode engine is the inference engine deployed on the second category of computing nodes. Alternatively, the resources (for example: computing resources, storage resources, network resources, etc.) of the computing nodes in the computing cluster can also be allocated as needed, so that part of these resources are used to perform the inference calculation of the Prefill stage, and the inference engine using this part of resources can be called the Prefill engine; another part of these resources is used to perform the inference calculation of the Decode stage, and the inference engine using this part of resources can be called the Decode engine.
[0059] Furthermore, at least one computing instance (which may be referred to as a first computing instance) may be deployed on a computing device or some of its resources used to perform reasoning calculations in the Prefill phase, and these first computing instances together constitute the above-mentioned Prefill engine. Similarly, at least one computing instance (which may be referred to as a second computing instance) may be deployed on a computing device or some of its resources used to perform reasoning calculations in the Decode phase, and these second computing instances together constitute the above-mentioned Decode engine. Among them, a computing instance refers to a dedicated resource allocated on demand through virtualization technology in a local data center or cloud computing platform. The resources of these computing instances may include computing resources, storage resources, network resources, etc. (for example: GPU, GPU memory, storage, network interface, etc.), so that each computing instance can use these resources to perform specific computing tasks. Computing instances can be created, started, stopped or deleted according to actual needs, and their resources can also be adjusted according to demand.
[0060] The computing resource of a computing instance may be a GPU of the computing device on which it is located. Specifically, a computing device may be equipped with multiple GPUs, and a GPU may be allocated to a computing instance as a computing resource of the computing instance. Alternatively, a computing device may be equipped with only one GPU, and the computing resource allocated to a computing instance may be a part of the computing unit, memory, etc. of the GPU.
[0061] In the inference engine based on the big model, each computing instance can call the big model to perform specific inference work, that is, the big model can be loaded into each computing instance, and each computing instance performs the inference operation of the big model. Specifically, the big model can be read from a local or cloud storage medium (e.g., local disk, cloud storage, etc.) to the GPU memory of each computing instance to perform inference based on the big model on the GPU resources of each computing instance.
[0062] In actual applications, the Model Runtime framework can be installed in the computing instance, and after the large model is read into the computing instance, the model service is configured according to the Model Runtime framework and the large model, so that the computing instance can be used as an independently run model service instance. Among them, the model runtime refers to the environment and framework used to perform model reasoning after the model is deployed. This environment usually includes a series of steps such as model loading, input data processing, model execution, and output result processing, covering the entire life cycle of the model from loading to execution and then unloading.
[0063] For the above-mentioned inference engine including at least one computing instance, a load-aware scheduling method may be adopted to implement scheduling of inference requests on the inference engine.
[0064] Load-aware scheduling is a strategy for allocating tasks in a computing system based on the current load of each node or resource. This scheduling method aims to optimize resource usage efficiency, reduce waiting time and processing delays, and improve the overall throughput of the system.
[0065] Load-aware scheduling is usually used in distributed systems, cloud computing platforms, data centers and other scenarios. Its purpose is to dynamically balance the workload of each node and avoid the situation where some nodes are overloaded while other nodes are idle. By monitoring the load of each node in real time, the load-aware scheduler can make more reasonable task allocation decisions.
[0066] For example, in the above-mentioned inference system, the above-mentioned API server can communicate with the above-mentioned first computing instance in the above-mentioned inference engine. In this case, the API server can obtain the inference request from the request queue, and select the first computing instance with a suitable GPU load according to the GPU load of each first computing instance, and schedule the obtained inference request to the first computing instance, so as to use the GPU resources of the first computing instance to perform the inference calculation of the Prefill stage for the inference request.
[0067] The first computing instance in the inference engine can communicate with the second computing instance. In this case, after the first computing instance completes the inference calculation of the Prefill stage for the inference request, it can select a second computing instance with a suitable GPU load according to the GPU load of each second computing instance, and schedule the obtained inference calculation result of the Prefill stage corresponding to the inference request to the second computing instance, so as to use the GPU resources of the second computing instance to perform the inference calculation of the Decode stage for the inference request, for example: continue the inference calculation of the Decode stage based on the inference calculation result of the Prefill stage.
[0068] By deploying computing instances on computing devices, on the one hand, the resources of the computing devices can be fully utilized to accelerate the execution of computing tasks, thereby improving the computing power of the inference engine; on the other hand, different computing instances are independent of each other, which can ensure the stability and security of computing tasks, and can achieve resource isolation and management, which is conducive to the effective use of resources and avoid resource requisition; on the third hand, for computing tasks that need to be processed in parallel, parallel computing can be used to speed up the completion time of computing tasks; on the fourth hand, elastic scaling can be achieved. When a large number of computing tasks need to be processed, computing instances can be added, and when the number of computing tasks that need to be processed decreases, some computing instances can be terminated to release some resources and realize flexible allocation of resources.
[0069] For the above-mentioned API server, the API server may include two components, namely, a request router component responsible for routing inference requests to forward the inference requests to the above-mentioned first computing instance in the above-mentioned inference engine, and a monitor component responsible for monitoring the operation-related information of all computing instances in the inference engine (for example: GPU memory utilization, GPU memory bandwidth utilization, the number of inference requests waiting in the scheduling queue and the number of running inference requests, etc.).
[0070] For the computing device or some of its resources used to deploy the Prefill engine part included in the above-mentioned inference engine, in order to facilitate the maintenance, management and scheduling of the above-mentioned first computing instance deployed on the computing device or some of its resources, a Prefill Agent can also be deployed on the computing device or some of its resources.
[0071] The API server can communicate with the Prefill agent. In this case, the API server can obtain the inference request from the request queue, and select the first computing instance with the appropriate GPU load according to the GPU load of each first computing instance, and send the obtained inference request to the Prefill agent deployed on the computing device where the first computing instance is located, and the Prefill agent further schedules the inference request to the first computing instance to use the GPU resources of the first computing instance to perform the inference calculation of the Prefill stage for the inference request.
[0072] For the computing device or some of its resources used to deploy the Decode engine part included in the above-mentioned inference engine, in order to facilitate the maintenance, management and scheduling of the above-mentioned second computing instance deployed on the computing device or some of its resources, a Decode Agent can also be deployed on the computing device or some of its resources.
[0073] The Prefill agent can communicate with the Decode agent. In this case, after the first computing instance managed by the Prefill agent completes the inference calculation of the Prefill stage for the inference request, it can select a second computing instance with a suitable GPU load according to the GPU load of each second computing instance, and send the obtained inference calculation result of the Prefill stage corresponding to the inference request to the Decode agent deployed on the computing device where the second computing instance is located, and the Decode agent further schedules the inference request to the second computing instance to use the GPU resources of the second computing instance to perform the inference calculation of the Decode stage for the inference request, for example: continue the inference calculation of the Decode stage based on the inference calculation result of the Prefill stage.
[0074] In practical applications, in order to ensure the reliability and correctness of the inference request scheduling, the above-mentioned agent (including the above-mentioned Prefill agent and the above-mentioned Decode agent) can maintain a scheduling queue (Schedule Queue).
[0075] A scheduling queue is a special type of task queue, which is mainly used to manage and schedule tasks with specific execution time or priority (in this application, it is the inference task specified by the inference request). The main functions of the scheduling queue include task scheduling, task management, and resource optimization. Among them, task scheduling means that the scheduling queue can ensure that the task is executed at a specified time point or time range, and support the scheduling of periodic tasks; and the priority of the task can be set to ensure that the high-priority tasks in the scheduling queue are executed first. Task management refers to the management of all tasks in one queue for easy monitoring and control, and the execution status of the task can be tracked, including unexecuted, executing, completed, and failed states. Resource optimization means that by reasonably arranging the execution time of tasks, the system can be prevented from being overloaded at a certain moment, and system resources can be ensured to be effectively utilized to avoid resource waste.
[0076] Through the above scheduling queue and the above computing instance, the above agent has queue and batch processing capabilities. That is, each inference request in an inference request set (also called a batch of inference requests) can be scheduled to each computing instance through the scheduling queue, so that each inference request is executed by each computing instance, which enables different inference requests in the inference request set to be executed in parallel.
[0077] In addition, the above agent can also contain two components, namely, the memory predictor and the dynamic memory management component. The memory predictor can continuously execute the memory consumption prediction algorithm and dynamically manage the GPU memory in combination with the dynamic memory management component.
[0078] For each computing instance, the computing instance may contain two components: a KV Cache component used to cache intermediate data and results of reasoning to avoid repeated calculations, and an Executor component responsible for calling the large model to execute reasoning tasks corresponding to the reasoning request.
[0079] Please Figure 1 and Figure 2 Based on the reference Figure 3 , Figure 3 It is a flowchart of a memory management method of an inference system shown in an exemplary embodiment of the present application.
[0080] In this embodiment, the memory management method of the above reasoning system can be applied to Figure 1 The inference engine shown.
[0081] It should be noted that in order to improve the accuracy and adaptability of GPU memory management, GPU memory management can be performed according to the execution stage of each inference request in the inference request set being executed in the scheduling queue, thereby realizing phased GPU memory management during the inference process.
[0082] As mentioned above, for the reasoning system based on the large model, the reasoning process of the large model may include the Prefill stage and the Decode stage. In the reasoning process of the large model, the Prefill stage and the Decode stage refer to different steps when generating text.
[0083] The prefill stage usually occurs at the initial stage of generating text, and its main purpose is to provide an initial context for the model so that the model can better understand and generate subsequent text. In the prefill stage, the model usually generates an initial text fragment that contains the starting part of the generated sequence. This stage may involve using some predefined strategies or algorithms to select the most appropriate starting vocabulary or phrases. For example, when dealing with a conditional generation task, the model may first generate a part of the text as a basis based on the input conditions (such as summary generation, question and answer, dialogue, etc.).
[0084] The Decode stage is the main stage of generating text. The model gradually generates new words or characters in this stage until a complete output sequence is generated. In the Decode stage, the model generates new content word by word or word by word based on the existing text. In each iteration, the model predicts the next most likely word and adds it to the current text sequence. This process continues until the preset end condition is reached (for example, reaching the maximum text length, generating a text end marker, etc.). Various strategies may be used in the Decode stage to optimize the quality of generation, such as Top-k sampling, Top-p (also known as Nucleus Sampling) sampling, and other techniques, which can help the model avoid generating content that is too mundane or meaningless.
[0085] like Figure 3 As shown, the memory management method of the above reasoning system may include the following steps:
[0086] Step 302: Determine a Prefill memory management time window based on the data processing duration associated with the Prefill inference request set being executed in the Prefill scheduling queue, calculate the GPU memory requirement corresponding to the Prefill inference request set within the Prefill memory management time window, and allocate GPU memory to the Prefill inference request set based on the GPU memory requirement.
[0087] In this embodiment, after each Prefill inference request in an inference request set (which may be referred to as a Prefill inference request set) is scheduled to each first computing instance for execution through the above-mentioned Prefill scheduling queue, the execution status of each Prefill inference request in the Prefill inference request set maintained by the Prefill scheduling queue can be updated from an unexecuted state to an executing state. For the Prefill inference request set being executed in the Prefill scheduling queue, the Prefill memory management time window can be determined based on the data processing duration associated with the Prefill inference request set. Subsequently, the Prefill memory management time window can be used as the time unit of the GPU memory management task in the Prefill stage, that is, the GPU memory management of the Prefill stage is performed once within the Prefill memory management time window.
[0088] In some embodiments, when determining the Prefill memory management time window based on the data processing duration associated with the Prefill inference request set being executed in the above-mentioned Prefill scheduling queue, the number of tokens contained in each Prefill inference request in the Prefill inference request set can be specifically compared to determine the largest number therein, that is, the maximum number of tokens corresponding to the Prefill inference requests in the Prefill inference request set, so that the token processing duration corresponding to the maximum number of tokens can be determined as the length of the above-mentioned Prefill memory management time window.
[0089] In this embodiment, when the Prefill memory time window is determined, the GPU memory requirement corresponding to the Prefill reasoning request set within the Prefill memory management time window can be calculated, so that GPU memory can be allocated to the Prefill reasoning request set according to the GPU memory requirement. In this way, since GPU memory can be allocated to the inference engine according to the actual GPU memory requirement during the process of the inference engine executing the reasoning request in the reasoning request set, it can not only ensure that the inference engine has enough GPU memory to use during the process of executing the reasoning request, but also avoid the waste of GPU memory.
[0090] In practical applications, the memory predictor in the above-mentioned Prefill agent can execute the memory consumption prediction algorithm, that is, determine the Prefill memory management time window according to the data processing duration associated with the Prefill reasoning request set being executed in the above-mentioned Prefill scheduling queue, and calculate the GPU memory demand corresponding to the Prefill reasoning request set within the Prefill memory management time window, and the dynamic memory management component in the Prefill agent performs specific GPU memory allocation, that is, allocates GPU memory to the Prefill reasoning request set according to the GPU memory demand.
[0091] It should be noted that the inference request in the Prefill stage is a computationally intensive task, and GPU memory is not a bottleneck. Therefore, GPU memory can be directly allocated to the above Prefill inference request set based on the predicted GPU memory demand.
[0092] Step 304: When the Prefill memory management time window ends, a next Prefill memory management time window corresponding to the Prefill memory management time window is determined again based on the data processing duration associated with the set of Prefill inference requests being executed in the Prefill scheduling queue.
[0093] In this embodiment, after the GPU memory management execution in the above-mentioned Prefill memory management time window ends, the end of the Prefill memory management time window can be waited. When the Prefill memory management time window ends, the next Prefill memory management time window corresponding to the Prefill memory management time window can be determined again based on the data processing duration associated with the Prefill reasoning request set being executed in the above-mentioned Prefill scheduling queue.
[0094] Specifically, assuming that the above-mentioned Prefill memory management time window is the Tth Prefill memory management time window determined according to the data processing duration associated with the Prefill reasoning request set being executed in the above-mentioned Prefill scheduling queue, the GPU memory demand corresponding to the Prefill reasoning request set in the Tth Prefill memory management time window can be calculated, and GPU memory can be allocated to the Prefill reasoning request set based on the GPU memory demand; at the end of the Tth Prefill memory management time window, the T+1th Prefill memory management time window can be re-determined based on the data processing duration associated with the Prefill reasoning request set being executed in the Prefill scheduling queue, so that the GPU memory demand corresponding to the Prefill reasoning request set in the T+1th Prefill memory management time window can be calculated, and GPU memory can be allocated to the Prefill reasoning request set based on the GPU memory demand; and so on.
[0095] In some embodiments, when calculating the GPU memory demand corresponding to the above-mentioned Prefill reasoning request set within the above-mentioned Prefill memory management time window, specifically, the static GPU memory usage corresponding to the Prefill reasoning request set within the Prefill memory management time window can be calculated first, and then the maximum GPU memory usage in the previous Prefill memory management time window corresponding to the Prefill memory management time window and the sum of the static GPU memory usage corresponding to the Prefill reasoning request set in each target Prefill memory management time window are determined, and the difference between the maximum GPU memory usage and the sum of the static GPU memory usage is determined as the dynamic GPU memory usage corresponding to the Prefill reasoning request set within the Prefill memory management time window, wherein the target Prefill memory management window is the Prefill memory management window located before the Prefill memory management time window during the execution of the Prefill reasoning request set, and finally the sum of the static GPU memory usage and the dynamic GPU memory usage is determined as the GPU memory demand corresponding to the Prefill reasoning request set within the Prefill memory management time window.
[0096] Specifically, assuming that the above-mentioned Prefill memory management time window is the Tth Prefill memory management time window determined according to the data processing duration associated with the Prefill inference request set being executed in the above-mentioned Prefill scheduling queue, the previous Prefill memory management time window corresponding to the Tth Prefill memory management time window is the T-1th Prefill memory management time window, and the target Prefill memory management window located before the Tth Prefill memory management time window during the execution of the Prefill inference request set is the 1st to T-1st Prefill memory management time windows. In this case, PrefillStaticResource[T] represents the static GPU memory usage within the Tth Prefill memory management time window, PrefillDynamicResource[T] represents the dynamic GPU memory usage within the Tth Prefill memory management time window, PrefillMaxResource[T] represents the maximum GPU memory usage within the Tth Prefill memory management time window, and PrefillNeedResource[T] represents the GPU memory demand within the Tth Prefill memory management time window, then the following formula is obtained:
[0097] PrefillDynamicResource[T] = PrefillMaxResource[T-1]–(PrefillStaticResource[0] + PrefillStaticResource[1]+ … +PrefillStaticResource[T-1]);
[0098] PrefillNeedResource[T] = PrefillStaticResource[T]+PrefillDynamicResource[T].
[0099] That is, for the current Prefill memory management time window mentioned above, the dynamic GPU memory usage within the Prefill memory management time window can be predicted based on the maximum GPU memory usage in the previous Prefill memory management time window corresponding to the Prefill memory management time window, and the sum of the static GPU memory usage in all Prefill memory management time windows before the Prefill memory management time window.
[0100] It should be noted that the above-mentioned static GPU memory usage refers to the GPU memory that is foreseeable to be used during the inference process, such as: the GPU memory usage required for the tokens contained in the Prompt, the GPU memory usage required for new tokens generated based on the Prompt, etc.; the above-mentioned dynamic GPU memory usage refers to the GPU memory that is unforeseeable to be used during the inference process, such as: the GPU memory usage required for the uncertain intermediate data generated during the inference process of a large model.
[0101] In practical applications, in order to improve the compatibility of the technical solution in this application, when the above-mentioned inference engine is started, a sufficient amount of virtual GPU memory can be reserved for the inference engine, and the above-mentioned memory management method is used in the subsequent inference process to allocate physical GPU memory by calling the API provided by CUDA (Compute Unified Device Architecture).
[0102] In some embodiments, when calculating the static GPU memory usage corresponding to the above-mentioned Prefill inference request set within the above-mentioned Prefill memory management time window, the GPU memory usage required for the token corresponding to each Prefill inference request in the above-mentioned Prefill inference request set (which may be referred to as the first GPU memory usage) and the GPU memory usage required to generate a token based on each Prefill inference request in the Prefill inference request set (which may be referred to as the second GPU memory usage) may be specifically determined, and the sum of the first GPU memory usage and the second GPU memory usage may be determined as the static GPU memory usage corresponding to the Prefill inference request set within the above-mentioned Prefill memory management time window. The GPU memory usage required for the token corresponding to a Prefill inference request may be the GPU memory usage required for the token included in the Prompt in the Prefill inference request; the GPU memory usage required to generate a token based on a Prefill inference request may be the GPU memory usage required to generate a token based on the Prompt included in the Prefill inference request.
[0103] Step 306: Determine a Decode memory management time window based on the data processing duration associated with the Decode inference request set being executed in the Decode scheduling queue, calculate the GPU memory requirement corresponding to the Decode inference request set within the Decode memory management time window, and allocate GPU memory to the Decode inference request set based on the GPU memory requirement.
[0104] In this embodiment, after each Decode inference request in an inference request set (which may be referred to as a Decode inference request set) is scheduled to each second computing instance for execution through the above-mentioned Decode scheduling queue, the execution status of each Decode inference request in the Decode inference request set maintained by the Decode scheduling queue can be updated from an unexecuted state to an executing state. For the Decode inference request set being executed in the Decode scheduling queue, the Decode memory management time window can be determined based on the data processing duration associated with the Decode inference request set. Subsequently, the Decode memory management time window can be used as the time unit of the GPU memory management task in the Decode stage, that is, the GPU memory management of the Decode stage is performed once within the Decode memory management time window.
[0105] In some embodiments, when determining the Decode memory management time window based on the data processing duration associated with the Decode inference request set being executed in the above-mentioned Decode scheduling queue, the token generation durations corresponding to each Decode inference request in the Decode inference request set can be specifically compared, where the token generation duration is the duration for generating a token based on the Prompt contained in the Decode inference request, so as to determine the longest token generation duration corresponding to the Decode inference request in the Decode inference request set, so that the longest token generation duration can be determined as the length of the above-mentioned Decode memory management time window.
[0106] In this embodiment, when the above-mentioned Decode memory time window is determined, the GPU memory requirement corresponding to the above-mentioned Decode reasoning request set within the Decode memory management time window can be calculated, so that the GPU memory can be allocated to the Decode reasoning request set according to the GPU memory requirement. In this way, since the GPU memory can be allocated to the reasoning engine according to the actual GPU memory requirement during the process of the reasoning engine executing the reasoning request in the reasoning request set, it can not only ensure that the reasoning engine has enough GPU memory to use during the process of executing the reasoning request, but also avoid the waste of GPU memory.
[0107] In practical applications, the memory predictor in the above-mentioned Decode agent can execute the memory consumption prediction algorithm, that is, determine the Decode memory management time window according to the data processing duration associated with the Decode inference request set being executed in the above-mentioned Decode scheduling queue, and calculate the GPU memory demand corresponding to the Decode inference request set within the Decode memory management time window, and the dynamic memory management component in the Decode agent performs specific GPU memory allocation, that is, allocates GPU memory to the Decode inference request set according to the GPU memory demand.
[0108] Step 308: When the Decode memory management time window ends, determine the next Decode memory management time window corresponding to the Decode memory management time window based on the data processing duration associated with the Decode inference request set being executed in the Decode scheduling queue.
[0109] In this embodiment, after the GPU memory management execution in the above-mentioned Decode memory management time window is completed, the end of the Decode memory management time window can be waited. When the Decode memory management time window ends, the next Decode memory management time window corresponding to the Decode memory management time window can be determined again based on the data processing duration associated with the Decode inference request set being executed in the above-mentioned Decode scheduling queue.
[0110] Specifically, assuming that the above-mentioned Decode memory management time window is the Tth Decode memory management time window determined according to the data processing duration associated with the Decode inference request set being executed in the above-mentioned Decode scheduling queue, the GPU memory requirement corresponding to the Decode inference request set in the Tth Decode memory management time window can be calculated, and GPU memory can be allocated to the Decode inference request set according to the GPU memory requirement; at the end of the Tth Decode memory management time window, the T+1th Decode memory management time window can be re-determined according to the data processing duration associated with the Decode inference request set being executed in the Decode scheduling queue, so that the GPU memory requirement corresponding to the Decode inference request set in the T+1th Decode memory management time window can be calculated, and GPU memory can be allocated to the Decode inference request set according to the GPU memory requirement; and so on.
[0111] In some embodiments, when calculating the GPU memory demand corresponding to the above-mentioned Decode reasoning request set within the above-mentioned Decode memory management time window, specifically, the static GPU memory usage corresponding to the Decode reasoning request set within the Decode memory management time window can be calculated first, and then the maximum GPU memory usage in the previous Decode memory management time window corresponding to the Decode memory management time window and the sum of the static GPU memory usage corresponding to the Decode reasoning request set in each target Decode memory management time window are determined, and the difference between the maximum GPU memory usage and the sum of the static GPU memory usage is determined as the dynamic GPU memory usage corresponding to the Decode reasoning request set within the Decode memory management time window, wherein the target Decode memory management window is the Decode memory management window located before the Decode memory management time window during the execution of the Decode reasoning request set, and finally the sum of the static GPU memory usage and the dynamic GPU memory usage is determined as the GPU memory demand corresponding to the Decode reasoning request set within the Decode memory management time window.
[0112] Specifically, assuming that the above-mentioned Decode memory management time window is the Tth Decode memory management time window determined according to the data processing duration associated with the Decode inference request set being executed in the above-mentioned Decode scheduling queue, the previous Decode memory management time window corresponding to the Tth Decode memory management time window is the T-1th Decode memory management time window, and the target Decode memory management window located before the Tth Decode memory management time window during the execution of the Decode inference request set is the 1st to T-1st Decode memory management time windows. In this case, DecodeStaticResource[T] represents the static GPU memory usage within the Tth Decode memory management time window, DecodeDynamicResource[T] represents the dynamic GPU memory usage within the Tth Decode memory management time window, DecodeMaxResource[T] represents the maximum GPU memory usage within the Tth Decode memory management time window, and DecodeNeedResource[T] represents the GPU memory demand within the Tth Decode memory management time window, then the following formula is obtained:
[0113] DecodeDynamicResource[T] = DecodeMaxResource[T-1]–(DecodeStaticResource[0] + DecodeStaticResource[1]+ … + DecodeStaticResource[T-1]);
[0114] DecodeNeedResource[T] = DecodeStaticResource[T]+DecodeDynamicResource[T].
[0115] That is, for the current Decode memory management time window mentioned above, the dynamic GPU memory usage within the Decode memory management time window can be predicted based on the maximum GPU memory usage in the previous Decode memory management time window corresponding to the Decode memory management time window, and the sum of the static GPU memory usage in all Decode memory management time windows before the Decode memory management time window.
[0116] It should be noted that the above-mentioned static GPU memory usage refers to the GPU memory that is foreseeable to be used during the inference process, such as: the GPU memory usage required for the tokens contained in the Prompt, the GPU memory usage required for new tokens generated based on the Prompt, etc.; the above-mentioned dynamic GPU memory usage refers to the GPU memory that is unforeseeable to be used during the inference process, such as: the GPU memory usage required for the uncertain intermediate data generated during the inference process of a large model.
[0117] In practical applications, in order to improve the compatibility of the technical solution in this application, when the above-mentioned inference engine is started, a sufficient amount of virtual GPU memory can be reserved for the inference engine, and the above-mentioned memory management method is used in the subsequent inference process to allocate physical GPU memory by calling the API provided by CUDA (Compute Unified Device Architecture).
[0118] In some embodiments, when calculating the static GPU memory usage corresponding to the Decode inference request set within the Decode memory management time window, the GPU memory usage required to generate a token based on each Decode inference request in the Decode inference request set can be specifically determined, and the GPU memory usage is determined as the static GPU memory usage corresponding to the Decode inference request set within the Decode memory management time window. The GPU memory usage required to generate a token based on a Decode inference request can be the GPU memory usage required to generate a token based on the Prompt included in the Decode inference request.
[0119] In some embodiments, in order to ensure the rationality and availability of GPU memory allocation, when allocating GPU memory for the Decode inference request set according to the GPU memory demand, the current GPU memory usage corresponding to the Decode inference request set can be obtained first. Then, according to the current GPU memory usage and the GPU memory demand, it can be determined whether the GPU memory allocation condition and the GPU memory release condition are met.
[0120] It should be noted that the above GPU memory allocation condition and the above GPU memory release condition are two mutually exclusive conditions. That is, if the GPU memory allocation condition is met, it means that the GPU memory release condition is not met; if the GPU release allocation condition is met, it means that the GPU memory allocation condition is not met.
[0121] If the above GPU memory allocation conditions are met, GPU memory can be allocated to the above Decode inference request set according to the above GPU memory demand.
[0122] If the above-mentioned GPU memory release condition is met, the GPU memory used to execute the above-mentioned Decode inference request set can be released, and the above-mentioned current CPU memory usage is updated, so that when it is determined that the above-mentioned GPU memory allocation condition is met based on the updated current GPU memory usage and the above-mentioned GPU memory demand, GPU memory is allocated to the Decode inference request set according to the GPU memory demand.
[0123] It should be noted that the amount of GPU memory used that has been released can be determined according to the executed GPU memory release task, and the released GPU memory usage amount can be subtracted from the above-mentioned current GPU memory usage amount to obtain the updated current GPU memory usage amount. However, in practical applications, in order to obtain a more accurate current GPU memory usage amount, the current GPU memory usage amount can be obtained in real time or periodically through tools such as DCGM (Data Center GPU Manager). DCGM is a tool for managing and monitoring GPU resources in a data center. It can help effectively manage large-scale GPU clusters and provide in-depth insights into GPU performance, health, and utilization.
[0124] In some embodiments, the above-mentioned memory allocation condition can be that the sum of the current GPU memory usage amount and the GPU memory demand amount is less than a preset threshold (which can be referred to as the first threshold); the above-mentioned GPU memory release condition can be that the current GPU memory usage amount is greater than a preset threshold (which can be referred to as the second threshold). Among them, the specific values of the first threshold and the second threshold can be the same or different, and the present application does not impose special restrictions on this.
[0125] Specifically, let DecodeNeedResource[T] represent the GPU memory demand amount in the T-th Decode memory management time window, DecodeCurrentResource[T] represent the current GPU memory usage amount obtained when performing GPU memory management in the T-th Decode memory management time window, Threshold1 represent the above-mentioned first threshold, and Threshold2 represent the above-mentioned second threshold. Then, it can be determined that the above-mentioned memory allocation condition is satisfied when DecodeCurrentResource[T] + DecodeNeedResource[T] < Threshold1, and it can be determined that the above-mentioned GPU memory release condition is satisfied when DecodeCurrentResource[T] > Threshold2.
[0126] In practical applications, the period when DecodeCurrentResource[T] + DecodeNeedResource[T] > Threshold1 and DecodeCurrentResource[T] < Threshold2 can be used as the buffer time for releasing GPU memory. If during this period, DecodeCurrentResource[T] changes on its own, such that the updated DecodeCurrentResource[T] + DecodeNeedResource[T] < Threshold1, then since the above memory allocation condition is satisfied, GPU memory can be allocated for the above Decode inference request set according to the above GPU memory requirement. If during this period, DecodeCurrentResource[T] changes on its own, such that DecodeCurrentResource[T] > Threshold2, then since the above memory release condition is satisfied, the GPU memory used to execute the above Decode inference request set can be released. In this way, the number of times of executing the GPU memory release task can be reduced to a certain extent.
[0127] In some embodiments, various methods can be used to implement the release of GPU memory. For example, when releasing the GPU memory used to execute the above Decode inference request set, the specific GPU memory release method can be selected according to the number of Decode inference requests in the Decode inference request set.
[0128] Specifically, it can be first determined whether the number of Decode inference requests in the above Decode inference request set is greater than a preset threshold (which can be referred to as the third threshold).
[0129] If the number of Decode inference requests in the above Decode inference request set is greater than the above third threshold, then the tokens corresponding to each Decode inference request in the Decode inference request set can be offloaded from the GPU memory to the CPU memory to release the GPU memory occupied by these tokens.
[0130] If the number of Decode reasoning requests in the above-mentioned Decode reasoning request set is not greater than the above-mentioned third threshold, the Decode reasoning request can be updated based on the tokens corresponding to each Decode reasoning request in the Decode reasoning request set to release the GPU memory occupied by these tokens. Wherein, updating the Decode reasoning request based on the token corresponding to a Decode reasoning request means using the token generated based on the Prompt contained in the Decode reasoning request to update the Prompt to form a new Prompt.
[0131] In the above technical solution, the inference engine in the inference system can use the GPU mounted on the computing device where it is located as a computing resource to execute inference requests and manage GPU memory. Specifically, the inference engine can maintain a scheduling queue for scheduling inference request sets, and determine the memory management time window according to the data processing duration associated with the inference request set being executed in the scheduling queue, and subsequently calculate the GPU memory demand corresponding to the inference request set within the memory management time window, and allocate GPU memory to the inference request set according to the GPU memory demand, and after the end of the memory management time window, the next memory management time window corresponding to the memory management time window can be determined again according to the data processing duration associated with the inference request set being executed in the scheduling queue, so as to perform GPU memory management again within the next memory management time window.
[0132] By adopting the above method, on the one hand, there is no need to reserve a large amount of GPU memory for the inference engine when it is started. Instead, in the process of the inference engine executing inference requests in batches, a memory management time window can be continuously set, and the GPU memory demand corresponding to the inference request set being executed within the memory management time window can be predicted, so as to allocate GPU memory to the inference request set according to the GPU memory demand. This not only ensures that the inference engine has sufficient GPU memory available during the process of batch executing inference requests, but also avoids waste of GPU memory. On the other hand, the memory management time window can be set according to different stages of the inference process, and the GPU demand within the set memory management time window can be predicted, so as to realize staged GPU memory management in the inference process, thereby improving the accuracy and adaptability of GPU memory management.
[0133] Corresponding to the aforementioned method embodiments, the present application also provides device embodiments.
[0134] Please refer to Figure 4 , Figure 4It is a structural diagram of a device shown in an exemplary embodiment of the present application. At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, a memory 408 and a non-volatile memory 410, and of course may also include other required hardware. One or more embodiments of the present application can be implemented based on software, such as the processor 402 reading the corresponding computer program from the non-volatile memory 410 into the memory 408 and then running it. Of course, in addition to the software implementation, one or more embodiments of the present application do not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic module, but can also be hardware or logic devices.
[0135] Please refer to Figure 5 , Figure 5 It is a block diagram of a memory management device of an inference system shown in an exemplary embodiment of the present application.
[0136] The memory management device of the above reasoning system can be applied to Figure 4 The device shown in the figure is used to implement the technical solution of the present application. Among them, the inference engine in the inference system can be deployed on the device, and the memory management device of the inference system can be specifically applied to the inference engine; the computing resources of the inference engine include a GPU mounted on the computing device for deploying the inference engine; the inference engine maintains a Prefill scheduling queue for scheduling the inference request set in the Prefill stage, and a Decode scheduling queue for scheduling the inference request set in the Decode stage; the device includes:
[0137] The first memory management module 502 determines a Prefill memory management time window according to the data processing duration associated with the Prefill reasoning request set being executed in the Prefill scheduling queue, calculates the GPU memory demand corresponding to the Prefill reasoning request set within the Prefill memory management time window, and allocates GPU memory to the Prefill reasoning request set according to the GPU memory demand; at the end of the Prefill memory management time window, re-determines a next Prefill memory management time window corresponding to the Prefill memory management time window according to the data processing duration associated with the Prefill reasoning request set being executed in the Prefill scheduling queue;
[0138] The second memory management module 504 determines a Decode memory management time window according to the data processing duration associated with the Decode inference request set being executed in the Decode scheduling queue, calculates the GPU memory requirement corresponding to the Decode inference request set within the Decode memory management time window, and allocates GPU memory to the Decode inference request set according to the GPU memory requirement; at the end of the Decode memory management time window, re-determines a next Decode memory management time window corresponding to the Decode memory management time window according to the data processing duration associated with the Decode inference request set being executed in the Decode scheduling queue.
[0139] In some embodiments, determining the Prefill memory management time window according to the data processing duration associated with the set of Prefill inference requests being executed in the Prefill scheduling queue includes:
[0140] Determine a maximum number of tokens corresponding to inference requests in the set of Prefill inference requests being executed in the Prefill scheduling queue, and determine a Prefill memory management time window according to a token processing duration corresponding to the maximum number of tokens.
[0141] In some embodiments, determining the Decode memory management time window according to the data processing duration associated with the Decode inference request set being executed in the Decode scheduling queue includes:
[0142] Determine a maximum token generation duration corresponding to an inference request in a Decode inference request set being executed in the Decode scheduling queue, and determine a Decode memory management time window according to the maximum token generation duration.
[0143] In some embodiments, calculating the GPU memory demand corresponding to the inference request set within the memory management time window includes:
[0144] Calculate the static GPU memory usage corresponding to the inference request set within the memory management time window;
[0145] Determine a maximum GPU memory usage in a previous memory management time window corresponding to the memory management time window, and a sum of static GPU memory usages corresponding to the inference request set in each target memory management time window, and determine a difference between the maximum GPU memory usage and the sum of the static GPU memory usages as a dynamic GPU memory usage corresponding to the inference request set in the memory management time window; wherein the target memory management window is a memory management window located before the memory management time window during the execution of the inference request set;
[0146] The sum of the static GPU memory usage and the dynamic GPU memory usage is determined as the GPU memory demand corresponding to the inference request set within the memory management time window.
[0147] In some embodiments, the calculating the static GPU memory usage corresponding to the inference request set within the memory management time window includes:
[0148] Determine a first GPU memory usage required for a token corresponding to each inference request in the Prefill inference request set, and a second GPU memory usage required to generate a token based on each inference request in the Prefill inference request set, and determine the sum of the first GPU memory usage and the second GPU memory usage as the static GPU memory usage corresponding to the Prefill inference request set within the Prefill memory management time window.
[0149] In some embodiments, the calculating the static GPU memory usage corresponding to the inference request set within the memory management time window includes:
[0150] Determine a GPU memory usage required to generate a token based on each inference request in the Decode inference request set, and determine the GPU memory usage as a static GPU memory usage corresponding to the Decode inference request set within the Decode memory management time window.
[0151] In some embodiments, allocating GPU memory for the Decode inference request set according to the GPU memory requirement includes:
[0152] Get the current GPU memory usage corresponding to the Decode inference request set;
[0153] Determine whether a GPU memory allocation condition and a GPU memory release condition are met according to the current GPU memory usage and the GPU memory demand;
[0154] If the GPU memory allocation condition is met, allocating GPU memory to the Decode inference request set according to the GPU memory requirement;
[0155] If the GPU memory release condition is met, the GPU memory used to execute the Decode inference request set is released, and the current CPU memory usage is updated, so that when it is determined that the GPU memory allocation condition is met based on the updated current GPU memory usage and the GPU memory requirement, GPU memory is allocated to the Decode inference request set according to the GPU memory requirement.
[0156] In some embodiments, the GPU memory allocation condition includes:
[0157] The sum of the current GPU memory usage and the GPU memory requirement is less than a preset first threshold;
[0158] The GPU memory release conditions include:
[0159] The current GPU memory usage is greater than a preset second threshold.
[0160] In some embodiments, releasing the GPU memory used to execute the Decode inference request set includes:
[0161] Determine whether the number of inference requests in the Decode inference request set is greater than a preset third threshold;
[0162] If the number is greater than the third threshold, unloading the token corresponding to each inference request in the Decode inference request set to the CPU memory to release the GPU memory occupied by the token;
[0163] If the number is not greater than the third threshold, based on the tokens corresponding to the respective inference requests in the Decode inference request set, the inference requests are updated to release the GPU memory occupied by the tokens.
[0164] For the device embodiment, it basically corresponds to the method embodiment, so the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the technical solution of the present application.
[0165] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver, a game console, a tablet computer, a wearable device or a combination of any of these devices.
[0166] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0167] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0168] Computer readable media include permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0169] It should be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0170] The above specific embodiments of the present application are described. Other embodiments are within the scope of the present application. In some cases, the actions or steps recorded in the present application can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0171] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present application. The singular forms of "a", "said" and "the" are also intended to include plural forms, unless the context clearly indicates other meanings. The term "and / or" refers to and includes any or all possible combinations of one or more associated listed items.
[0172] The description of the terms "one embodiment", "some embodiments", "example", "specific example" or "an implementation method" etc. used in one or more embodiments of the present application means that the specific features or characteristics described in conjunction with the embodiment are included in at least one embodiment of the present application. The schematic description of these terms is not necessarily for the same embodiment. Moreover, the specific features or characteristics described can be combined in a suitable manner in one or more embodiments of the present application. In addition, different embodiments and specific features or characteristics in different embodiments can be combined without contradicting each other.
[0173] It should be understood that, although the terms first, second, third, etc. may be used to describe various information in one or more embodiments of the present application, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determination".
[0174] The above description is merely a preferred embodiment of one or more embodiments of the present application and is not intended to limit one or more embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application shall be included in the scope of protection of one or more embodiments of the present application.
[0175] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
Claims
1. A memory management method for an inference system, applied to an inference engine in the inference system; the computing resources of the inference engine include a GPU mounted on a computing device for deploying the inference engine; the inference engine maintains a Prefill scheduling queue for scheduling an inference request set in a Prefill stage, and a Decode scheduling queue for scheduling an inference request set in a Decode stage; the method comprises: Determine a Prefill memory management time window according to a data processing duration for data processing of a set of Prefill inference requests being executed in the Prefill scheduling queue, calculate a GPU memory requirement corresponding to the set of Prefill inference requests within the Prefill memory management time window, and allocate GPU memory to the set of Prefill inference requests according to the GPU memory requirement; At the end of the Prefill memory management time window, re-determine a next Prefill memory management time window corresponding to the Prefill memory management time window according to the data processing duration for the Prefill reasoning request set being executed in the Prefill scheduling queue; Determine a Decode memory management time window according to a data processing duration for data processing of a Decode inference request set being executed in the Decode scheduling queue, calculate a GPU memory requirement corresponding to the Decode inference request set within the Decode memory management time window, and allocate GPU memory to the Decode inference request set according to the GPU memory requirement; When the Decode memory management time window ends, a next Decode memory management time window corresponding to the Decode memory management time window is determined again based on the data processing duration for the Decode inference request set being executed in the Decode scheduling queue.
2. The method according to claim 1, wherein determining the Prefill memory management time window according to the data processing duration associated with the set of Prefill inference requests being executed in the Prefill scheduling queue comprises: Determine a maximum number of tokens corresponding to inference requests in the set of Prefill inference requests being executed in the Prefill scheduling queue, and determine a Prefill memory management time window according to a token processing duration corresponding to the maximum number of tokens.
3. The method according to claim 1, wherein determining the Decode memory management time window according to the data processing duration associated with the Decode inference request set being executed in the Decode scheduling queue comprises: Determine a maximum token generation duration corresponding to an inference request in a Decode inference request set being executed in the Decode scheduling queue, and determine a Decode memory management time window according to the maximum token generation duration.
4. The method according to claim 1, calculating the GPU memory demand corresponding to the inference request set within the memory management time window, comprising: Calculate the static GPU memory usage corresponding to the inference request set within the memory management time window; Determine a maximum GPU memory usage in a previous memory management time window corresponding to the memory management time window, and a sum of static GPU memory usages corresponding to the inference request set in each target memory management time window, and determine a difference between the maximum GPU memory usage and the sum of the static GPU memory usages as a dynamic GPU memory usage corresponding to the inference request set in the memory management time window; wherein the target memory management window is a memory management window located before the memory management time window during the execution of the inference request set; The sum of the static GPU memory usage and the dynamic GPU memory usage is determined as the GPU memory demand corresponding to the inference request set within the memory management time window.
5. The method according to claim 4, wherein the step of calculating the static GPU memory usage corresponding to the inference request set within the memory management time window comprises: Determine a first GPU memory usage required for a token corresponding to each inference request in the Prefill inference request set, and a second GPU memory usage required to generate a token based on each inference request in the Prefill inference request set, and determine the sum of the first GPU memory usage and the second GPU memory usage as the static GPU memory usage corresponding to the Prefill inference request set within the Prefill memory management time window.
6. The method according to claim 4, wherein the step of calculating the static GPU memory usage corresponding to the inference request set within the memory management time window comprises: Determine a GPU memory usage required to generate a token based on each inference request in the Decode inference request set, and determine the GPU memory usage as a static GPU memory usage corresponding to the Decode inference request set within the Decode memory management time window.
7. The method according to claim 1, allocating GPU memory for the Decode inference request set according to the GPU memory requirement, comprising: Get the current GPU memory usage corresponding to the Decode inference request set; Determine whether a GPU memory allocation condition and a GPU memory release condition are met according to the current GPU memory usage and the GPU memory demand; If the GPU memory allocation condition is met, allocating GPU memory to the Decode inference request set according to the GPU memory requirement; If the GPU memory release condition is met, the GPU memory used to execute the Decode inference request set is released, and the current CPU memory usage is updated, so that when it is determined that the GPU memory allocation condition is met based on the updated current GPU memory usage and the GPU memory requirement, GPU memory is allocated to the Decode inference request set according to the GPU memory requirement.
8. The method according to claim 7, wherein the GPU memory allocation condition comprises: The sum of the current GPU memory usage and the GPU memory requirement is less than a preset first threshold; The GPU memory release conditions include: The current GPU memory usage is greater than a preset second threshold.
9. The method according to claim 7, wherein releasing GPU memory used to execute the Decode inference request set comprises: Determine whether the number of inference requests in the Decode inference request set is greater than a preset third threshold; If the number is greater than the third threshold, unloading the token corresponding to each inference request in the Decode inference request set to the CPU memory to release the GPU memory occupied by the token; If the number is not greater than the third threshold, based on the tokens corresponding to the respective inference requests in the Decode inference request set, the inference requests are updated to release the GPU memory occupied by the tokens.
10. A memory management device for an inference system, applied to an inference engine in the inference system; the computing resources of the inference engine include a GPU mounted on a computing device for deploying the inference engine; the inference engine maintains a Prefill scheduling queue for scheduling an inference request set in a Prefill stage, and a Decode scheduling queue for scheduling an inference request set in a Decode stage; the device comprises: a first memory management module, determining a Prefill memory management time window according to a data processing duration for data processing on a set of Prefill inference requests being executed in the Prefill scheduling queue, calculating a GPU memory requirement corresponding to the set of Prefill inference requests within the Prefill memory management time window, and allocating GPU memory to the set of Prefill inference requests according to the GPU memory requirement; and determining a subsequent Prefill memory management time window corresponding to the Prefill memory management time window according to a data processing duration for data processing on a set of Prefill inference requests being executed in the Prefill scheduling queue at the end of the Prefill memory management time window; a second memory management module, which determines a Decode memory management time window according to a data processing duration for data processing for the Decode inference request set being executed in the Decode scheduling queue, calculates a GPU memory requirement corresponding to the Decode inference request set within the Decode memory management time window, and allocates GPU memory to the Decode inference request set according to the GPU memory requirement; and when the Decode memory management time window ends, re-determines a next Decode memory management time window corresponding to the Decode memory management time window according to the data processing duration for data processing for the Decode inference request set being executed in the Decode scheduling queue.
11. An electronic device, comprising: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 9 by running the executable instructions.
12. A computer-readable storage medium having computer instructions stored thereon, wherein the instructions are executed by a processor to implement the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Memory allocation method and device for inference card, electronic equipment and storage medium
CN115495248A
Large language model reasoning system, method and equipment without perception of server
CN116702907A