Memory Management Method and Device for Inference System

By dynamically managing GPU memory in the inference engine, determining the memory management time window and allocating memory based on the data processing time of inference requests, the problem of inefficient resource management in the existing technology is solved, and efficient utilization of GPU memory and performance improvement of the inference engine are achieved.

CN119248522BActive Publication Date: 2025-05-27ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411783366.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-27
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

Existing inference systems have problems with inefficiency in resource management, especially in the reservation and utilization of GPU memory by the inference engine, which leads to waste of resources and performance degradation.

Method used

By maintaining the scheduling queue in the inference engine, the memory management time window is determined based on the data processing time of the executed inference request set, and the GPU memory requirement in the time window is calculated, and the GPU memory is dynamically allocated.

Benefits of technology

It realizes that in the inference engine performs inference requests in batches, it can manage GPU memory flexibly and efficiently, avoid resource waste, and ensures the performance of the inference engine under high load conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119248522B_ABST
    Figure CN119248522B_ABST
Patent Text Reader

Abstract

One or more embodiments of the present application provide a memory management method and apparatus for an inference system. The method is applied to an inference engine in the inference system. The computing resources of the inference engine include a GPU installed on a computing device for deploying the inference engine. The inference engine maintains a scheduling queue for scheduling a set of inference requests. The method includes: determining a memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue; calculating the GPU memory requirement corresponding to the set of inference requests within the memory management time window, and allocating GPU memory for the set of inference requests according to the GPU memory requirement; at the end of the memory management time window, re-determining a subsequent memory management time window corresponding to the memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method and device for memory management of an inference system. Background Art

[0002] An inference system is a computer program that uses logical rules and known facts to draw new conclusions or make decisions. The inference system is an important part of the field of artificial intelligence and is mainly used to simulate the human decision-making process. It is based on a set of defined knowledge bases and inference engines to derive conclusions. The inference system can execute the inference requests it receives and output the corresponding inference results.

[0003] A typical inference system usually consists of the following parts: a knowledge base, an inference engine, a user interface, and an explanation facility. Among them, the knowledge base includes all the facts and rules known to the system. These facts can be about the state of the world, object attributes, etc., and the rules are logical expressions that describe how to draw new conclusions from known facts. The inference engine is the core component of the inference system. It is responsible for performing logical operations in the inference process, that is, drawing new conclusions or making decisions from the given knowledge base; the inference engine uses a series of rules and known facts to derive new knowledge, thereby helping the system solve problems or make decisions. The user interface allows users to interact with the system, input queries, or observe the results of the inference process. The explanation facility is used to explain how the system arrives at a specific conclusion, which is very important for transparency and trust.

[0004] The inference engine usually uses resources such as computing resources, storage resources, and network resources (for example: GPU, GPU memory, storage, network interface, etc.) to execute inference tasks. The efficient use of these resources directly affects the performance of the inference engine. Therefore, it is desirable to be able to manage the resources used by the inference engine better and more flexibly. Summary of the Invention

[0005] One or more embodiments of the present application provide the following technical solutions:

[0006] The present application provides a method for memory management of an inference system, which is applied to an inference engine in the inference system; the computing resources of the inference engine include a GPU mounted on a computing device for deploying the inference engine; the inference engine maintains a scheduling queue for scheduling a set of inference requests; the method includes:

[0007] Determine a memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue;

[0008] Calculate the GPU memory requirement corresponding to the set of inference requests within the memory management time window, and allocate GPU memory to the set of inference requests according to the GPU memory requirement;

[0009] At the end of the memory management time window, re-determine the next memory management time window corresponding to the memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue.

[0010] This application also provides a memory management device for an inference system, which is applied to an inference engine in the inference system; the computing resources of the inference engine include the GPU installed on the computing device for deploying the inference engine; the inference engine maintains a scheduling queue for scheduling a set of inference requests; the device includes:

[0011] A time window determination module, which determines a memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue;

[0012] A GPU memory allocation module, which calculates the GPU memory requirement corresponding to the set of inference requests within the memory management time window, and allocates GPU memory to the set of inference requests according to the GPU memory requirement;

[0013] The time window determination module is further configured to, at the end of the memory management time window, re-determine the next memory management time window corresponding to the memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue.

[0014] This application also provides an electronic device, including:

[0015] A processor;

[0016] A memory for storing instructions executable by the processor;

[0017] Wherein, the processor realizes the steps of the method as described in any one of the above by running the executable instructions.

[0018] This application also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any one of the above are realized.

[0019] In the above technical solution, the inference engine in the inference system can use the GPU installed on the computing device where it is located as the computing resource to execute inference requests and manage the GPU memory. Specifically, the inference engine can maintain a scheduling queue for scheduling a set of inference requests, and determine a memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue. Subsequently, it can calculate the GPU memory demand corresponding to the set of inference requests within the memory management time window, and allocate GPU memory to the set of inference requests according to the GPU memory demand. After the end of the memory management time window, it can re-determine the next memory management time window corresponding to the memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue, so as to perform GPU memory management again within the next memory management time window.

[0020] By adopting the above method, there is no need to reserve a large amount of GPU memory for the inference engine when it starts. Instead, during the process of the inference engine batch-executing inference requests, the memory management time window can be continuously set, and the GPU memory demand corresponding to the set of inference requests being executed within the memory management time window can be predicted, so as to allocate GPU memory to the set of inference requests according to the GPU memory demand. This can not only ensure that there is sufficient GPU memory for the inference engine to use during the process of batch-executing inference requests, but also avoid waste of GPU memory. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings required for the description of the exemplary embodiments will be described below, where:

[0022] Figure 1 is a schematic diagram of an inference system shown in an exemplary embodiment of the present application.

[0023] Figure 2 is a schematic diagram of another inference system shown in an exemplary embodiment of the present application.

[0024] Figure 3 is a flowchart of a memory management method for an inference system shown in an exemplary embodiment of the present application.

[0025] Figure 4 is a schematic structural diagram of a device shown in an exemplary embodiment of the present application.

[0026] Figure 5 is a block diagram of a memory management device for an inference system shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of the present application. On the contrary, they are merely examples consistent with some aspects of one or more embodiments of the present application.

[0028] It should be noted that in other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in the present application. In some other embodiments, the steps included in the method may be more or fewer than those described in the present application. In addition, a single step described in the present application may be decomposed into multiple steps for description in other embodiments; and multiple steps described in the present application may also be combined into a single step for description in other embodiments.

[0029] In the present application, the inference engine in the inference system can be deployed on a computing device and use the resources of the computing device (such as computing resources, storage resources, network resources, etc., for example: GPU, GPU memory, storage, network interface, etc.) to perform calculations during the inference process and execute inference tasks. Alternatively, the inference engine in the inference system can be deployed on a computing cluster to meet the needs of processing large-scale data, high-concurrency requests, or high-performance computing, thereby achieving higher computing power, better fault tolerance, and more flexible resource management.

[0030] Generally, the inference engine can use the CPU as the computing resource, where the CPU can be the CPU installed on the computing device for deploying the inference engine.

[0031] With the development of artificial intelligence technology, inference systems based on large models (such as: large language model Large Language Model) are increasingly widely used.

[0032] A large model refers to a machine learning model with a large number of parameters, such as various variants under the Transformer architecture, including but not limited to natural language processing models such as GPT and BERT. These models achieve powerful representation learning capabilities through a large amount of training data and complex architectures.

[0033] In an inference system based on a large model, the large model can be regarded as a huge knowledge base, which contains information learned from a large amount of data. The large model learns a large number of patterns and features during the training process, and these patterns and features represent the complex relationships in the data. Therefore, to a certain extent, the information stored inside the large model can be regarded as a form of knowledge.

[0034] The inference engine designed for large models is responsible for calling large models to perform specific inference tasks, that is, it can manage the loading of large models, execute the inference operations of large models, and manage the interaction with hardware (such as: CPU, GPU or other accelerators). The inference engine can also include optimization algorithms to improve the speed and efficiency of inference.

[0035] In practical applications, large models can be deployed separately from the inference system, or large models can be integrated into the inference system, and the inference engine therein can be used to call large models to efficiently execute inference tasks.

[0036] Since large models are AI applications that require a large amount of computing resources, the inference engine usually uses GPUs as computing resources when calling large models to perform inference tasks.

[0037] GPUs are designed for parallel processing and can process multiple data points simultaneously. This is especially useful for deep learning models as they often need to perform the same operations on large amounts of data. Modern AI models, especially neural network models, require a large number of floating-point operations. GPUs generally perform better than CPUs in floating-point operations, especially when dealing with large-scale matrix operations, which are common operations in deep learning. GPUs usually have higher memory bandwidth than CPUs, which means they can read and write data from memory faster. This is very important for applications that need to frequently access large amounts of data. Many GPU vendors (such as: NVIDIA) have optimized their hardware and software stacks for machine learning tasks. For example, they provide hardware specifically designed to accelerate tensor operations, as well as programming models like CUDA to fully utilize these hardware features. GPUs can improve the inference speed. For applications deployed on a large scale, using GPUs can reduce latency and improve the response speed of the system.

[0038] During the operation of the inference engine, whether it can efficiently utilize computing resources, storage resources, network resources and other resources of computing devices usually directly affects the performance of the inference engine.

[0039] For example, during the inference process of large models, KV Cache (Key-Value Cache) is usually used to store the generated tokens (Token).

[0040] It should be noted that the inference process of a large model refers to the process of using a pre-trained large machine learning model to generate corresponding output results (i.e., inference results) based on the given input data (i.e., the data included in the inference request). In the field of natural language processing, especially in large language models, the inference process usually involves generating a continuous text sequence based on the given input prompt (Prompt), with one Token generated in each iteration. Here, a Token usually refers to a series of small units into which text is segmented in natural language processing, such as words or characters, depending on the tokenization method of the model.

[0041] Specifically, a piece of text can be provided to the model as input, and this piece of text is called a Prompt; the model receives this Prompt and converts it into a form that can be processed internally, usually by converting each Token in the Prompt into a corresponding numerical representation, for example: an embedding vector. An embedding vector is a numerical vector in a high-dimensional space used to represent the semantic information of each Token; based on the current Prompt, the model predicts the next most likely Token to appear through complex calculations (such as: multi-layer neural network operations); the model adds the predicted Token to the existing Prompt to form a new Prompt; the above steps can be repeated multiple times, and each repetition is called an iteration. In each iteration, the model will predict the next Token based on the latest Prompt until a set stopping condition is reached, such as: a certain number of Tokens have been generated or a specific end flag has been encountered.

[0042] KV Cache is a data storage method usually used to accelerate data access. It stores data in the form of key-value pairs, where the key is the unique identifier for looking up data, and the value is the data associated with it or information pointing to the data location. When an application requests specific data, the system can quickly retrieve the corresponding value through the key, thereby improving the system's response speed.

[0043] In large models, especially in large models used to perform natural language processing tasks, KV Cache is used to store previous calculation results for quick reuse during the decoding process, thereby accelerating the inference process.

[0044] In practical applications, using a KV Cache to store the generated tokens consumes a large amount of GPU memory. For an inference engine based on a large model, when the inference engine starts, a certain amount of GPU memory needs to be reserved for the inference engine according to past experience and actual situations for subsequent model inference in the inference engine, where the allocation ratio is usually 80% to 90% of the total GPU memory. Moreover, due to the uncertainty of the intermediate data generated during the inference process of the large model, a certain amount of GPU memory also needs to be reserved for this intermediate data.

[0045] However, the traffic of the inference engine is not fixed, that is, the number of inference tasks to be executed by the inference engine may vary over time. For example, there are more inference requests sent to the inference engine in the current time period, while there are fewer inference requests sent to the inference engine in the next time period. In this case, the GPU memory reserved for the inference engine will be idle when the traffic is small and cannot be used by other services; in addition, the remaining 10% to 20% of the GPU memory not reserved for the inference engine cannot be used by the inference engine either. Therefore, the GPU memory of the computing device used to deploy the inference engine cannot be efficiently utilized, which will affect the performance of the inference engine.

[0046] In the technical solution provided by one or more embodiments of the present application, the inference engine in the inference system can use the GPU installed on the computing device where it is located as a computing resource to execute inference requests and manage the GPU memory. Specifically, the inference engine can maintain a scheduling queue for scheduling a set of inference requests, and determine a memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue. Subsequently, the GPU memory demand corresponding to the set of inference requests within the memory management time window can be calculated, and the GPU memory can be allocated to the set of inference requests according to the GPU memory demand. After the memory management time window ends, the next memory management time window corresponding to the memory management time window can be determined again according to the data processing duration associated with the set of inference requests being executed in the scheduling queue for GPU memory management in the next memory management time window.

[0047] By adopting the above method, it is not necessary to reserve a large amount of GPU memory for the inference engine when it starts. Instead, during the process of the inference engine batch-executing inference requests, the memory management time window can be continuously set, and the GPU memory demand corresponding to the set of inference requests being executed within the memory management time window can be predicted, so as to allocate the GPU memory to the set of inference requests according to the GPU memory demand. This can not only ensure that the inference engine has enough GPU memory to use during the process of batch-executing inference requests, but also avoid waste of GPU memory.

[0048] Please refer to Figure 1 and Figure 2 , Figure 1 and Figure 2 which are respectively schematic diagrams of an inference system shown in an exemplary embodiment of the present application.

[0049] As Figure 1 shown, the above-mentioned inference system may include an API server (API Server) and an inference engine.

[0050] The API server is a server specifically designed to process client requests and return responses, and is usually an indispensable component in an application. It is responsible for processing client requests, executing business logic, generating responses, and ensuring the security, scalability, and performance of the system.

[0051] Specifically, the API server usually serves as a unified entry of the inference system, processes all requests from clients, and can distribute requests to different backend services to achieve decoupling between services. The API server can be responsible for authentication and authorization, ensuring that only legitimate requests can access system resources, and implementing fine-grained permission control to protect sensitive data and operations. The API server can also cooperate with a load balancer to evenly distribute requests to multiple backend instances, improving the availability and performance of the inference system.

[0052] The above-mentioned inference engine can be deployed on an independent computing device, or can be deployed on a computing cluster (Cluster) composed of at least one computing node (Node). Among them, the computing device can be a physical or virtual local computing device, or a physical or virtual cloud computing device.

[0053] A computing cluster is a system in which at least one computing device is connected through a network to work collaboratively. These computing devices jointly complete computing tasks, providing stronger computing power and higher availability than a single computing device. A computing node is one of the core components of a computing cluster, usually referring to a computing device used to execute computing tasks, and the GPU installed on it can be used as a computing resource.

[0054] Correspondingly, the above-mentioned API server can be deployed on the computing device where the above-mentioned inference engine is located; or, the above-mentioned API server can be deployed on the head node in the above-mentioned computing cluster, where the head node usually refers to a computing node responsible for the management and scheduling of the computing cluster, and can also be deployed on other computing devices outside the computing cluster.

[0055] As Figure 2As shown, for each computing device in an independent computing device or computing cluster used to deploy the above-mentioned inference engine, at least one computing instance can be deployed on the computing device. Among them, a computing instance refers to dedicated resources allocated on demand through virtualization technology in a local data center or a cloud computing platform. The resources of these computing instances can include computing resources, storage resources, network resources, etc. (for example: GPU, GPU memory, storage, network interface, etc.), enabling each computing instance to use these resources to perform specific computing tasks. Computing instances can be created, started, stopped, or deleted according to actual needs, and their resources can also be adjusted according to requirements.

[0056] The computing resources of a computing instance can be the GPU of the computing device where it is located. Specifically, a computing device can carry multiple GPUs, and one GPU can be allocated to one computing instance as the computing resources of this computing instance. Or, a computing device can carry only one GPU, and the computing resources allocated to a computing instance can be a part of the computing unit, memory, etc. of this GPU.

[0057] In an inference engine based on a large model, each computing instance can call the large model to perform specific inference work, that is, the large model can be loaded into each computing instance, and each computing instance executes the inference operation of the large model. Specifically, the large model can be read from a local or cloud storage medium (for example: local disk, cloud storage, etc.) into the GPU memory of each computing instance to perform inference based on the large model on the GPU resources of each computing instance.

[0058] In practical applications, a Model Runtime framework can be installed in the computing instance, and after the large model is read into the computing instance, the model service can be configured according to the model runtime framework and the large model, so that the computing instance can be used as an independently running model service instance. Among them, model runtime refers to the environment and framework used to execute model inference after the model is deployed. This environment usually includes a series of steps such as model loading, input data processing, model execution, and output result processing, covering the entire life cycle of the model from loading to execution and then to unloading.

[0059] For the above-mentioned inference engine including at least one computing instance, a Load-aware Scheduling method can be adopted to implement the scheduling of inference requests on the inference engine.

[0060] Load-aware Scheduling refers to a strategy of allocating tasks according to the current load conditions of each node or resource in a computing system. This scheduling method aims to optimize the resource usage efficiency, reduce waiting time and processing latency, and improve the overall throughput of the system.

[0061] Load-aware scheduling is usually applied in scenarios such as distributed systems, cloud computing platforms, and data centers. Its purpose is to achieve dynamic balance of the workloads of each node and avoid the situation where some nodes are overloaded while others are idle. By monitoring the load conditions of each node in real time, the load-aware scheduler can make more reasonable task allocation decisions.

[0062] For example, in the above inference system, communication can occur between the above API server and the above inference engine. In this case, the API server can obtain inference requests from the request queue and, based on the load conditions of the GPUs of each computing instance, select a computing instance with appropriate GPU load conditions and schedule the obtained inference requests to this computing instance to use the GPU resources of this computing instance to execute the inference requests.

[0063] By deploying computing instances on computing devices, firstly, the resources of the computing devices can be fully utilized to accelerate the execution of computing tasks, thereby improving the computing power of the inference engine; secondly, different computing instances are independent of each other, which can ensure the stability and security of computing tasks and can achieve resource isolation and management, contributing to the effective utilization of resources and avoiding resource appropriation; thirdly, for computing tasks that require parallel processing, the completion time of the computing tasks can be accelerated through parallel computing; fourthly, elastic scaling can be achieved. When a large number of computing tasks need to be processed, computing instances can be added, and when the number of computing tasks to be processed decreases, some computing instances can be terminated to release some resources and achieve flexible allocation of resources.

[0064] For the above API server, two components can be included inside the API server, namely a Request Router component responsible for routing inference requests to forward the inference requests to the above inference engine, and a Monitor component responsible for monitoring the running-related information of all computing instances in the inference engine (such as: GPU memory utilization rate, GPU memory bandwidth utilization rate, the number of inference requests waiting in the scheduling queue and the number of running inference requests, etc.).

[0065] For the computing device used to deploy the above inference engine, in order to facilitate the maintenance, management, and scheduling of the computing instances deployed on this computing device, an Agent can also be deployed on this computing device.

[0066] The above API server can communicate with the agent deployed on the above computing device. In this case, the API server can obtain the inference request from the request queue, and select a computing instance with appropriate GPU load according to the GPU load of each computing instance, and send the obtained inference request to the agent deployed on the computing device where the computing instance is located. The agent further schedules the inference request to the computing instance to use the GPU resources of the computing instance to execute the inference request.

[0067] In practical applications, to ensure the reliability and correctness of inference request scheduling, for the above agent, it can maintain a schedule queue.

[0068] The schedule queue is a special type of task queue, mainly used to manage and schedule tasks with specific execution times or priorities (in this application, it is the inference task specified for the inference request). The main functions of the schedule queue include task scheduling, task management, and resource optimization. Among them, task scheduling means that the schedule queue can ensure that tasks are executed at the specified time point or within the specified time range, and at the same time support the scheduling of periodic tasks; and, the priority of tasks can be set to ensure that high-priority tasks in the schedule queue are executed first. Task management means that all tasks are centralized in a queue for management, which is convenient for monitoring and control, and at the same time, the execution status of tasks can be tracked, including unexecuted, executing, completed, and failed statuses. Resource optimization means that by reasonably arranging the execution time of tasks, the system can be prevented from being overloaded at a certain moment, and the system resources can be effectively utilized to avoid resource waste.

[0069] Through the above schedule queue and the above computing instance, the above agent has queue and batch processing capabilities. That is, each inference request in an inference request set (also called a batch of inference requests) can be scheduled to each computing instance through the schedule queue, so that each computing instance executes each inference request, which enables different inference requests in the inference request set to be executed in parallel.

[0070] In addition, the above agent can also include two components, namely a memory predictor (MemoryPredictor) and a dynamic memory management component (Dynamic Memory Manager). Among them, the memory predictor can continuously execute the memory consumption prediction algorithm and dynamically manage the GPU memory in combination with the dynamic memory management component.

[0071] For each computing instance, two components can be included inside the computing instance, namely, a KV Cache component for caching intermediate data and results of inference to avoid repeated calculations, and an Executor component responsible for calling the large model to execute the inference task corresponding to the inference request.

[0072] Please, on the basis of Figure 1 and Figure 2 , refer to Figure 3 . Figure 3 FIG.

[0073] In this embodiment, the above-mentioned memory management method of the inference system can be applied to an inference engine as shown in Figure 1 .

[0074] As shown in Figure 3 , the above-mentioned memory management method of the inference system may include the following steps:

[0075] Step 302: Determine a memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue.

[0076] In this embodiment, after scheduling each inference request in a set of inference requests to each computing instance for execution through the above-mentioned scheduling queue, the execution status of each inference request in the set of inference requests maintained by the scheduling queue can be updated from the unexecuted state to the executing state. For the set of inference requests being executed in the scheduling queue, a memory management time window can be determined according to the data processing duration associated with the set of inference requests. Subsequently, the memory management time window can be used as the time unit for the GPU memory management task, that is, the GPU memory management is executed once within the memory management time window.

[0077] In some embodiments, in order to improve the accuracy and adaptability of the determined memory management time window, when determining the memory management time window according to the data processing duration associated with the set of inference requests being executed in the above-mentioned scheduling queue, the corresponding data processing duration can be first determined through the execution stages of each inference request in the set of inference requests, and then the memory management time window can be determined according to the data processing duration.

[0078] It should be noted that the inference process may include multiple stages in some cases. For example, for an inference system based on a large model, the inference process of the large model (especially the generative model in natural language processing) may include a Prefill stage and a Decode stage.

[0079] During the inference process of large models, the Prefill stage and the Decode stage refer to different steps in text generation.

[0080] The Prefill stage usually occurs in the initial stage of text generation. Its main purpose is to provide the model with an initial context so that the model can better understand and generate subsequent text. In the Prefill stage, the model usually generates an initial text segment, which contains the starting part of the generated sequence. This stage may involve using some predefined strategies or algorithms to select the most appropriate starting words or phrases. For example, when dealing with a conditional generation task, the model may generate a part of the text as a basis based on the input conditions (such as: abstract generation, question answering, dialogue, etc.).

[0081] The Decode stage is the main stage of text generation. In this stage, the model gradually generates new words or characters until a complete output sequence is generated. In the Decode stage, the model generates new content word by word or character by character based on the existing text. In each iteration, the model predicts the next most likely word and adds it to the current text sequence. This process continues until a preset end condition is reached (such as: reaching the maximum text length, generating a text end marker, etc.). The Decode stage may use various strategies to optimize the quality of generation, such as techniques like Top-k sampling, Top-p (also known as Nucleus Sampling) sampling, etc., which can help the model avoid generating overly ordinary or meaningless content.

[0082] In practical applications, the Prefill and Decode stages are closely connected. The Prefill stage provides the initial text basis, while the Decode stage is responsible for gradually expanding on this basis to generate a complete text output. These two stages together determine the quality and coherence of the finally generated text.

[0083] Specifically, it is possible to first determine whether there is an inference request in the Prefill stage in the set of inference requests being executed in the above scheduling queue.

[0084] If there is an inference request in the Prefill stage, then the number of tokens included in each inference request in the above set of inference requests can be compared to determine the maximum number among them, which is the maximum token number corresponding to the inference requests in this set of inference requests. Thus, the token processing duration corresponding to this maximum token number can be determined as the length of the above memory management time window.

[0085] If there is no inference request in the Prefill stage, it means that all inference requests in the above inference request set are in the Decode stage. In this case, the token generation durations corresponding to each inference request in the inference request set can be compared, where the token generation duration is the duration for generating one token based on the Prompt included in the inference request, to determine the longest token generation duration corresponding to the inference requests in the inference request set, so that the longest token generation duration can be determined as the length of the above memory management time window.

[0086] Step 304: Calculate the GPU memory requirement corresponding to the inference request set within the memory management time window, and allocate GPU memory for the inference request set according to the GPU memory requirement.

[0087] In this embodiment, when the above memory time window is determined, the GPU memory requirement corresponding to the above inference request set within the memory management time window can be calculated, so that GPU memory can be allocated for the inference request set according to the GPU memory requirement. In this way, since GPU memory can be allocated according to the actual GPU memory requirement during the process of the inference engine executing the inference requests in the inference request set, it can not only ensure that there is enough GPU memory for the inference engine to use during the execution of the inference requests, but also avoid waste of GPU memory.

[0088] In practical applications, the above memory predictor can execute a memory consumption prediction algorithm, that is, determine the memory management time window according to the data processing duration associated with the inference request set being executed in the above scheduling queue, and calculate the GPU memory requirement corresponding to the inference request set within the memory management time window, while the above dynamic memory management component performs specific GPU memory allocation, that is, allocate GPU memory for the inference request set according to the GPU memory requirement.

[0089] Step 306: At the end of the memory management time window, re-determine the next memory management time window corresponding to the memory management time window according to the data processing duration associated with the inference request set being executed in the scheduling queue.

[0090] In this embodiment, after the execution of GPU memory management within the above memory management time window is completed, it can wait for the end of the memory management time window. At the end of the memory management time window, the next memory management time window corresponding to the memory management time window can be re-determined according to the data processing duration associated with the inference request set being executed in the above scheduling queue.

[0091] Specifically, assuming that the above-mentioned memory management time window is the T-th memory management time window determined according to the data processing duration associated with the set of inference requests being executed in the above-mentioned scheduling queue, the GPU memory requirement corresponding to the set of inference requests within the T-th memory management time window can be calculated, and GPU memory can be allocated to the set of inference requests according to the GPU memory requirement; at the end of the T-th memory management time window, the (T + 1)-th memory management time window can be determined again according to the data processing duration associated with the set of inference requests being executed in the scheduling queue, so that the GPU memory requirement corresponding to the set of inference requests within the (T + 1)-th memory management time window can be calculated, and GPU memory can be allocated to the set of inference requests according to the GPU memory requirement; and so on.

[0092] In some embodiments, when calculating the GPU memory requirement corresponding to the set of inference requests within the above-mentioned memory management time window, specifically, the static GPU memory usage corresponding to the set of inference requests within the memory management time window can be calculated first, then the maximum GPU memory usage within the previous memory management time window corresponding to the memory management time window, and the sum of the static GPU memory usages corresponding to the set of inference requests within each target memory management time window are determined, and the difference between the maximum GPU memory usage and the sum of the static GPU memory usages is determined as the dynamic GPU memory usage corresponding to the set of inference requests within the memory management time window, where the target memory management time window is the memory management time window before the memory management time window during the execution of the set of inference requests, and finally, the sum of the static GPU memory usage and the dynamic GPU memory usage is determined as the GPU memory requirement corresponding to the set of inference requests within the memory management time window.

[0093] Specifically, assume that the above memory management time window is the T-th memory management time window determined according to the data processing duration associated with the set of inference requests being executed in the above scheduling queue. Then, the previous memory management time window corresponding to the T-th memory management time window is the (T - 1)-th memory management time window, and the target memory management time windows before the T-th memory management time window during the execution of this set of inference requests are the 1st to (T - 1)-th memory management time windows. In this case, let StaticResource[T] represent the static GPU memory usage within the T-th memory management time window, DynamicResource[T] represent the dynamic GPU memory usage within the T-th memory management time window, MaxResource[T] represent the maximum GPU memory usage within the T-th memory management time window, and NeedResource[T] represent the GPU memory demand within the T-th memory management time window. Then, there are the following formulas:

[0094] DynamicResource[T] = MaxResource[T - 1] – (StaticResource[0] + StaticResource[1] + … + StaticResource[T - 1]);

[0095] NeedResource[T] = StaticResource[T] + DynamicResource[T].

[0096] That is to say, for the current above-mentioned memory management time window, the dynamic GPU memory usage within this memory management time window can be predicted based on the maximum GPU memory usage within the previous memory management time window corresponding to this memory management time window and the sum of the static GPU memory usage within all memory management time windows before this memory management time window.

[0097] It should be noted that the above-mentioned static GPU memory usage refers to the GPU memory that can be foreseen to be used during the inference process. For example: the GPU memory usage required for the tokens included in the Prompt, the GPU memory usage required for the new tokens generated based on the Prompt, etc.; the above-mentioned dynamic GPU memory usage refers to the GPU memory that cannot be foreseen to be used during the inference process. For example: the GPU memory usage required for the uncertain intermediate data generated during the inference process of the large model.

[0098] In practical applications, to improve the compatibility of the technical solution in this application, when the above-mentioned inference engine is started, a sufficient amount of virtual GPU memory can be reserved for the inference engine, and then the above-mentioned memory management method can be adopted during the subsequent inference process to allocate physical GPU memory by calling the API provided by CUDA (Compute Unified Device Architecture).

[0099] In some embodiments, to improve the accuracy and adaptability of the determined static GPU memory usage, when calculating the static GPU memory usage corresponding to the above-mentioned inference request set within the above-mentioned memory management time window, the static GPU memory usage within the memory management time window can be determined according to the execution stages of the respective inference requests in the inference request set.

[0100] Specifically, it can first be determined whether there is an inference request in the inference request set being executed in the above-mentioned scheduling queue that is in the Prefill stage.

[0101] If there is an inference request in the Prefill stage, the GPU memory usage required for the tokens corresponding to the respective inference requests in the above-mentioned inference request set (which can be referred to as the first GPU memory usage) and the GPU memory usage required for generating a token based on the respective inference requests in the inference request set (which can be referred to as the second GPU memory usage) can be determined, and the sum of the first GPU memory usage and the second GPU memory usage can be determined as the static GPU memory usage corresponding to the inference request set within the above-mentioned memory management time window. Among them, the GPU memory usage required for the tokens corresponding to an inference request can be the GPU memory usage required for the tokens included in the Prompt in the inference request; the GPU memory usage required for generating a token based on an inference request can be the GPU memory usage required for generating a token based on the Prompt included in the inference request.

[0102] If there is no inference request in the Prefill stage, it means that all the inference requests in the above-mentioned inference request set are in the Decode stage. In this case, the GPU memory usage required for generating a token based on the respective inference requests in the inference request set can be determined, and this GPU memory usage can be determined as the static GPU memory usage corresponding to the inference request set within the memory management time window.

[0103] In some embodiments, to ensure the rationality and availability of GPU memory allocation, when allocating GPU memory for the above-mentioned inference request set according to the above-mentioned GPU memory demand, specifically, the current GPU memory usage corresponding to the inference request set can be obtained first. Then, based on the current GPU memory usage and the GPU memory demand, it can be determined whether the GPU memory allocation condition and the GPU memory release condition are met.

[0104] It should be noted that the above-mentioned GPU memory allocation condition and the above-mentioned GPU memory release condition are two mutually exclusive conditions. That is, if the GPU memory allocation condition is met, it means that the GPU memory release condition is not met; if the GPU release allocation condition is met, it means that the GPU memory allocation condition is not met.

[0105] If the above-mentioned GPU memory allocation condition is met, GPU memory can be allocated for the above-mentioned inference request set according to the above-mentioned GPU memory demand.

[0106] If the above-mentioned GPU memory release condition is met, the GPU memory used to execute the above-mentioned inference request set can be released, and the above-mentioned current GPU memory usage can be updated. In the case where it is determined that the above-mentioned GPU memory allocation condition is met based on the updated current GPU memory usage and the above-mentioned GPU memory demand, GPU memory can be allocated for the inference request set according to the GPU memory demand.

[0107] It should be noted that the amount of released GPU memory usage can be determined according to the executed GPU memory release task, and the released GPU memory usage can be subtracted from the above-mentioned current GPU memory usage to obtain the updated current GPU memory usage. However, in practical applications, to obtain a more accurate current GPU memory usage, the current GPU memory usage can be obtained in real time or periodically through tools such as DCGM (Data Center GPU Manager). DCGM is a tool for managing and monitoring GPU resources in a data center. It can help effectively manage large-scale GPU clusters and provide in-depth insights into GPU performance, health, and utilization.

[0108] In some embodiments, the above-mentioned memory allocation condition can be that the sum of the current GPU memory usage and the GPU memory demand is less than a preset threshold (which can be referred to as the first threshold); the above-mentioned GPU memory release condition can be that the current GPU memory usage is greater than a preset threshold (which can be referred to as the second threshold). Among them, the specific values of the first threshold and the second threshold can be the same or different, and the present application does not impose special restrictions on this.

[0109] Specifically, let NeedResource[T] represent the GPU memory demand within the T-th memory management time window, CurrentResource[T] represent the current GPU memory usage obtained when performing GPU memory management within the T-th memory management time window, Threshold1 represent the above-mentioned first threshold, and Threshold2 represent the above-mentioned second threshold. Then, it can be determined that the above-mentioned memory allocation condition is satisfied when CurrentResource[T] + NeedResource[T] < Threshold1, and it can be determined that the above-mentioned memory release condition is satisfied when CurrentResource[T] > Threshold2.

[0110] In practical applications, the period when CurrentResource[T] + NeedResource[T] > Threshold1 and CurrentResource[T] < Threshold2 can be used as the buffer time for releasing GPU memory. If during this period, CurrentResource[T] changes by itself, such that the updated CurrentResource[T] + NeedResource[T] < Threshold1, then since the above-mentioned memory allocation condition is satisfied, GPU memory can be allocated to the above-mentioned set of inference requests according to the above-mentioned GPU memory demand. If during this period, CurrentResource[T] changes by itself, such that CurrentResource[T] > Threshold2, then since the above-mentioned memory release condition is satisfied, the GPU memory used to execute the above-mentioned set of inference requests can be released. In this way, the number of times of executing the GPU memory release task can be reduced to a certain extent.

[0111] In some embodiments, various methods can be adopted to implement the release of GPU memory. For example, when releasing the GPU memory used to execute the above-mentioned set of inference requests, the specific GPU memory release method can be selected according to the number of inference requests in the set of inference requests.

[0112] Specifically, it can first be determined whether the number of inference requests in the above-mentioned set of inference requests is greater than a preset threshold (which can be referred to as the third threshold).

[0113] If the number of inference requests in the above-mentioned set of inference requests is greater than the above-mentioned third threshold, then the tokens corresponding to each inference request in the set of inference requests can be offloaded from the GPU memory to the CPU memory to release the GPU memory occupied by these tokens.

[0114] If the number of inference requests in the above-mentioned set of inference requests is not greater than the above-mentioned third threshold, the inference requests can be updated based on the tokens corresponding to each inference request in the set of inference requests to release the GPU memory occupied by these tokens. Among them, updating an inference request based on the token corresponding to an inference request means using the tokens generated based on the Prompt included in the inference request to update the Prompt to form a new Prompt.

[0115] In the above technical solution, the inference engine in the inference system can use the GPU installed on the computing device where it is located as the computing resource to execute inference requests and manage the GPU memory. Specifically, the inference engine can maintain a scheduling queue for scheduling the set of inference requests, and determine the memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue. Subsequently, the GPU memory requirement corresponding to the set of inference requests within this memory management time window can be calculated, and the GPU memory can be allocated to the set of inference requests according to this GPU memory requirement. After the end of this memory management time window, the next memory management time window corresponding to this memory management time window can be determined again according to the data processing duration associated with the set of inference requests being executed in the scheduling queue, so as to perform GPU memory management within the next memory management time window.

[0116] By adopting the above method, there is no need to reserve a large amount of GPU memory for the inference engine when it starts. Instead, during the process of the inference engine batch-executing inference requests, the memory management time window can be continuously set, and the GPU memory requirement corresponding to the set of inference requests being executed within this memory management time window can be predicted, so as to allocate the GPU memory to the set of inference requests according to this GPU memory requirement. This can not only ensure that the inference engine has sufficient GPU memory to use during the process of batch-executing inference requests, but also avoid waste of GPU memory.

[0117] Corresponding to the embodiment of the foregoing method, the present application also provides an embodiment of a device.

[0118] Please refer to Figure 4 , Figure 4It is a schematic structural diagram of a device shown in an exemplary embodiment of the present application. At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, a memory 408, and a non-volatile memory 410. Of course, it may also include other required hardware. One or more embodiments of the present application can be implemented in a software manner. For example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into the memory 408 and then runs it. Of course, in addition to the software implementation, one or more embodiments of the present application do not exclude other implementation manners, such as a logic device or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic module, and can also be hardware or a logic device.

[0119] Please refer to Figure 5 , Figure 5 It is a block diagram of a memory management device of an inference system shown in an exemplary embodiment of the present application.

[0120] The above-mentioned memory management device of the inference system can be applied to a device as shown in Figure 4 to implement the technical solution of the present application. Among them, an inference engine in the inference system can be deployed on the device, and the memory management device of the inference system can be specifically applied to the inference engine; the computing resources of the inference engine include the GPU installed on the computing device for deploying the inference engine; the inference engine maintains a scheduling queue for scheduling a set of inference requests; the device includes:

[0121] A time window determination module 502 determines a memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue;

[0122] A GPU memory allocation module 504 calculates the GPU memory demand corresponding to the set of inference requests within the memory management time window, and allocates GPU memory for the set of inference requests according to the GPU memory demand;

[0123] The time window determination module 502 is further configured to, at the end of the memory management time window, re-determine the next memory management time window corresponding to the memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue.

[0124] In some embodiments, the determining the memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue includes:

[0125] In the set of inference requests being executed in the scheduling queue, it is determined whether there is an inference request in the Prefill stage;

[0126] If there is an inference request in the Prefill stage, determine the maximum number of tokens corresponding to the inference requests in the set of inference requests, and determine the memory management time window according to the token processing duration corresponding to the maximum number of tokens;

[0127] If there is no inference request in the Prefill stage, determine the longest token generation duration corresponding to the inference requests in the set of inference requests, and determine the memory management time window according to the longest token generation duration.

[0128] In some embodiments, calculating the GPU memory requirement corresponding to the set of inference requests within the memory management time window includes:

[0129] Calculate the static GPU memory usage corresponding to the set of inference requests within the memory management time window;

[0130] Determine the maximum GPU memory usage within the previous memory management time window corresponding to the memory management time window, and the sum of the static GPU memory usages corresponding to the set of inference requests within each target memory management time window, and determine the difference between the maximum GPU memory usage and the sum of the static GPU memory usages as the dynamic GPU memory usage corresponding to the set of inference requests within the memory management time window; wherein, the target memory management time window is the memory management time window before the memory management time window during the execution of the set of inference requests;

[0131] Determine the sum of the static GPU memory usage and the dynamic GPU memory usage as the GPU memory requirement corresponding to the set of inference requests within the memory management time window.

[0132] In some embodiments, calculating the static GPU memory usage corresponding to the set of inference requests within the memory management time window includes:

[0133] In the set of inference requests being executed in the scheduling queue, determine whether there is an inference request in the Prefill stage;

[0134] If there is an inference request in the Prefill stage, determine the first GPU memory usage required for the tokens corresponding to each inference request in the set of inference requests, and the second GPU memory usage required for generating one token based on each inference request in the set of inference requests, and determine the sum of the first GPU memory usage and the second GPU memory usage as the static GPU memory usage corresponding to the set of inference requests within the memory management time window;

[0135] If there is no inference request in the Prefill stage, determine the GPU memory usage required to generate a token based on each inference request in the set of inference requests, and determine the GPU memory usage as the static GPU memory usage corresponding to the set of inference requests within the memory management time window.

[0136] In some embodiments, allocating GPU memory for the set of inference requests according to the GPU memory demand includes:

[0137] Obtain the current GPU memory usage corresponding to the set of inference requests;

[0138] Based on the current GPU memory usage and the GPU memory demand, determine whether the GPU memory allocation condition and the GPU memory release condition are satisfied;

[0139] If the GPU memory allocation condition is satisfied, allocate GPU memory for the set of inference requests according to the GPU memory demand;

[0140] If the GPU memory release condition is satisfied, release the GPU memory used to execute the set of inference requests, and update the current GPU memory usage. When it is determined that the GPU memory allocation condition is satisfied based on the updated current GPU memory usage and the GPU memory demand, allocate GPU memory for the set of inference requests according to the GPU memory demand.

[0141] In some embodiments, the GPU memory allocation condition includes:

[0142] The sum of the current GPU memory usage and the GPU memory demand is less than a preset first threshold;

[0143] The GPU memory release condition includes:

[0144] The current GPU memory usage is greater than a preset second threshold.

[0145] In some embodiments, releasing the GPU memory used to execute the set of inference requests includes:

[0146] Determine whether the number of inference requests in the set of inference requests is greater than a preset third threshold;

[0147] If the number is greater than the third threshold, unload the tokens corresponding to each inference request in the set of inference requests to the CPU memory to release the GPU memory occupied by the tokens;

[0148] If the quantity is not greater than the third threshold, update the inference requests based on the tokens corresponding to the respective inference requests in the set of inference requests to release the GPU memory occupied by the tokens.

[0149] For the apparatus embodiments, they basically correspond to the method embodiments. Therefore, for the relevant parts, refer to the partial descriptions of the method embodiments. The apparatus embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the technical solution of this application.

[0150] The systems, apparatuses, modules or units illustrated in the above embodiments may be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer may be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0151] In a typical configuration, a computer includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0152] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0153] Computer readable media include permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0154] It should be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0155] The above specific embodiments of the present application are described. Other embodiments are within the scope of the present application. In some cases, the actions or steps recorded in the present application can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the process depicted in the accompanying drawings does not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0156] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present application. The singular forms of "a", "said" and "the" are also intended to include plural forms, unless the context clearly indicates other meanings. The term "and / or" refers to and includes any or all possible combinations of one or more associated listed items.

[0157] The description of terms such as "one embodiment", "some embodiments", "example", "specific example", or "a kind of implementation manner" used in one or more embodiments of this application means that the specific features or characteristics described in connection with the embodiment are included in at least one embodiment of this application. The schematic descriptions of these terms do not necessarily refer to the same embodiment. Moreover, the specific features or characteristics described can be combined in a suitable manner in one or more embodiments of this application. In addition, without contradiction, different embodiments and the specific features or characteristics in different embodiments can be combined.

[0158] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to a determination".

[0159] The above description is only the preferred embodiment of one or more embodiments of this application, and is not intended to limit one or more embodiments of this application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of this application shall be included within the scope protected by one or more embodiments of this application.

[0160] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

Claims

1. A memory management method for an inference system, applied to an inference engine in the inference system; the computing resources of the inference engine include a GPU mounted on a computing device for deploying the inference engine; The inference engine maintains a scheduling queue for scheduling a set of inference requests; the method comprises: Determining a memory management time window according to a data processing duration associated with a set of inference requests being executed in the scheduling queue; Calculating a static GPU memory usage corresponding to the inference request set within the memory management time window; Determine a maximum GPU memory usage in a previous memory management time window corresponding to the memory management time window, and a sum of static GPU memory usages corresponding to the inference request set in each target memory management time window, and determine a difference between the maximum GPU memory usage and the sum of the static GPU memory usages as a dynamic GPU memory usage corresponding to the inference request set in the memory management time window; wherein the target memory management time window is a memory management time window located before the memory management time window during the execution of the inference request set; Determine the sum of the static GPU memory usage and the dynamic GPU memory usage as the GPU memory demand corresponding to the inference request set within the memory management time window, and allocate GPU memory to the inference request set according to the GPU memory demand; When the memory management time window ends, a subsequent memory management time window corresponding to the memory management time window is determined again based on the data processing duration associated with the inference request set being executed in the scheduling queue.

2. The method according to claim 1, wherein determining the memory management time window according to the data processing duration associated with the set of inference requests being executed in the scheduling queue comprises: In the set of inference requests being executed in the scheduling queue, determining whether there is an inference request in the Prefill stage; If there is an inference request in the Prefill phase, determine a maximum number of tokens corresponding to the inference request in the inference request set, and determine a memory management time window according to a token processing duration corresponding to the maximum number of tokens; If there is no inference request in the Prefill phase, a longest token generation duration corresponding to the inference requests in the inference request set is determined, and a memory management time window is determined according to the longest token generation duration.

3. The method according to claim 1, wherein the step of calculating the static GPU memory usage corresponding to the inference request set within the memory management time window comprises: In the set of inference requests being executed in the scheduling queue, determining whether there is an inference request in the Prefill stage; If there are inference requests in the Prefill stage, determine a first GPU memory usage required for a token corresponding to each inference request in the inference request set, and a second GPU memory usage required to generate a token based on each inference request in the inference request set, and determine the sum of the first GPU memory usage and the second GPU memory usage as a static GPU memory usage corresponding to the inference request set within the memory management time window; If there is no inference request in the Prefill stage, determine the GPU memory usage required to generate a token based on each inference request in the inference request set, and determine the GPU memory usage as the static GPU memory usage corresponding to the inference request set within the memory management time window.

4. The method according to claim 1, allocating GPU memory to the inference request set according to the GPU memory requirement, comprising: Obtain the current GPU memory usage corresponding to the inference request set; Determine whether a GPU memory allocation condition and a GPU memory release condition are met according to the current GPU memory usage and the GPU memory demand; If the GPU memory allocation condition is met, allocating GPU memory to the inference request set according to the GPU memory requirement; If the GPU memory release condition is met, the GPU memory used to execute the inference request set is released, and the current GPU memory usage is updated, so that when it is determined that the GPU memory allocation condition is met based on the updated current GPU memory usage and the GPU memory requirement, GPU memory is allocated to the inference request set according to the GPU memory requirement.

5. The method according to claim 4, wherein the GPU memory allocation condition comprises: The sum of the current GPU memory usage and the GPU memory requirement is less than a preset first threshold; The GPU memory release conditions include: The current GPU memory usage is greater than a preset second threshold.

6. The method according to claim 4, wherein releasing GPU memory used to execute the set of inference requests comprises: Determining whether the number of inference requests in the inference request set is greater than a preset third threshold; If the number is greater than the third threshold, unloading the token corresponding to each inference request in the inference request set to the CPU memory to release the GPU memory occupied by the token; If the number is not greater than the third threshold, based on the tokens corresponding to the respective inference requests in the inference request set, the inference requests are updated to release the GPU memory occupied by the tokens.

7. A memory management device for an inference system, applied to an inference engine in the inference system; the computing resources of the inference engine include a GPU mounted on a computing device for deploying the inference engine; The inference engine maintains a scheduling queue for scheduling an inference request set; the device comprises: A time window determination module determines a memory management time window according to a data processing duration associated with a set of inference requests being executed in the scheduling queue; A GPU memory allocation module is configured to calculate a static GPU memory usage corresponding to the inference request set within the memory management time window; determine a maximum GPU memory usage within a previous memory management time window corresponding to the memory management time window, and a sum of static GPU memory usages corresponding to the inference request set within each target memory management time window, and determine a difference between the maximum GPU memory usage and the sum of the static GPU memory usages as a dynamic GPU memory usage corresponding to the inference request set within the memory management time window; wherein the target memory management time window is a memory management time window located before the memory management time window during the execution of the inference request set; determine the sum of the static GPU memory usage and the dynamic GPU memory usage as a GPU memory demand corresponding to the inference request set within the memory management time window, and allocate GPU memory to the inference request set according to the GPU memory demand; The time window determination module is also used to determine, when the memory management time window ends, a subsequent memory management time window corresponding to the memory management time window based on the data processing duration associated with the set of inference requests being executed in the scheduling queue.

8. An electronic device, comprising: processor; a memory for storing processor-executable instructions; The processor implements the method according to any one of claims 1 to 6 by running the executable instructions.

9. A computer-readable storage medium having computer instructions stored thereon, wherein the instructions are executed by a processor to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Scheduling engine, scheduling method, electronic device, storage medium and program product

    CN118885305A