Load-Aware Scheduling Method for Inference System and Inference System

By introducing a global scheduler into the inference system, scheduling inference requests based on GPU load information, the efficient scheduling problem of inference requests in the computing cluster is solved, GPU resource usage is optimized, delay is reduced and system throughput is improved.

CN119149252BActive Publication Date: 2025-07-18ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411646359.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-07-18
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

When deploying an inference system on a computing cluster, how to effectively schedule inference requests to optimize GPU resource usage efficiency and reduce latency, especially in high concurrency and high-performance computing scenarios, it is difficult for the prior art to realize load-aware scheduling.

Method used

A global scheduler is introduced to maintain the dynamically updated GPU load information of each computing instance. Based on this information, the inference request is scheduled to the target computing instance that meets the preset conditions for execution, and load-aware scheduling is realized.

Benefits of technology

Optimizes the efficiency of GPU resource usage, reduces the waiting time and processing delay of inference requests, improves the overall throughput of the inference system, and is versatile and scalable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119149252B_ABST
    Figure CN119149252B_ABST
Patent Text Reader

Abstract

One or more embodiments of the present application provide a load-aware scheduling method for an inference system and an inference system. The method is applied to a global scheduler in the inference system. The inference system further includes an inference engine. The inference engine includes at least one computing instance deployed on each computing node in a computing cluster. The computing resources of the computing instance include the GPUs installed on the computing node where it is located. The global scheduler maintains dynamically updated GPU load information of each computing instance. The method includes: obtaining a target inference request to be executed; determining a target computing instance whose GPU load meets a preset condition based on the maintained GPU load information of each computing instance; and sending the target inference request to the target computing instance so that the target computing instance executes the target inference request.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present application relate to the field of artificial intelligence technology, and particularly to a load-aware scheduling method for an inference system and an inference system. Background Art

[0002] An inference system is a computer program that uses logical rules and known facts to draw new conclusions or make decisions. The inference system is an important part of the field of artificial intelligence and is mainly used to simulate the human decision-making process. It derives conclusions based on a set of defined knowledge bases and inference engines. The inference system can execute the inference requests it obtains and output corresponding inference results.

[0003] A typical inference system usually consists of the following parts: a knowledge base, an inference engine, a user interface, and an explanation facility. Among them, the knowledge base includes all the facts and rules known to the system. These facts can be about the state of the world, object attributes, etc., and the rules are logical expressions describing how to draw new conclusions from known facts. The inference engine is the core component of the inference system. It is responsible for performing logical operations in the inference process, that is, drawing new conclusions or making decisions from the given knowledge base; the inference engine uses a series of rules and known facts to derive new knowledge, thereby helping the system solve problems or make decisions. The user interface allows users to interact with the system, input queries or observe the results of the inference process. The explanation facility is used to explain how the system arrives at a specific conclusion, which is very important for transparency and trust.

[0004] In cases where large-scale data, high-concurrency requests, or high-performance computing are required, the inference engine is usually deployed on a computing cluster. By deploying the inference engine on a computing cluster, higher computing power, better fault tolerance, and more flexible resource management can be achieved. However, this introduces the problem of how to schedule inference requests on the inference engine at the cluster level. Summary of the Invention

[0005] One or more embodiments of the present application provide the following technical solutions:

[0006] The present application provides a load-aware scheduling method for an inference system, which is applied to a global scheduler in the inference system; the inference system further includes an inference engine; the inference engine includes at least one computing instance deployed on each computing node in a computing cluster; the computing resources of the computing instance include the GPUs installed on the computing nodes where they are located; the global scheduler maintains dynamically updated GPU load information of each computing instance.

[0007] The method includes:

[0008] Obtain a target inference request to be executed;

[0009] Based on the maintained GPU load information of each computing instance, determine a target computing instance whose GPU load meets a preset condition;

[0010] Send the target inference request to the target computing instance for the target computing instance to execute the target inference request.

[0011] The present application further provides an inference system, which includes a global scheduler and an inference engine; the inference engine includes at least one computing instance deployed on each computing node in a computing cluster; the computing resources of the computing instance include the GPUs installed on the computing nodes where they are located; the global scheduler maintains dynamically updated GPU load information of each computing instance.

[0012] The global scheduler is configured to:

[0013] Obtain a target inference request to be executed;

[0014] Based on the maintained GPU load information of each computing instance, determine a target computing instance whose GPU load meets a preset condition;

[0015] Send the target inference request to the target computing instance for the target computing instance to execute the target inference request.

[0016] The present application further provides an electronic device, including:

[0017] A processor;

[0018] A memory for storing executable instructions of the processor;

[0019] Wherein, the processor runs the executable instructions to implement the steps of the method as described in any one of the above.

[0020] The present application further provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any one of the above are implemented.

[0021] In the above technical solution, the inference engine in the inference system may include at least one computing instance deployed on each computing node in the computing cluster. The computing resources of each computing instance may include the GPUs installed on the corresponding nodes. The global scheduler in the inference system may maintain the dynamically updated GPU load information of each computing instance, and when obtaining a target inference request to be executed, first determine a target computing instance whose GPU load meets the preset conditions based on the maintained GPU load information of each computing instance, and then send the target inference request to the target computing instance for the target computing instance to execute the target inference request, thereby implementing the scheduling for the target inference request.

[0022] By adopting the above method, by adding a global scheduler in the inference system, the global scheduler can schedule the inference request to a computing instance with appropriate GPU load in the inference engine for execution, realizing load-aware scheduling of the inference request on the inference engine at the cluster level. Thereby, the utilization efficiency of GPU resources can be optimized, the waiting time and processing delay for the execution of the inference request can be reduced, and the overall throughput of the inference system can be improved. In addition, by scheduling the inference request to the inference engine through the global scheduler instead of directly scheduling the inference request within the inference engine, the scheduling method can be decoupled from the specific inference engine, having certain generality and scalability. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings required for the description of the exemplary embodiments will be described below, where:

[0024] Figure 1 is a schematic diagram of an inference system shown in an exemplary embodiment of the present application.

[0025] Figure 2 is a schematic diagram of another inference system shown in an exemplary embodiment of the present application.

[0026] Figure 3 is a flowchart of a load-aware scheduling method for an inference system shown in an exemplary embodiment of the present application.

[0027] Figure 4 is a schematic structural diagram of a device shown in an exemplary embodiment of the present application.

[0028] Figure 5 is a block diagram of an inference system shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of the present application. On the contrary, they are merely examples consistent with some aspects of one or more embodiments of the present application.

[0030] It should be noted that in other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in the present application. In some other embodiments, the steps included in the method may be more or fewer than those described in the present application. In addition, a single step described in the present application may be decomposed into multiple steps for description in other embodiments; and multiple steps described in the present application may also be combined into a single step for description in other embodiments.

[0031] In the present application, the inference engine in the inference system can be deployed on a computing cluster to meet the requirements of processing large-scale data, high-concurrency requests, or high-performance computing, thereby achieving higher computing power, better fault tolerance, and more flexible resource management. In such an inference system, a load-aware scheduling method can be adopted to schedule inference requests on the inference engine at the cluster level.

[0032] Load-aware scheduling refers to a strategy of allocating tasks according to the current load conditions of each node or resource in a computing system. This scheduling method aims to optimize the resource utilization efficiency, reduce the waiting time and processing delay, and improve the overall throughput of the system.

[0033] Load-aware scheduling is usually applied to scenarios such as distributed systems, cloud computing platforms, and data centers, aiming to achieve dynamic balance of the workloads of each node and avoid the situation where some nodes are overloaded while others are idle. By real-time monitoring the load conditions of each node, the load-aware scheduler can make more reasonable task allocation decisions.

[0034] Generally, the inference engine can use the CPU as the computing resource, and the CPU can be the CPU installed on the device for deploying the inference engine. The load metrics of the CPU usually include the CPU usage rate and the memory usage rate. That is, for the inference engine deployed on the computing cluster, by calculating the CPU usage rate and the memory usage rate to represent the load condition of the CPU, load-aware scheduling of inference requests can be realized on such an inference engine.

[0035] With the development of artificial intelligence technology, inference systems based on large models (e.g., Large Language Model) are being applied more and more widely.

[0036] Large models refer to machine learning models with a large number of parameters, such as various variants under the Transformer architecture, including but not limited to natural language processing models like GPT and BERT. These models achieve powerful representation learning capabilities through a large amount of training data and complex architectures.

[0037] In an inference system based on a large model, the large model can be regarded as a huge knowledge base that contains information learned from a large amount of data. The large model learns a large number of patterns and features during the training process, and these patterns and features represent the complex relationships in the data. Therefore, to some extent, the information stored inside the large model can be regarded as a form of knowledge.

[0038] The inference engine designed for the large model is responsible for calling the large model to perform specific inference tasks, that is, it can manage the loading of the large model, execute the inference operations of the large model, and manage the interaction with hardware (e.g., CPU, GPU, or other accelerators). The inference engine can also include optimization algorithms to improve the speed and efficiency of inference.

[0039] In practical applications, the large model can be deployed separately from the inference system, or the large model can be integrated into the inference system and use the inference engine therein to call the large model to efficiently execute inference tasks.

[0040] Since large models are artificial intelligence applications that require a large amount of computing resources, when the inference engine calls the large model to execute inference tasks, GPUs are usually used as computing resources.

[0041] GPUs are designed specifically for parallel processing and can process multiple data points simultaneously. This is especially useful for deep learning models because they often need to perform the same operations on a large amount of data. Modern AI models, especially neural network models, require a large number of floating-point operations. GPUs usually have better performance in floating-point operations than CPUs, especially when dealing with large-scale matrix operations, which are common operations in deep learning. GPUs usually have a higher memory bandwidth than CPUs, which means they can read and write data from memory faster. This is very important for applications that need to access a large amount of data frequently. Many GPU vendors (e.g., NVIDIA) have optimized their hardware and software stacks for machine learning tasks. For example, they provide hardware specifically designed to accelerate tensor operations, as well as programming models like CUDA to fully utilize these hardware features. GPUs can improve the inference speed. For applications deployed on a large scale, using GPUs can reduce latency and improve the response speed of the system.

[0042] There are many choices for the load metrics of the GPU. Commonly used metrics include SM Clock (Streaming Multiprocessor Clock), SM Activity (Streaming Multiprocessor Activity), Memory Utilization (Memory Usage), etc. Among them, SM Clock refers to the clock frequency of the streaming multiprocessors on the GPU; the streaming multiprocessors are the computing units on the GPU and are responsible for executing parallel computing tasks. A higher SM Clock frequency means that the computing units of the GPU run faster, which usually leads to higher computing performance but may also increase power consumption and heat. SM Activity refers to the degree of activity of the streaming multiprocessors on the GPU, that is, what proportion of the time these computing units are actually executing tasks within a period of time. If the SM Activity is high, it means that most of the computing units of the GPU are working actively; conversely, if it is low, it may mean that there is a computing bottleneck or a long waiting time. Memory Utilization refers to the usage of the GPU video memory and usually represents the proportion of the currently used video memory in the total video memory. A higher Memory Utilization means that more video memory is occupied, which may lead to a slowdown in data exchange speed or affect performance due to insufficient video memory. When the video memory is close to full load, it may trigger data overflow to the host memory, which will further reduce performance.

[0043] For the inference engine deployed on the computing cluster, by calculating the load metrics of the GPU to represent the load situation of the GPU, it is possible to achieve load-aware scheduling of inference requests on such inference engines.

[0044] In the technical solution provided by one or more embodiments of the present application, the inference engine in the inference system may include at least one computing instance deployed on each computing node in the computing cluster. The computing resources of each computing instance may include the GPUs installed on the corresponding nodes. The global scheduler in the inference system can maintain the dynamically updated GPU load information of each computing instance, and in the case of obtaining a target inference request to be executed, first determine a target computing instance whose GPU load meets the preset conditions based on the maintained GPU load information of each computing instance, and then send the target inference request to the target computing instance so that the target computing instance executes the target inference request, thereby realizing the scheduling for the target inference request.

[0045] In the above - mentioned manner, by adding a global scheduler to the inference system, the global scheduler can schedule inference requests to the computing instances with appropriate GPU loads in the inference engine for execution. This realizes load - aware scheduling of inference requests on the inference engine at the cluster level, thereby optimizing the utilization efficiency of GPU resources, reducing the waiting time and processing latency of inference request execution, and improving the overall throughput of the inference system. In addition, by scheduling inference requests to the inference engine through the global scheduler instead of directly scheduling inference requests within the inference engine, the scheduling method can be decoupled from the specific inference engine, having a certain degree of generality and scalability.

[0046] Please refer to Figure 1 and Figure 2 , Figure 1 and Figure 2 which are respectively schematic diagrams of an inference system shown in an exemplary embodiment of the present application.

[0047] As Figure 1 shown, the above - mentioned inference system may include a global scheduler (Global Scheduler) and an inference engine. Among them, the inference engine may be deployed on a computing cluster (Cluster) composed of at least one computing node (Node).

[0048] A computing cluster is a system in which multiple computers or servers are connected through a network to work collaboratively. These nodes jointly complete computing tasks, providing stronger computing power and higher availability than a single computer. A computing node is one of the core components of a computing cluster, usually referring to a computer or server used to execute compute - intensive tasks. A computing node can undertake computing tasks, and the GPU installed on it can be used as computing resources. Among them, the computer or server can be a physical or virtual local computer or local server, or a physical or virtual cloud computer or cloud server.

[0049] In practical applications, the above - mentioned global scheduler can be deployed on the head node in the above - mentioned computing cluster, where the head node usually refers to the node responsible for the management and scheduling of the computing cluster, or can also be deployed on other computing devices outside the computing cluster. The present application does not limit this.

[0050] In the above - mentioned inference system, communication can be carried out between the above - mentioned global scheduler and the above - mentioned inference engine. For example, direct communication can be carried out between the two through the HTTP (HyperText Transfer Protocol) protocol.

[0051] Specifically, the above-mentioned global scheduler can communicate with the applications and components deployed on each computing node in the above-mentioned computing cluster. In this case, the global scheduler can obtain inference requests from the request queue, and based on the load conditions of the GPUs installed on each computing node, select a computing node with appropriate GPU load conditions, and schedule the obtained inference requests to this computing node to use the GPU installed on this computing node to execute the inference requests.

[0052] As Figure 2 shown, for each computing node in the above-mentioned computing cluster, at least one computing instance can be deployed on this computing node. Among them, a computing instance refers to dedicated resources allocated on demand through virtualization technology in a local data center or a cloud computing platform. The resources of these computing instances can include computing resources, storage resources, network resources, etc. (for example: GPUs, GPU memory, storage, network interfaces, etc.), enabling each computing instance to use these resources to execute specific computing tasks. Computing instances can be created, started, stopped, or deleted according to actual needs, and their resources can also be adjusted according to requirements.

[0053] In an inference engine for large models, the computing resources of a computing instance can be the GPUs of the computing node where it is located.

[0054] In practical applications, multiple GPUs can be installed on a computing node, and one GPU can be allocated to a computing instance as the computing resources of this computing instance. Alternatively, only one GPU can be installed on a computing node, and the computing resources allocated to a computing instance can be a part of the computing units, memory, etc. of this GPU.

[0055] In an inference engine for large models, each computing instance can call the large model to perform specific inference work, that is, the large model can be loaded into each computing instance, and each computing instance executes the inference operations of the large model. Specifically, the large model can be read from a local or cloud storage medium (for example: local disk, cloud storage, etc.) into the GPU memory of each computing instance to perform inference based on the large model on the GPU computing resources of each computing instance.

[0056] In practical applications, a Model Runtime framework can be installed in a computing instance. After reading a large model into the computing instance, a model service can be configured according to the Model Runtime framework and the large model, so that the computing instance can serve as an independently running model service instance. Here, the model runtime refers to the environment and framework used to perform model inference after the model is deployed. This environment usually includes a series of steps such as model loading, input data processing, model execution, and output result processing, covering the entire life cycle of the model from loading to execution and then to unloading.

[0057] By deploying computing instances on computing nodes, firstly, the resources of the computing nodes can be fully utilized to accelerate the execution of computing tasks, thereby improving the computing power of the inference engine; secondly, different computing instances are independent of each other, which can ensure the stability and security of computing tasks, and can achieve resource isolation and management, contributing to the effective utilization of resources and avoiding resource contention; thirdly, for computing tasks that need to be processed in parallel, the completion time of the computing tasks can be accelerated through parallel computing; fourthly, elastic scaling can be achieved. When a large number of computing tasks need to be processed, computing instances can be added, and when the number of computing tasks to be processed decreases, some computing instances can be terminated to release some resources, realizing flexible allocation of resources.

[0058] To facilitate load-aware scheduling of inference requests, the global scheduler can maintain the GPU load information of each computing instance, where the GPU load information can indicate the load condition of the GPU. In this case, the global scheduler can obtain an inference request from the request queue and, based on the maintained GPU load information of each computing instance, select a computing instance with an appropriate GPU load condition and schedule the obtained inference request to this computing instance to execute the inference request using the GPU that is the computing resource of this computing instance.

[0059] It should be noted that the GPU load information of each computing instance maintained by the above global scheduler can be dynamically updated. For example, after scheduling an inference request to a computing instance and the computing instance starts to execute the inference request, GPU load information can be generated according to the current GPU load condition of this computing instance, and the generated GPU load information can be reported to the global scheduler, and the global scheduler updates the originally maintained GPU load information of this computing instance based on this GPU load information.

[0060] For each computing node in the above computing cluster, to facilitate the maintenance, management, and scheduling of the computing instances deployed on this computing node, a Local Scheduler can also be deployed on this computing node.

[0061] The above-mentioned global scheduler can communicate with the local schedulers deployed on each computing node in the above-mentioned computing cluster. In this case, the global scheduler can obtain inference requests from the request queue, and based on the GPU load information of each computing instance it maintains, select a computing instance with an appropriate GPU load situation, and send the obtained inference request to the local scheduler deployed on the computing node where the computing instance is located. The local scheduler further schedules the inference request to the computing instance to execute the inference request using the GPU as the computing resource of the computing instance.

[0062] In practical applications, to ensure the reliability and correctness of inference request scheduling, for each local scheduler, it can maintain a sequential queue, and the sequential queue can be a first-in-first-out (FIFO) sequence of inference requests, that is, the inference requests in the sequential queue are processed in the order in which they enter the sequential queue.

[0063] In addition, for each computing node in the above-mentioned computing cluster, the local scheduler deployed on the computing node can monitor the GPU load situation of each computing instance deployed on the computing node, generate corresponding GPU load information, and report the generated GPU load information to the above-mentioned global scheduler, so that the global scheduler can maintain the dynamically updated GPU load information of each computing instance.

[0064] Furthermore, in practical applications, the above-mentioned global scheduler can internally include two components, namely a request router component responsible for routing inference requests, and a load manager component responsible for obtaining and maintaining the GPU load information of all computing instances.

[0065] For each local scheduler, the local scheduler can internally include two components, namely a load management component responsible for calculating the GPU load information of all computing instances deployed on the computing node where it is located, and a monitor component responsible for monitoring the running-related information of all computing instances deployed on the computing node where it is located (such as: GPU memory utilization rate, GPU memory bandwidth utilization rate, the number of inference requests waiting in the sequential queue, and the number of running inference requests, etc.).

[0066] For each computing instance, the computing instance can internally include two components, namely a cache engine component for caching the results of inference to avoid repeated calculations, and an executor component responsible for calling the large model to execute the inference task corresponding to the inference request.

[0067] Please, based on Figure 1 and Figure 2 , refer to Figure 3 . Figure 3 is a flowchart of a load-aware scheduling method for an inference system shown in an exemplary embodiment of the present application.

[0068] In this embodiment, the above-mentioned load-aware scheduling method for the inference system can be applied to a global scheduler as shown in Figure 1 .

[0069] As shown in Figure 3 , the above-mentioned load-aware scheduling method for the inference system may include the following steps:

[0070] Step 302: Obtain a target inference request to be executed.

[0071] In this embodiment, the above-mentioned global scheduler may first obtain an inference request to be executed (which can be called a target inference request). For example, the global scheduler may obtain an inference request in sequence from a request queue for temporarily storing all inference requests to be executed as the target inference request.

[0072] Step 304: Based on the GPU load information of each maintained computing instance, determine a target computing instance whose GPU load meets a preset condition.

[0073] In this embodiment, as described above, the above-mentioned global scheduler may maintain the GPU load information of each computing instance, and the GPU load information may indicate the load condition of the GPU. In this case, the global scheduler may determine a computing instance whose GPU load meets a preset condition (which can be called a target computing instance) based on the maintained GPU load information of each computing instance.

[0074] In some embodiments, the computing instance whose GPU load meets the preset condition may be the computing instance with the minimum CPU load.

[0075] Step 306: Send the target inference request to the target computing instance so that the target computing instance executes the target inference request.

[0076] In this embodiment, when the above-mentioned global scheduler determines the above-mentioned target computing instance, it can send the above-mentioned target inference request to the target computing instance through the communication connection with the above-mentioned inference engine, so that the target computing instance executes the target inference request. That is, the global scheduler can schedule the target inference request to the target computing instance whose GPU load meets the preset condition, thereby realizing load-aware scheduling of inference requests on the inference engine at the cluster level, optimizing resource usage efficiency, reducing waiting time and processing latency, and improving the overall throughput of the system.

[0077] In practical applications, when the above-mentioned global scheduler maintains the GPU load information of each computing instance, it can associate and store the instance information of each computing instance with the GPU load information, so that it can locate the target computing instance according to the instance information of the target computing instance whose determined GPU load meets the preset conditions, and send the above-mentioned target inference request to the target computing instance.

[0078] In some embodiments, as described above, in addition to including at least one computing instance deployed on each computing node in the computing cluster, the above-mentioned inference engine may further include a local scheduler deployed on each computing node in the computing cluster. Correspondingly, when sending the above-mentioned target inference request to the above-mentioned target computing instance, the target inference request may be specifically sent to the target local scheduler, so that the target local scheduler sends the target inference request to the target computing instance; wherein, the target local scheduler is specifically the local scheduler deployed on the computing node where the target computing instance is located.

[0079] In some embodiments, in order to make the dynamic update of the GPU load information of each computing instance maintained by the above-mentioned global scheduler easy to implement, accurate and reliable, for each computing node in the computing cluster, the local scheduler deployed on the computing node may periodically count the GPU load information of each computing instance deployed on the computing node (for example: generate GPU load information according to the current GPU load situation of the computing instance) at a preset time period, and report the counted GPU load information to the global scheduler, so that the global scheduler can update the GPU load information of each computing instance deployed on the computing node maintained based on the GPU load information reported by the local scheduler.

[0080] In some embodiments, the above-mentioned GPU load information may specifically be manifested as instance resource utilization. It should be noted that for each computing instance, the instance resource utilization of the computing instance may be an indicator for indicating the GPU load amount, and specifically may be an indicator calculated by the local scheduler deployed on the computing node where the computing instance is located based on the GPU memory utilization and GPU memory bandwidth utilization of the computing instance. Among them, the GPU memory utilization can be obtained by reading the hardware information of the GPU, and the GPU memory bandwidth utilization can be obtained through the Data Center GPUManager (DCGM) service provided by NVIDIA. At this time, when determining the target computing instance whose GPU load meets the preset conditions based on the maintained GPU load information of each computing instance, the target computing instance with the smallest instance resource utilization may be specifically determined based on the maintained each computing instance.

[0081] In practical applications, the above GPU memory utilization rate can be the utilization rate of memory blocks in the GPU. Here, a block refers to a group of threads that can be organized together for easy management and communication. All threads in a block can cooperate to execute certain operations and can use shared memory to facilitate this cooperation. Each block can access a specific amount of shared memory and other resources. Therefore, the number of blocks and the number of threads in each block jointly determine the total amount of resources used.

[0082] It should be noted that when scheduling a certain inference request, if the inference request is for the Prefill (pre-filling) stage during the inference process of a large model, the above block utilization rate can specifically be the number of blocks required for the prompt of the large model; if the inference request is for the Decode (decoding) stage during the inference process of a large model, the above block utilization rate can specifically be the actual GPU memory utilization rate.

[0083] During the inference process of large models (especially generative models in natural language processing), the Prefill (pre-filling) stage and the Decode (decoding) stage refer to different steps when generating text.

[0084] In the Prefill stage, the model usually generates an initial text segment that contains the starting part of the generated sequence. The goal of the Prefill stage is to provide a good start for subsequent generation. This stage may involve using some predefined strategies or algorithms to select the most appropriate starting words or phrases. For example, when dealing with a conditional generation task, the model may generate a part of the text as a basis based on the input conditions (such as summary generation, question answering, etc.).

[0085] In the main process of text generation in the Decode stage, the model generates new content word by word or token by token based on the existing text. In each iteration, the model predicts the next most likely word and adds it to the current text sequence. This process continues until a preset end condition is reached (for example, reaching the maximum text length or generating a text end marker). The Decode stage may use various strategies to optimize the quality of generation, such as techniques like Top-k sampling, Top-p (also known as Nucleus Sampling) sampling, etc., which can help the model avoid generating overly ordinary or meaningless content.

[0086] In practical applications, the Prefill and Decode stages are closely connected. The Prefill stage provides the initial text basis, while the Decode stage is responsible for gradually expanding on this basis to generate the complete text output. These two stages jointly determine the quality and coherence of the finally generated text.

[0087] In some embodiments, the instance resource utilization can be calculated based on the GPU memory utilization and GPU memory bandwidth utilization of the computing instance through the following formula:

[0088] ;

[0089] ;

[0090] ;

[0091] where represents the instance resource utilization of the computing instance, represents the total GPU memory size of the computing instance, represents the composite load of the computing instance, represents the number of tokens generated by the computing instance in each iteration, represents the GPU memory utilization of the computing instance, represents the composite coefficient, represents the preset first control coefficient, represents the preset second control coefficient, , represents the preset second threshold, represents the GPU memory bandwidth utilization of the computing instance.

[0092] Through the above composite load, the above GPU memory utilization and the above GPU memory bandwidth utilization can be combined for load-aware scheduling. In an actual application scenario, the GPU memory bandwidth utilization is negatively correlated with the GPU memory utilization. To prevent the data access performance between the GPU memory and the main memory from dropping sharply due to excessive GPU memory bandwidth utilization, a threshold (i.e., the above second threshold) can be set for the GPU memory bandwidth utilization. When the GPU memory bandwidth utilization is lower than the threshold, it tends to increase the GPU memory bandwidth utilization, thereby obtaining a larger composite load; when the GPU memory bandwidth utilization exceeds the threshold, it tends to increase the GPU memory utilization while reducing the GPU memory bandwidth utilization, thereby obtaining a smaller composite load. In this way, the GPU memory utilization can be improved while maintaining a lower GPU memory bandwidth utilization.

[0093] The above and All are control coefficients that can be obtained through experiments and are used to control 's response sensitivity. The value of should be greater than 's value, indicating that it is desired to increase the GPU memory bandwidth utilization when the GPU memory bandwidth utilization is lower than the threshold.

[0094] The above can represent the number of new tokens generated by all inference requests in each iteration (TokensGenerated Per Iteration). Therefore, can represent how many iterations are required to consume the remaining GPU memory of the computing instance at the current token generation rate. It can be seen that the smaller the instance resource utilization rate of the computing instance, the more iterations the remaining GPU memory of the computing instance can support, and the smaller the load of the computing instance; the larger the instance resource utilization rate of the computing instance, the fewer iterations the remaining GPU memory of the computing instance can support, and the larger the load of the computing instance. That is to say, the load size of the computing instance is positively correlated with the instance resource utilization rate.

[0095] It should be noted that the inference process of the large model refers to the process of using a trained large machine learning model to generate corresponding output results based on the given input data (i.e., the data included in the inference request). In the field of natural language processing, especially in large language models, the inference process usually involves generating a continuous text sequence according to the given input prompt (Prompt), and generating one Token in each iteration. Here, the "token" usually refers to a series of small units into which the text is segmented in natural language processing, such as words or characters, etc., depending on the tokenization method of the model. Specifically, a piece of text can be provided to the model as input, and this piece of text is called a Prompt; the model receives this Prompt and converts it into an internally processable form, usually converting each word into a corresponding numerical representation (e.g., WordEmbedding); based on the current Prompt, the model predicts the next most likely token to appear through complex calculations (e.g., multi-layer neural network operations); the model adds the predicted token to the existing Prompt to form a new Prompt; the above steps can be repeated multiple times, and each repetition is called an iteration. In each iteration, the model will predict the next token based on the latest Prompt until the set stop condition is reached, such as generating a certain number of Tokens or encountering a specific end flag.

[0096] In some embodiments, the above GPU load information may further include the above GPU memory utilization rate. When determining the target computing instance with the minimum instance resource utilization rate based on the GPU load information of each computing instance maintained, the global scheduler may specifically first determine, based on the GPU load information of each computing instance maintained, the computing instances (referred to as candidate computing instances) with a GPU memory utilization rate less than a preset threshold (which may be referred to as the first threshold), and then may sort the candidate computing instances according to the instance resource utilization rate to determine the target computing instance with the minimum instance resource utilization rate from the candidate computing instances. In this way, when the global scheduler determines the target computing instance whose GPU load meets the preset conditions based on the GPU load information of each computing instance maintained, the number of size comparisons of the GPU loads of different computing instances can be reduced, the selection efficiency of the target computing instance can be improved, and thus the scheduling efficiency can be improved.

[0097] In some embodiments, in the case where a computing instance is in the termination process, the instance resource utilization rate of the computing instance may be set to infinity, so as to avoid scheduling an inference request to a computing instance that is in the termination process.

[0098] In the above technical solution, the inference engine in the inference system may include at least one computing instance deployed on each computing node in the computing cluster. The computing resources of each computing instance may include the GPUs installed on the corresponding nodes. The global scheduler in the inference system may maintain the dynamically updated GPU load information of each computing instance, and in the case of obtaining a target inference request to be executed, may first determine the target computing instance whose GPU load meets the preset conditions based on the GPU load information of each computing instance maintained, and then send the target inference request to the target computing instance for the target computing instance to execute the target inference request, so as to implement the scheduling for the target inference request.

[0099] By adopting the above method, by adding a global scheduler in the inference system, the global scheduler can schedule the inference request to a computing instance with an appropriate GPU load in the inference engine for execution, realizing load-aware scheduling of the inference request on the inference engine at the cluster level. Thus, the usage efficiency of GPU resources can be optimized, the waiting time and processing delay for the execution of the inference request can be reduced, and the overall throughput of the inference system can be improved. In addition, by scheduling the inference request to the inference engine through the global scheduler instead of directly scheduling the inference request within the inference engine, the scheduling method can be decoupled from the specific inference engine, and has a certain degree of generality and scalability.

[0100] Corresponding to the embodiment of the foregoing method, the present application also provides an embodiment of a device.

[0101] Please refer to Figure 4, Figure 4 It is a schematic structural diagram of a device shown in an exemplary embodiment of the present application. At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, a memory 408, and a non-volatile memory 410. Of course, other required hardware may also be included. One or more embodiments of the present application may be implemented in a software manner. For example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into the memory 408 and then runs it. Of course, in addition to the software implementation manner, one or more embodiments of the present application do not exclude other implementation manners, such as a logic device or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic module, and may also be hardware or a logic device.

[0102] Please refer to Figure 5 , Figure 5 It is a block diagram of an inference system shown in an exemplary embodiment of the present application.

[0103] The above-mentioned inference system can be applied to Figure 4 the device shown in the figure to implement the technical solution of the present application. Among them, the inference system includes a global scheduler 502 and an inference engine 504; the inference engine includes at least one computing instance deployed on each computing node in the computing cluster; the computing resources of the computing instance include the GPUs installed on the computing node where it is located; the global scheduler maintains the GPU load information of each computing instance that is dynamically updated.

[0104] The global scheduler is used for:

[0105] Obtain a target inference request to be executed;

[0106] Based on the GPU load information of each computing instance maintained, determine a target computing instance whose GPU load meets a preset condition;

[0107] Send the target inference request to the target computing instance so that the target computing instance executes the target inference request.

[0108] In some embodiments, the preset condition is that the GPU load in the computing instance is the smallest.

[0109] In some embodiments, the inference engine further includes a local scheduler deployed on each computing node in the computing cluster;

[0110] The sending the target inference request to the target computing instance includes:

[0111] Send the target inference request to the target local scheduler, so that the target local scheduler sends the target inference request to the target computing instance; wherein, the target local scheduler is a local scheduler deployed on the computing node where the target computing instance is located.

[0112] In some embodiments, the global scheduler is further configured to:

[0113] Obtain the GPU load information of each computing instance on each computing node sent periodically by the local scheduler on each computing node according to a preset time period;

[0114] Update the GPU load information of each computing instance maintained based on the obtained GPU load information.

[0115] In some embodiments, the GPU load information includes instance resource utilization rate; wherein, the instance resource utilization rate is an index for indicating GPU load calculated by the local scheduler based on the GPU memory utilization rate and GPU memory bandwidth utilization rate of the computing instance;

[0116] Determining a target computing instance whose GPU load meets a preset condition based on the GPU load information of each computing instance maintained includes:

[0117] Based on the GPU load information of each computing instance maintained, determine the target computing instance with the smallest instance resource utilization rate.

[0118] In some embodiments, the GPU load information further includes GPU memory utilization rate;

[0119] Determining the target computing instance with the smallest instance resource utilization rate based on the GPU load information of each computing instance maintained includes:

[0120] Based on the GPU load information of each computing instance maintained, determine candidate computing instances whose GPU memory utilization rate is less than a preset first threshold;

[0121] Sort the candidate computing instances according to the instance resource utilization rate to determine the target computing instance with the smallest instance resource utilization rate from the candidate computing instances.

[0122] In some embodiments, the instance resource utilization rate is calculated based on the GPU memory utilization rate and GPU memory bandwidth utilization rate of the computing instance through the following formula:

[0123] ;

[0124] ;

[0125] ;

[0126] Among them, represents the instance resource utilization rate of the computing instance, represents the total GPU memory size of the computing instance, represents the composite load of the computing instance, represents the number of tokens generated by the computing instance in each iteration, represents the GPU memory utilization rate of the computing instance, represents the composite coefficient, represents a preset first control coefficient, represents a preset second control coefficient, , represents a preset second threshold, represents the GPU memory bandwidth utilization rate of the computing instance.

[0127] In some embodiments, the instance resource utilization rate of the computing instance is set to infinity when the computing instance is in the termination process.

[0128] For the apparatus embodiments, they basically correspond to the method embodiments, so for the relevant parts, please refer to the partial descriptions of the method embodiments. The apparatus embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the technical solution of this application.

[0129] The systems, apparatuses, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0130] In a typical configuration, a computer includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0131] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0132] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transitory medium that can store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0133] It should be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0134] The above describes specific embodiments of the present application. Other embodiments are within the scope of the present application. In some cases, the acts or steps recited in the present application may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0135] The terms used in one or more embodiments of the present application are for the purpose of describing particular embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the", and "said" are also intended to include the plural forms unless the context clearly dictates otherwise. The term "and / or" means and includes any and all possible combinations of one or more of the associated listed items.

[0136] The description of terms such as "one embodiment", "some embodiments", "example", "specific example", or "a kind of implementation manner" used in one or more embodiments of the present application means that the specific features or characteristics described in connection with the embodiment are included in at least one embodiment of the present application. The schematic description of these terms does not necessarily refer to the same embodiment. Moreover, the specific features or characteristics described can be combined in a suitable manner in one or more embodiments of the present application. In addition, different embodiments and the specific features or characteristics in different embodiments can be combined without contradiction.

[0137] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of the present application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0138] The above description is only the preferred embodiment of one or more embodiments of the present application, and is not intended to limit one or more embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of the present application shall be included in the protection scope of one or more embodiments of the present application.

[0139] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

Claims

1. A load-aware scheduling method for an inference system, which is applied to a global scheduler in the inference system; the inference system further includes an inference engine; the inference engine includes at least one computing instance deployed on each computing node in a computing cluster, and a local scheduler deployed on each computing node in the computing cluster; the computing resources of the computing instance include the GPUs installed on the computing nodes where they are located; the global scheduler maintains dynamically updated GPU load information of each computing instance; the GPU load information includes instance resource utilization rate; The instance resource utilization rate is an index for indicating the GPU load calculated by the local scheduler based on the GPU memory utilization rate and the GPU memory bandwidth utilization rate of the computing instance; The method includes: Obtaining a target inference request to be executed; Based on the GPU load information of each computing instance maintained, determining a target computing instance with the minimum instance resource utilization rate; Sending the target inference request to a target local scheduler, so that the target local scheduler sends the target inference request to the target computing instance, and the target computing instance executes the target inference request; wherein, the target local scheduler is a local scheduler deployed on the computing node where the target computing instance is located.

2. The method according to claim 1, the method further includes: Obtaining the GPU load information of each computing instance on each computing node periodically sent by the local scheduler on each computing node according to a preset time period; Updating the GPU load information of each computing instance maintained based on the obtained GPU load information.

3. The method according to claim 1, the GPU load information further includes the GPU memory utilization rate; The determining, based on the GPU load information of each computing instance maintained, a target computing instance with the minimum instance resource utilization rate includes: Based on the GPU load information of each computing instance maintained, determining candidate computing instances with the GPU memory utilization rate less than a preset first threshold; Sorting the candidate computing instances according to the instance resource utilization rate to determine a target computing instance with the minimum instance resource utilization rate from the candidate computing instances.

4. The method according to claim 1, calculating the instance resource utilization rate based on the GPU memory utilization rate and the GPU memory bandwidth utilization rate of the computing instance through the following formula: ; ; ; Among them, Indicates the instance resource utilization rate of the computing instance, Indicates the total GPU memory size of the computing instance, Indicates the composite load of the computing instance, Indicates the number of tokens generated by the computing instance in each iteration, Indicates the GPU memory utilization rate of the computing instance, Indicates the composite coefficient, Indicates the preset first control coefficient, Indicates the preset second control coefficient, , Indicates the preset second threshold, Indicates the GPU memory bandwidth utilization rate of the computing instance.

5. The method according to claim 4, the instance resource utilization rate of the computing instance is set to infinity when the computing instance is in the termination process.

6. An inference system, the inference system includes a global scheduler and an inference engine; the inference engine includes at least one computing instance deployed on each computing node in a computing cluster, and a local scheduler deployed on each computing node in the computing cluster; the computing resources of the computing instance include the GPU carried on the computing node where it is located; the global scheduler maintains the GPU load information of each computing instance updated dynamically; the GPU load information includes instance resource utilization rate; The instance resource utilization rate is an index for indicating the GPU load calculated by the local scheduler based on the GPU memory utilization rate and the GPU memory bandwidth utilization rate of the computing instance; The global scheduler is used for: Obtaining a target inference request to be executed; Based on the GPU load information of each computing instance maintained, determining a target computing instance with the minimum instance resource utilization rate; Sending the target inference request to a target local scheduler, so that the target local scheduler sends the target inference request to the target computing instance, and the target computing instance executes the target inference request; wherein, the target local scheduler is a local scheduler deployed on the computing node where the target computing instance is located.

7. An electronic device, including: A processor; A memory for storing processor-executable instructions; Wherein, the processor realizes the method according to any one of claims 1 to 5 by running the executable instructions.

8. A computer-readable storage medium having stored thereon computer instructions which, when executed by a processor, implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Distributed training task scheduling method, system and device for intelligent computing

    CN115248728A

  • Heterogeneous GPU (Graphics Processing Unit) energy consumption perception scheduling method and device oriented to AI reasoning

    CN117742946A

  • Load balancing method and device and electronic equipment

    CN118819841A