Data scheduling method, apparatus, and computer device
Patent Information
- Application Number
- PCT/CN2025/141668
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-12-11
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025141668_01102026_PF_FP_ABST
Abstract
Description
Data scheduling methods, apparatus and computer equipment
[0001] This application claims priority to Chinese Patent Application No. 202510381264.4, filed on March 27, 2025, entitled “Data Scheduling Method, Apparatus and Computer Equipment”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of model technology, and in particular to a data scheduling method, apparatus, and computer equipment. Background Technology
[0003] Large language model (LLM) inference typically involves two phases: prefill and decoding, used for understanding input and generating responses, respectively. The prefill phase computes key-value data based on inference requests and stores it in a key-value cache (KV cache). The decoding phase loads this key-value data to perform decoding tasks, iteratively generating output tokens and continuously updating the key-value data until completion. Prefilling is computationally intensive and can be processed in parallel, while decoding is memory-intensive, requiring each token to be processed individually. Due to the different computational requirements of the prefill and decoding phases, an architecture called the prefill-decode (PD) architecture has been proposed. This architecture deploys the prefill and decoding phases on separate servers. Specifically, the prefill server performs the prefill task to obtain key-value data and then transmits the prefilled key-value data to the decoding server for decoding. However, there are currently no clear guidelines on the principles and timing of transmitting key-value data from the prefill server to the decoding server. Improper execution of transmission rules will lead to low efficiency in the inference process. Therefore, there is an urgent need for a data scheduling method to improve inference efficiency. Summary of the Invention
[0004] This application provides a data scheduling method, apparatus, and computer device that can improve inference efficiency.
[0005] Firstly, this application provides a data scheduling method that, in addition to a pre-filling server and a decoding server, adds a scheduling controller. This controller can select appropriate inference requests based on the actual operating status of the pre-filling and decoding servers, and schedule the key-value data of the inference request from the pre-filling server to a specific decoding server at the appropriate time to execute the decoding task. The scheduling controller can obtain the status of the pre-filling server, such as the status of multiple inference requests that have completed pre-filling, and the load status of multiple decoding servers. This allows it to act as an intermediate monitoring point to monitor the pre-filling and decoding servers separately, and to perform scheduling based on the monitoring results, determining which inference request should be scheduled to which decoding server for the decoding task. Because the scheduling considers not only the status of inference requests on the pre-filling server but also the load status of the decoding server, it can reduce the probability of insufficient memory on the decoding server under high concurrency conditions and improve inference efficiency.
[0006] In the above process, the key-value data of which inference request is sent can be determined by the status of the already filled inference request. Through status monitoring, the service level, remaining processing time and size of the key-value data of the inference request can be obtained, so as to further determine the target inference request. For example, the target inference request has the highest service level or the shortest remaining processing time.
[0007] In the process of determining the target inference request described above, different scheduling strategies can be used to determine the target inference request in different ways. For example, if the scheduling strategy is service level, i.e., user priority, then the target inference request can be determined based on its service level or its remaining processing time. The higher the service level of the inference request, the greater its probability of being prioritized for scheduling. Conversely, the shorter the remaining processing time, the greater its probability of being prioritized for scheduling. Furthermore, under the condition that the pre-filling processing time is equal, the specified processing time for inference requests with high service levels is shorter than that for inference requests with low service levels. Therefore, after pre-filling processing, the remaining processing time for inference requests with high service levels is still relatively short. Thus, the probability of inference requests with high service levels being prioritized for scheduling remains high. Moreover, the remaining processing time can also indicate the urgency of the inference request itself. Therefore, this approach considers not only service level but also inference requests with urgent processing needs, preventing inference request timeouts and ensuring the normal operation of the inference service.
[0008] In some embodiments, when determining the target inference request, the bandwidth between servers is further considered, that is, the transmission capacity between the pre-filling server and multiple decoding servers. Accordingly, the process of determining the target inference request includes: determining the target inference request from multiple inference requests based on the state of the pre-filling server (that is, the size of the key-value data of the inference request and the remaining processing time of each inference request) and the bandwidth between the pre-filling server and multiple decoding servers. Data transmission between servers is also time-consuming. Considering the size of the key-value data being transmitted and the bandwidth used for transmission, timeouts of inference requests can be further avoided, thereby ensuring the normal operation of the inference service. By considering the transmission time consumption, inference requests with larger key-value data can be prioritized to avoid timeouts.
[0009] Once the target inference request is determined, the decoding server to which the key-value data of the target inference request should be sent can be determined by the load status of the decoding server. By monitoring the status, the number of decoding tasks running on the decoding server and the remaining memory can be obtained, thereby further determining the decoding server. The determined target decoding server can meet the decoding task requirements of the target inference request.
[0010] In some embodiments, when determining the target decoding server, it is selected from multiple decoding servers based on the remaining memory of the decoding server and the size of the key-value data in the target inference request. Whether the remaining memory size can accommodate the key-value data of the target inference request is an important condition for ensuring the normal operation of the service. Therefore, considering the remaining memory and key-value data when considering the processing capacity of the decoding server can ensure effective scheduling and inference efficiency.
[0011] In a second aspect, a data scheduling apparatus is provided, comprising at least one functional module for implementing the method provided in the first aspect or any alternative method thereof. In some embodiments, the module in the data scheduling apparatus is implemented in hardware or firmware.
[0012] Thirdly, a computer device is provided, comprising a processor and a memory connected to the memory, the computer device being used to implement the method provided in the first aspect or any alternative method of the first aspect.
[0013] Fourthly, a computer-readable storage medium is provided for storing at least one piece of program code for implementing the method provided in the first aspect or any alternative method of the first aspect.
[0014] Fifthly, a computer program product is provided, which is used to implement the method provided in the first aspect or any alternative method of the first aspect.
[0015] In a sixth aspect, a computing system is provided, the computing system comprising a plurality of pre-filled servers, a plurality of decoding servers, and a scheduling controller, wherein the plurality of pre-filled servers are used to pre-fill inference requests, the plurality of decoding servers are used to perform decoding tasks, and the scheduling controller is used to execute a data scheduling method as provided in the first aspect or any alternative method of the first aspect. Attached Figure Description
[0016] Figure 1 is a schematic diagram of the structure of a large language model provided in an embodiment of this application;
[0017] Figure 2 is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0018] Figure 3 is a schematic diagram of the structure of a computer device provided in an embodiment of this application;
[0019] Figure 4 is a schematic diagram of the system configuration provided in an embodiment of this application;
[0020] Figure 5 is a flowchart of a data scheduling process provided in an embodiment of this application;
[0021] Figure 6 is a schematic diagram of the structure of a data scheduling device provided in an embodiment of this application. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. It should be noted that all information (including but not limited to interfaces), data (including but not limited to files, directories, etc.) and signals involved in this application are authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the data involved in this application were obtained under fully authorized conditions.
[0023] To facilitate understanding, the key terms and concepts involved in this application will be explained below.
[0024] Large language model (LLM): refers to a computer model that can process and generate natural language. LLM is an artificial intelligence technology based on deep learning that can predict the next word or sentence by learning the statistical patterns and semantic information of language data.
[0025] Figure 1 shows a schematic diagram of the LLM structure. The LLM consists of multiple network layers: a word embedding layer, N transformer layers, and an output layer. Here, N is a positive integer. The word embedding layer and the N transformer layers are connected sequentially. The outputs of the word embedding layer and the N transformer layers are called hidden states, which represent the intermediate representation of the model's input after transformation by the network layers.
[0026] The word embedding layer is used to embed the input of LLM (referred to as model input) to obtain the hidden state of the model input in the word embedding layer. The model input is the input sequence corresponding to the text. The input sequence includes the numerical representation of each word in the text. The input sequence is obtained based on the word segmentation of the text.
[0027] Each transformer layer includes an attention layer and a feed-forward network (FFN) layer, with the attention layer and FFN layer adjacent to each other. The functions of the attention layer, FFN layer, and output layer will be described below in (1), (2), and (3), respectively.
[0028] (1) Attention layer
[0029] The attention layer can be a multi-head attention layer. For example, the attention layer includes H attention heads and an integration unit connected to the H attention heads. H is an integer greater than 1. Each attention head includes a key-value data calculation unit and an attention calculation unit.
[0030] The key-value data calculation unit is used to normalize and linearly transform the hidden states output by the previous network layer to obtain the processing result; multiply the processing result by the query (Q) weight matrix to obtain the query vector corresponding to the model input; multiply the processing result by the key (K) weight matrix to obtain the key (K) data; multiply the processing result by the value (V) weight matrix to obtain the value cache (V) data corresponding to the model input; and output the query vector, key cache, and value data corresponding to the model input. The query vector, key data, and value data also correspond to their respective attention heads, and the key data and value data together form key-value data. In some embodiments, the key data is also called key cache data, the value data is also called value cache data, and the key-value data is also called key-value cache data.
[0031] The attention calculation unit is used to perform attention operations on the query vector, key data, and value data output by the key-value data calculation unit, obtain the attention operation result corresponding to the attention head, and output the attention operation result.
[0032] For example, the attention calculation unit performs attention calculation according to the following formula (1) to obtain the attention calculation result.
[0033] In formula (1), Q, K, and V represent the query vector, key data, and value data, respectively. Attention(Q, K, V) represents the attention operation result, d k Indicates the feature dimension.
[0034] The integration unit is used to combine the attention operation results output by H attention heads into a long vector. Then, a linear transformation is performed on the concatenated vector to integrate the information from different attention heads, so as to obtain the final attention operation result of the attention layer. The final attention operation result is output to the FFN of the transformer for subsequent operations.
[0035] In other embodiments, the attention layer is a single-head attention layer, which includes an attention head. The integration unit performs a linear transformation on the attention operation result output by the attention head to obtain the final attention operation result of the attention layer.
[0036] (2) FFN layer
[0037] The FFN layer processes the attention operation results output by the attention layer to obtain the hidden state of the model input in the corresponding transformer layer, and outputs that hidden state. The next transformer layer processes this hidden state to obtain a new hidden state, until the last transformer layer outputs the hidden state.
[0038] (3) Output layer
[0039] The output layer performs a lexical probability mapping on the hidden states output from the last transformer layer, outputting a dictionary probability matrix. This dictionary probability matrix includes the probability of the numerical representation of each word in the dictionary. Subsequent sampling operations are performed on the dictionary probability matrix, and the sampling results are segmented to obtain a single word. This word is generated by performing one forward computation on the LLM.
[0040] A token, also known as a label or mark, corresponds to a character or word in a text during natural language processing. The numerical representation of a token is its identification (ID). For example, in English, "apple" can be a token; for some morphologically rich languages, words may be broken down into smaller sub-word units as tokens.
[0041] The inference process is the process of performing inference tasks using an LLM (Limited Language Model). The inference process using KV (Key-Value) caching technology consists of two phases: a prefill phase and a decoding phase. During inference, K rounds of inference are performed based on the LLM for the inference task. The first round of inference occurs in the prefill phase, and the subsequent K-1 rounds occur in the decoding phase. In other words, for the inference task, one round of inference is performed based on the LLM in the prefill phase, and K-1 rounds of inference are performed based on the LLM in the decoding phase. Each round of inference performs a forward computation on the LLM, generating a token. The numerical representations of the K tokens generated from the K rounds of inference are combined to form an output sequence. These K tokens are used as the inference result corresponding to the inference task. The numerical representation of the i-th token in the output sequence is the numerical representation of the token generated in the i-th round of inference. K is an integer greater than 1, and i is an integer greater than 0 and less than or equal to K.
[0042] Pre-filling stage: In the pre-filling stage, the input sequence corresponding to the inference request is input into the LLM model for one forward computation, generating a token and completing one round of inference. During each round of inference, each attention layer of the LLM generates key-value data, which is stored in the KV cache for use in the decoding stage. The key-value data generated by an attention layer includes the key-value data generated by each attention head of that attention layer.
[0043] Decoding Phase: The decoding phase is where the model generates the output sequence. During each round of inference in the decoding phase, the input sequence and the numerical representations of the generated tokens are combined to form a context sequence. The most recently generated token (i.e., the token generated in the previous round of inference) is used as the current token. The numerical representation of the current token in the context sequence is input into the LLM for a forward computation to generate a new token. For example, the attention layer in the LLM generates the query vector and key-value data corresponding to the current token. The query vector corresponding to the current token includes the query vectors generated by each attention head of the attention layer, and the key-value data corresponding to the current token includes the key-value data generated by each attention head of the attention layer. During the text generation inference process, the model typically generates tokens one by one. Except for the most recently generated token, the key-value data corresponding to previously processed tokens remains unchanged in subsequent inference steps. Therefore, the key-value data corresponding to the current token generated by the attention layer is added to the KV cache. This allows the key-value data of the current token in the KV cache, along with the key-value data generated by the attention layer in previous inference rounds, to form the context sequence's key-value data for that attention layer, thus updating the KV cache. Next, the attention layer performs attention operations based on the query vector corresponding to the current token and the context sequence's key-value data for that attention layer, obtaining the attention calculation result for the context sequence and outputting the attention calculation result. During this process, it is not necessary to recalculate the key-value data corresponding to previous tokens; only the key-value data of the latest token needs to be calculated, thereby reducing a significant amount of redundant computation. Then, the connected FFN layer further processes this attention calculation result until the LLM outputs a new token.
[0044] The implementation environment of this application will be described below with reference to Figure 2. Figure 2 is a schematic diagram of a computing system provided in an embodiment of this application. This computing system can be a data center, such as an artificial intelligence (AI) inference cluster. Referring to Figure 2, taking an inference cluster as an example, the computing system includes multiple pre-fill servers 201, multiple decoding servers 202, and a scheduling controller 203. The pre-fill servers 201 are used to perform pre-filling tasks to pre-fill the input inference request, obtain key-value data, and transmit the key-value data to the decoding servers under the scheduling of the scheduling controller. The decoding servers 202 are used to load the key-value data, execute the decoding task corresponding to the inference request, output a token iteratively, and continuously transmit key-value data until the end. The scheduling controller 203 is used to perform task scheduling based on the status of the inference requests on the multiple pre-fill servers 201 and the load status of the multiple decoding servers 202, that is, to determine which inference request will be scheduled to which decoding server for decoding. It should be noted that the scheduling controller 203 can be implemented in hardware or software. It can be implemented as a standalone server or as part of a server in the system. This application embodiment does not limit this.
[0045] In some embodiments, the pre-fill server 201 includes a pre-fill service controller for maintaining the state of pre-filled inference requests within the server. This state includes the size of the key-value data for each inference request and may also include the service level objective (SLO) of the inference request. In some embodiments, the decoding server 202 includes a decoding service controller for maintaining the load state of the server, such as the memory occupied by each inference request being decoded, the number of decoding tasks running simultaneously on the server, and the remaining memory of the server. It should be noted that the pre-fill service controller and the decoding service controller can be implemented in hardware or software. They can be implemented as independent devices or as part of a server in a system; this application embodiment does not limit this.
[0046] In some embodiments, the scheduling controller 203 includes a pre-fill server status monitor and a decoding server status monitor, etc., for periodically monitoring the status of inference requests on the pre-fill servers and the load status of the decoding servers. The scheduling controller may include an inference request selection algorithm, i.e., selecting which pre-fill server and which inference request's key-value data to send next. The scheduling controller 203 may include a decoding server selection algorithm, i.e., selecting which decoding server to receive the key-value data corresponding to the inference request selected by the selection algorithm. The scheduling controller can communicate with each pre-fill server and decoding server to obtain the status of each server, and issue commands to each server, such as requesting the pre-fill servers and decoding servers to update their status, or specifying a pre-fill server to send a block of key-value data to a certain decoding server, etc.
[0047] The aforementioned computing system can be used in conversational AI systems such as chatbots, where the model needs to generate a corresponding response for each user input. Using key-value caching technology allows the model to generate responses faster, making the dialogue smoother and more natural, and improving the user experience. This computing system can also be used for text generation tasks, such as article writing and story creation, where the model needs to generate text word by word. Leveraging inference key-value caching can accelerate the text generation process and improve the efficiency of content production.
[0048] This application does not limit the type of server. For example, the server can be a general-purpose server, an AI server, a cloud server, a baseboard management controller (BMC) board, a smart network card, a V2X device, a B2C device, a B2B device, etc.
[0049] The servers described above are interconnected. In some embodiments, the servers, etc., support standard communication technologies and / or protocols. The network includes, but is not limited to, Transmission Control Protocol / Internet Protocol (TCP / IP) networks and RDMA networks such as RoCE networks, InfiniBand (IB) networks, Storage Area Networks (SANs), Local Area Networks (LANs), Metropolitan Area Networks (MANs), Wide Area Networks (WANs), mobile, wired or wireless networks, private networks, or any combination of virtual private networks. In some implementations, technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or part of the link. In other embodiments, custom and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0050] The aforementioned scheduling controller 203 can be implemented as a computer device. Figure 3 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Referring to Figure 3, the computer device 300 may include a processor 301 and a memory 302. The computer device 300 also includes an external interface 303 and a universal serial bus (USB), etc. It is understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the computer device 300. In other embodiments of this application, the computer device 300 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0051] Processor 301 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. The different processing units may be independent devices or integrated into one or more processors.
[0052] The processor 301 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 301 is a cache memory. This memory can store instructions or data that the processor 301 has just used or that are used repeatedly. If the processor 301 needs to use the instruction or data again, it can retrieve it directly from this memory. This avoids repeated accesses, reduces the waiting time of the processor 301, and thus improves the efficiency of the system.
[0053] Memory 302 can be used to store computer executable program code, which includes instructions. Memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, and universal flash storage (UFS). Processor 301 executes various functional applications and data processing of computer device 300 by running instructions stored in memory 302 and / or instructions stored in memory disposed in the processor.
[0054] External interface 303 can be used to connect to external devices, thereby enabling communication with external devices.
[0055] It is understood that the above structure is only one hardware possibility for implementing the scheduling device provided in the embodiments of this application, and the embodiments of this application do not specifically limit it.
[0056] The above embodiments described the implementation environment and possible hardware structures. The principles of the embodiments of this application are described below. The above model inference process can include two stages: a prefill stage and a decoding stage. In the prefill stage, based on a given user-requested Prompt token list: X = [1:s], with s tokens as input, a series of KV (key-value) pairs are calculated and stored as key-value data. In the decoding stage, the key-value data is loaded, and output tokens are generated iteratively, continuously updating the key-value data until the end.
[0057] Analysis reveals that the two stages exhibit some heterogeneity. The pre-filling stage is computationally intensive, requiring the model to process the entire input sequence and compute the key-value data for each token, a process involving numerous matrix-vector multiplications. Since the computation of key-value data for each token is independent (due to the parallelism of self-attention), it does not need to be executed sequentially, fully utilizing parallel computing capabilities. The decoding stage, however, is memory-intensive. In the token generation stage, the model generates tokens sequentially, computed only one token at a time. Therefore, each time a token is generated, the model accesses and updates the key-value data, resulting in numerous memory read / write operations. Consequently, the performance of this stage is typically limited by memory bandwidth. Furthermore, since each generated token requires storing its corresponding key-value data, the required GPU memory increases rapidly with the number of generated tokens, making the memory resource demands far higher than in the pre-filling stage.
[0058] Based on the characteristics of different stages, the pre-filling and decoding stages of model inference are separated into different servers, a process known as PD separation. PD separation minimizes system idle time and increases throughput. However, in the current PD separation architecture, there is no direct communication mechanism between the pre-filling server and the decoding server, meaning they cannot know each other's status. While this design reduces mutual interference, it also prevents them from coordinating scheduling: since the two servers operate independently, they cannot perform collaborative dynamic scheduling based on optimization goals. Furthermore, the decoding server may run out of memory. Different users request different numbers of tokens for inference, resulting in significant differences in the size of the key-value data after pre-filling. If the decoding server is unaware of the size of the next key-value data to be transmitted, it may request the pre-filling server to send the key-value data when memory is insufficient, leading to memory shortage and inability to fully receive the key-value data, thus impacting inference efficiency. In this embodiment, a scheduling controller is added to the pre-filling server and decoding server. This scheduling controller can select appropriate inference requests based on the actual operating status of the pre-filling and decoding servers, and schedule the key-value data of the inference request from the pre-filling server to a specific decoding server at the appropriate time to execute the decoding task. The scheduling controller can obtain the status of the pre-filling server, such as the status of multiple inference requests that have completed pre-filling, and the load status of multiple decoding servers. This allows it to act as an intermediate monitoring point to monitor the pre-filling and decoding servers separately, and to perform scheduling based on the monitoring results, determining which inference request should be scheduled to which decoding server for the decoding task. Because the scheduling considers not only the status of inference requests on the pre-filling server but also the load status of the decoding server, the probability of insufficient memory on the decoding server under high concurrency can be reduced, thus improving inference efficiency.
[0059] In the above process, the key-value data of which inference request is sent can be determined by the status of the already filled inference request. Through status monitoring, the service level, remaining processing time and size of the key-value data of the inference request can be obtained, so as to further determine the target inference request. For example, the target inference request has the highest service level or the shortest remaining processing time.
[0060] In the process of determining the target inference request described above, different scheduling strategies can be used to determine the target inference request in different ways. For example, if the scheduling strategy is service level, i.e., user priority, then the target inference request can be determined based on its service level or its remaining processing time. The higher the service level of the inference request, the greater its probability of being prioritized for scheduling. Conversely, the shorter the remaining processing time, the greater its probability of being prioritized for scheduling. Furthermore, under the condition that the pre-filling processing time is equal, the specified processing time for inference requests with high service levels is shorter than that for inference requests with low service levels. Therefore, after pre-filling processing, the remaining processing time for inference requests with high service levels is still relatively short. Thus, the probability of inference requests with high service levels being prioritized for scheduling remains high. Moreover, the remaining processing time can also indicate the urgency of the inference request itself. Therefore, this approach considers not only service level but also inference requests with urgent processing needs, preventing inference request timeouts and ensuring the normal operation of the inference service.
[0061] In some embodiments, when determining the target inference request, the bandwidth between servers is further considered, that is, the transmission capacity between the pre-filling server and multiple decoding servers. Accordingly, the process of determining the target inference request includes: determining the target inference request from multiple inference requests based on the state of the pre-filling server (that is, the size of the key-value data of the inference request and the remaining processing time of each inference request) and the bandwidth between the pre-filling server and multiple decoding servers. Data transmission between servers is also time-consuming. Considering the size of the key-value data being transmitted and the bandwidth used for transmission, timeouts of inference requests can be further avoided, thereby ensuring the normal operation of the inference service. By considering the transmission time consumption, inference requests with larger key-value data can be prioritized to avoid timeouts.
[0062] Once the target inference request is determined, the decoding server to which the key-value data of the target inference request should be sent can be determined by the load status of the decoding server. By monitoring the status, the number of decoding tasks running on the decoding server and the remaining memory can be obtained, thereby further determining the decoding server. The determined target decoding server can meet the decoding task requirements of the target inference request.
[0063] In some embodiments, when determining the target decoding server, it is selected from multiple decoding servers based on the remaining memory of the decoding server and the size of the key-value data in the target inference request. Whether the remaining memory size can accommodate the key-value data of the target inference request is an important condition for ensuring the normal operation of the service. Therefore, considering the remaining memory and key-value data when considering the processing capacity of the decoding server can ensure effective scheduling and inference efficiency.
[0064] For computing systems, different scheduling strategies can be employed during scheduling, and the method for determining the target inference request and the target decoding server is related to the scheduling strategy. That is, during scheduling, different methods can be used to determine the target inference request and the target decoding server based on different scheduling strategies to meet corresponding optimization objectives. Scheduling strategies include service level-based scheduling, load balancing, etc. The process of service level-based scheduling is described in the above embodiments, while load balancing refers to load balancing in terms of computing power, i.e., among the decoding servers. The following specific examples illustrate scheduling based on different scheduling strategies.
[0065] Refer to Figure 4, which illustrates the various servers and their maintenance status. These servers include pre-populated servers P1 to P2. M Scheduling controller and decoding servers D1 to D2 N .
[0066] In some embodiments, pre-fill servers P1 to P M This includes a pre-population service controller, which maintains the state of inference requests already populated on the pre-population server and updates it to the scheduling controller in a timely manner. This state can be in the form of metadata, denoted as PS. i Where i is an integer greater than or equal to 1. The state of the i-th pre-filled server is determined by PS. i ={P i,1 ,P i,2 ,…,P i,mi} indicates that m i This indicates that the i-th pre-filled server contains m. i One inference request. For each P i,j , including (SLO) i,j TTS i,j C i,j ), of which SLO i,j This inference requests P. i,j Service level, the higher the level, the higher the SLO i,j The smaller the value, the unit is seconds; TTS i,j This inference request P i,jThe remaining processing time, that is, how much time remains before a SLO is violated. i,j The unit of measurement is seconds. Because the pre-filled server received the inference request P... i,j Next, it is necessary to request P for this reasoning. i,j Complete the pre-filling task, therefore TTS i,j <SLO i,j C i,j This represents the size of the key-value data, measured in gigabytes, where j is less than or equal to m. i Positive integers.
[0067] In some embodiments, decoding servers D1 to D N This includes a decoding service controller, which maintains the load status of the decoding server. The load status can also be in the form of metadata, denoted as D. i This metadata may include, but is not limited to: the number of running decoding tasks (numJobs), the memory (M) occupied by each decoding task, and the remaining memory (M) of the server. res The load status of the i-th decoding server is determined by M. i,res =M i -Σ(M(D i,j )) represents. Where M i D represents the total memory of the i-th decoding server. i,j M(D) represents the j-th decoding task of the i-th server. i,j ) represents D i,j The memory occupied by the task. Therefore, M i,res This represents the remaining memory for the i-th decoding server. numJobs i This represents the number of decoding tasks running simultaneously on the i-th server.
[0068] The scheduling controller includes a pre-filled state monitor and a decoding state monitor, which are used to monitor the status of the two types of servers mentioned above, respectively. After information from multiple pre-filled servers is aggregated into the scheduling controller, the pre-filled state monitor maintains PS = {PS1, PS2, ...} = {P...}. i,j After information from multiple decoding servers is aggregated to the scheduling controller, the decoding status detector of the scheduling controller maintains D = {D1, D2, ...}. The scheduling controller also includes a bandwidth monitor to monitor the bandwidth between any pair of pre-filled servers and decoding servers, denoted as B = {B...}. i,k}, where B i,k Represents any pre-filled server P i With any decoding server D k The bandwidth between.
[0069] The following explanation uses a service level-based scheduling strategy as an example. This scheduling includes two parts: selecting inference requests and selecting decoding servers. The process of selecting inference requests is illustrated in steps 501 to 502 of Figure 5, and the process of selecting decoding servers is illustrated in step 503. Referring to Figure 5, this data scheduling process includes:
[0070] 501. The scheduling controller periodically obtains the status of multiple pre-filled servers, the load status of multiple decoding servers, and the bandwidth between multiple pairs of pre-filled servers and decoding servers.
[0071] In some embodiments, when acquiring the status of the pre-filled server and the decoding server, the scheduling controller periodically receives the status actively sent by the pre-filled server and the decoding server. Optionally, to avoid status update failures, the scheduling controller can detect status updates. If no status is received from a server within a target interval, remedial measures such as fault detection can be triggered to ensure the normal operation of the system service. In other embodiments, the scheduling controller periodically polls each server to obtain the status maintained on each server. Optionally, the scheduling controller can detect status updates. If the status of any server has not changed compared to the previously polled status, it indicates a possible server failure, and remedial measures such as fault detection are triggered to ensure the normal operation of the system service. In the above process, when the server actively sends status to the scheduling controller, the sending can be performed in parallel to ensure the timeliness of the status. When the scheduling controller polls, its polling of the pre-filled server and the decoding server can be performed in parallel to ensure the accuracy of the status. For example, the above process is executed by the pre-filled status monitor and the decoding status monitor included on the scheduling controller.
[0072] In some embodiments, when acquiring bandwidth between multiple pairs of pre-filled servers and decoding servers, the scheduling controller can perform bandwidth probing by controlling the transmission of test messages between the pre-filled servers or decoding servers. For example, a test command can be sent to pre-filled server P1, instructing P1 to send a test message to decoding server D2. Decoding server D2, upon receiving the test message, returns a test response to test the bandwidth between pre-filled server P1 and decoding server D2. After obtaining the bandwidth, pre-filled server P1 sends the obtained bandwidth to the scheduling controller. Of course, this bandwidth test can also be performed in other ways, such as sending test commands to the decoding server, etc., which is not limited in this embodiment. In addition, when determining which pair of pre-filled servers and decoding servers to test the bandwidth, pairing can be random or based on the server status to reduce the testing burden. For example, for some decoding servers, their load status is already poor, making it difficult to handle more decoding tasks. Skipping such decoding servers during pairing not only reduces the amount of computation required for selection but also simplifies the selection process. For example, the above process is executed by a bandwidth monitor included on the scheduling controller.
[0073] 502. The scheduling controller determines the target inference request from multiple inference requests on multiple pre-filled servers based on the bandwidth between multiple pairs of pre-filled servers and decoding servers, the key-value data size of the inference request, and the remaining processing time of each inference request.
[0074] When the pre-filled state monitor is not empty, that is, when there is an inference request to be scheduled, the scheduling controller can start selecting the key-value data of the next inference request to enter the decoding server.
[0075] The following explanation uses service level scheduling as an example to illustrate the process of step 502. The main idea behind this scheduling strategy in selecting inference requests is: (1) to satisfy the SLO (Service Level Requirement) of each inference request as much as possible; (2) if two inference requests have the same remaining processing time, to slightly favor the request with the larger amount of key-value data to avoid timeouts. To ensure a certain degree of randomness in the selection, this process is implemented through probabilistic random sampling. Based on point (2) of the above selection strategy, inference requests with smaller remaining processing time are given a higher probability of being selected.
[0076] Accordingly, step 502 above includes the following steps: For each pre-filled server, determine the minimum bandwidth between the pre-filled server and the decoding server from the bandwidth between multiple pairs of pre-filled servers and decoding servers, thereby finding the worst bandwidth condition, that is, the bandwidth with the worst transmission capacity. Its mathematical representation can be Worst_BW.i =min(B i,k ), for any k. Where Worst_BW i This indicates the minimum bandwidth.
[0077] For inference requests on multiple pre-filled servers, the mathematical representation of the transfer time required for each inference request to transmit key-value data under its worst-case bandwidth conditions can be expressed as Worst_transfer_time for P. i,j :Z i,j =C i,j / Worst_BW i .
[0078] Based on the remaining processing time and transmission time of multiple inference requests, the maximum waiting time for each inference request within the pre-filled server is calculated and determined, which can be mathematically represented as T. i,j ={TTS i,j -Z i,j}, where T ij Indicates a reasoning request P ij The maximum waiting time is within the pre-filled server.
[0079] The selection probability of each inference request is determined based on this waiting time. The selection probability refers to the probability that the key-value data of the inference request is selected and transmitted to the decoding server, and its mathematical representation can be Prob. i,j =Softmax(-T) i,j ).
[0080] After obtaining the probabilities, a random number between 0 and 1 is generated. Then, using the cumulative distribution function, the probability of each inference request (Prob) is calculated based on this random number. i,j Select one inference request from the inference requests as the target inference request.
[0081] The above process ensures that inference requests with higher service levels have a greater probability of being selected. Since inference requests with higher service levels typically have a shorter SLO, if we assume that two inference requests with different service levels have gone through the same pre-fill time, the TTS of the inference request with higher service levels will also be correspondingly smaller. Therefore, when undergoing softmax transformation, a shorter waiting time will generate a higher probability, which ensures that inference requests with higher service levels have a greater probability of being selected.
[0082] The process of selecting target inference requests described above is guided by the service level of the user's inference request and by referring to the real-time state of the pre-populated server. It ensures that requests with high service levels can be processed in a timely manner through the coordination of the task scheduling controller, which satisfies the request priority of high-value users. At the same time, it must also take into account the needs of low-value users or free users to achieve effective scheduling of inference requests.
[0083] 503. The scheduling controller determines the target decoding server from multiple decoding servers based on the load status of multiple decoding servers and the status of the target inference request.
[0084] The process of step 503 will be explained below, taking the scheduling strategy of minimizing memory waste as an example. The main idea of this scheduling strategy in selecting the decoding server is: (1) If there are multiple decoding servers with remaining memory greater than the key-value data size of the target inference request, then select the decoding server with the smallest remaining memory; (2) If there are no such servers, then select the one with the closest remaining memory to the key-value data size of the target inference request, that is, the one with the shortest average waiting time.
[0085] Accordingly, step 503 above includes the following steps:
[0086] For multiple decoding servers, determining which decoding server has more remaining memory than the key-value data size of the target inference request can be achieved using the following code: For select{D s},in which M s,res >C i,j .
[0087] If there are decoding servers with more remaining memory than the key-value data size of the target inference request, select the decoding server with the least remaining memory as the target decoding server. This process can be implemented using the following code: If Then:select D k ,in which k=argmin(M s,res ).
[0088] If no decoding server has remaining memory smaller than the key-value data size of the target inference request, then the decoding server with the largest remaining memory is selected as the target decoding server. This process can be implemented using the following code: else:select D k ,in which k=argmax(M resIn this process, the decoding server with the largest remaining memory is identified as the target decoding server. Alternatively, the process can involve selecting a decoding server whose remaining memory is closest to the size of the key-value data in the target inference request to avoid wasting excessive resources. In other embodiments, if no decoding server has remaining memory smaller than the size of the key-value data in the target inference request, it means that all decoding servers are operating at full capacity. In this case, the scheduling controller monitors each decoding server until the remaining memory of one of the decoding servers exceeds the size of the key-value data in the target inference request, and then identifies that server as the target decoding server, ensuring timely and normal operation of the decoding task. For example, the scheduling controller monitors the decoding server whose remaining memory is closest to the size of the key-value data. Once it detects that the remaining memory of that decoding server reaches the size of the key-value data in the target inference request, it triggers the instruction sending process in step 504, thereby transmitting the key-value data of the target inference request to that decoding server.
[0089] The following explanation uses load balancing as an example to illustrate the process of step 503. This scheduling strategy aims to schedule the target inference request to the decoding server with the lowest task load when selecting the decoding server. The main idea is: (1) Select the group of decoding servers running the fewest decoding tasks simultaneously. (2) If there is only one decoding server in the group, select that decoding server as the target decoding server; if there are multiple decoding servers in the group, select the decoding server with the largest remaining memory as the target decoding server.
[0090] Accordingly, step 503 above includes the following steps:
[0091] For multiple decoding servers, the decoding server group Ds with the fewest decoding tasks running simultaneously is determined based on the number of decoding tasks running on each decoding server. This process can be implemented by the following code: select{Ds}, in which Ds = argmin(numJobsi);
[0092] If there is only one decoding server in the decoding server group Ds, then that decoding server is selected as the target decoding server. If there are multiple decoding servers in the decoding server group Ds, then the decoding server with the largest remaining memory among all decoding servers is selected as the target decoding server based on the remaining memory of each decoding server.
[0093] The process can be implemented using the following code:
[0094] If|{D s}=1,Then:select the only element in D s
[0095] else:select D k ,in which k=argmax(M s,res )
[0096] By implementing a load-balancing-based scheduling strategy, the probability of insufficient memory on the decoding server in high-concurrency scenarios can be effectively reduced. Furthermore, this scheduling strategy prevents the decoding server from being overloaded, resulting in high overall operating efficiency, stable operation, and ensuring the stability and security of the entire decoding process.
[0097] It should be noted that the strategy for selecting inference requests and the strategy for selecting decoding servers can be decoupled. That is, based on different optimization goals of the system, different scheduling strategies can be set for different stages to achieve the corresponding optimization goals.
[0098] 504. The scheduling controller sends a transmission instruction to the pre-fill server, which instructs the pre-fill server to transmit the key-value data of the target inference request to the target decoding server.
[0099] Through the above process, the decoding server that can perform the decoding task in the current state can be identified. After the target decoding server is determined, the scheduling controller can quickly send a transmission instruction to the pre-filled server, requesting it to send the key-value data of the target inference request to the target decoding server. After receiving the key-value data of the target inference request, the target decoding server can perform the decoding task based on the key-value data to output the inference result of the inference request.
[0100] In some embodiments, after sending the transmission command, the scheduling controller updates the status of each detector again; if necessary, it executes the inference request selection algorithm again. For example, if the pre-filled status monitors are not empty, the scheduling controller can select the inference request and the above process again.
[0101] Steps 501 to 504 above illustrate the process of first selecting an inference request and then selecting a target decoding server. However, under other scheduling strategies, these two processes can be integrated into a single process. For example, a scheduling strategy could optimize for the consistency of user waiting time, ensuring that all inference requests have similar or identical user waiting times, resulting in a consistent user experience. User waiting time (time to first token, TTFT) refers to the time from when the pre-filled server receives the user's inference request to when the first token is output after decoding begins. This is a crucial metric during inference, indicating the time the user waits for the LLM to provide an answer. "Time consistency" means ensuring that all inference requests have similar or identical TTFTs, allowing users to perceive service consistency. Understandably, in a PD-separated architecture, TTFT includes the pre-filling time, the time spent waiting in the pre-filled server for transmission to the decoding server, the actual transmission time of the key-value data, and the time to generate the first token in the decoding server. The decoding time is negligible.
[0102] The main idea of this algorithm is to simultaneously select between the inference request and the target decoding server, aiming to make the TTFT of the inference request as close as possible to the target TTFT. Assuming the target TTFT is F (in seconds), and the inference request P... i,j The time elapsed since the inference request was received is S. i,j Its key-value data size is C i,j The bandwidth between the servers is B. i,k Steps 502 and 503 may include: for each decoding server, based on the key-value data size of each inference request and the bandwidth between server pairs, obtaining the predicted time required for the key-value data transfer of each inference request to the decoding server; based on the predicted values of each inference request relative to multiple decoding servers, determining the corresponding average predicted value for each inference request; then, based on the average predicted value corresponding to the inference request and the time elapsed since the inference request was received, calculating the predicted TTFT of the inference request; and finally, based on the time difference between the predicted TTFT and the target TTFT of each inference request, selecting the inference request whose predicted TTFT is closest to the target TTFT, and determining the decoding server with the smallest predicted value of the inference request as the target decoding server. The code for this algorithm is as follows:
[0103] Z i,j,k =C i,j / B i,k / / Z i,j,k Indicates transmission P i,j The predicted time required to transmit the key-value data to the k-th decoding server;
[0104] Z i,j =mean(Z) i,j,k ) / / Z i,j Indicates transmission P i,j The average time required for the KV cache to predict values.
[0105] T i,j =abs(F-(S) i,j +Z i,j )) / / T i,j P represents i,j The time difference between the predicted TTFT and the target TTFT.
[0106] Choose i and j such that [i,j] = argmin(T) i,j / / Select the inference request that is closest to the target TTFT.
[0107] The scheduling strategy described above can ensure that inference requests have similar or identical user waiting times. In other words, the time difference between each user sending an inference request and seeing the corresponding inference result (i.e., the response) is not large, resulting in a consistent user experience.
[0108] This application also provides a data scheduling device. As shown in FIG6, FIG6 is a schematic diagram of the structure of a data scheduling device provided in an embodiment of this application. As shown in FIG6, the device includes a pre-filling state acquisition module 601, a decoding state acquisition module 602, and a sending module 603.
[0109] The pre-fill status acquisition module 601 is used to acquire the status of the pre-fill server, the status of the pre-fill server including the status of multiple inference requests that have been pre-filled;
[0110] The decoding status acquisition module 602 is used to acquire the load status of multiple decoding servers;
[0111] The sending module 603 is used to send a transmission instruction to the pre-filled server based on the status of the pre-filled server and the load status of the plurality of decoding servers;
[0112] The transmission instruction instructs the pre-filled server to transmit the key-value data of the target inference request to the target decoding server among the plurality of decoding servers.
[0113] In some embodiments, the sending module includes:
[0114] The first determining unit is configured to determine the target inference request from the plurality of inference requests based on the state of the pre-filled server;
[0115] The second determining unit is used to determine the target decoding server from the plurality of decoding servers based on the load status of the plurality of decoding servers and the status of the target inference request;
[0116] A sending unit is used to send the transmission instruction to the pre-filled server.
[0117] In some embodiments, the status of the inference request includes the remaining processing time for the inference request.
[0118] The first determining unit is used to determine the target inference request from the plurality of inference requests based on the remaining processing time of the plurality of inference requests.
[0119] In some embodiments, the state of the inference request includes the service level of the inference request.
[0120] The first determining unit is configured to determine a target inference request from the plurality of inference requests based on the service level of the plurality of inference requests.
[0121] In some embodiments, the first determining unit is configured to determine a target inference request from the plurality of inference requests based on the state of the pre-filled server and the bandwidth between the pre-filled server and the plurality of decoding servers.
[0122] In some embodiments, the load status includes the remaining memory of the decoding server, and the second determining unit is configured to select the target decoding server from the plurality of decoding servers according to the remaining memory and the key-value data size of the target inference request.
[0123] In some embodiments, the load status includes at least one of the number of decoding tasks of the decoding server and the remaining memory. The second determining unit is configured to select a group of decoding servers from the plurality of decoding servers according to the number of decoding tasks; and select the target decoding server from the group of decoding servers according to the remaining memory and the key-value data size of the target inference request.
[0124] In some embodiments, the determination of the target inference request and the target decoding server is related to a scheduling strategy.
[0125] In some embodiments, the data scheduling device described above is also used to coordinate the implementation of other steps executed by the scheduling controller in the foregoing embodiments. It should be understood that the device provided in the above embodiments is only illustrated by the division of the above functional units when performing data scheduling. In practical applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the data scheduling device and the data scheduling method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0126] This application provides a computer-readable storage medium for storing at least one piece of program code for implementing the above-described data scheduling method.
[0127] This application provides a computer program product for implementing the above-described data scheduling method.
[0128] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items that have substantially the same function and purpose. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are merely used to distinguish one element from another. For example, without departing from the various examples described, a first key can be referred to as a second key, and similarly, a second key can be referred to as a first key. Both the first key and the second key can be keys, and in some cases, they can be separate and different keys.
[0129] In this application, the term "at least one" means one or more, and the term "multiple" means two or more. For example, multiple keys means two or more keys.
[0130] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0131] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of program structure information. This program structure information includes one or more program instructions. When these program instructions are loaded and executed on a computing device, the processes or functions according to the embodiments of this application are generated, in whole or in part.
[0132] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0133] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A data scheduling method, characterized by, Applied to a data scheduling device, the method includes: Obtain the status of the pre-populated server, which includes the status of multiple inference requests that have been pre-populated; Get the load status of multiple decoding servers; Based on the status of the pre-filled server and the load status of the plurality of decoding servers, a transmission instruction is sent to the pre-filled server; The transmission instruction instructs the pre-filled server to transmit the key-value data of the target inference request to the target decoding server among the plurality of decoding servers.
2. The method of claim 1, wherein, Sending a transmission instruction to the pre-fill server based on the status of the pre-fill server and the load status of the plurality of decoding servers includes: Based on the state of the pre-filled server, the target inference request is determined from the plurality of inference requests; Based on the load status of the plurality of decoding servers and the status of the target inference request, the target decoding server is determined from the plurality of decoding servers; Send the transmission instruction to the pre-filled server.
3. The method of claim 2, wherein, The status of the inference request includes the remaining processing time for the inference request. The step of determining the target inference request from the plurality of inference requests based on the state of the pre-filled server includes: determining the target inference request from the plurality of inference requests based on the remaining processing time of the plurality of inference requests; or, The state of the inference request includes the service level of the inference request, and determining the target inference request from the plurality of inference requests based on the state of the pre-filled server includes: determining the target inference request from the plurality of inference requests based on the service level of the plurality of inference requests; or, The target inference request is determined from the plurality of inference requests based on the state of the pre-filled server and the bandwidth between the pre-filled server and the plurality of decoding servers.
4. The method of claim 1, wherein, The load status includes the remaining memory of the decoding server. Determining the target decoding server from the plurality of decoding servers based on the load status of the plurality of decoding servers and the status of the target inference request includes: selecting the target decoding server from the plurality of decoding servers according to the remaining memory and the key-value data size of the target inference request. or, The load status includes at least one of the number of decoding tasks of the decoding server and the remaining memory. Determining the target decoding server from the plurality of decoding servers based on the load status of the plurality of decoding servers and the status of the target inference request includes: selecting a group of decoding servers from the plurality of decoding servers according to the number of decoding tasks; and selecting the target decoding server from the group of decoding servers according to the remaining memory and the key-value data size of the target inference request.
5. The method according to any one of claims 1 to 4, characterized in that, The method for determining the target inference request and the target decoding server is related to the scheduling strategy.
6. A data scheduling apparatus, characterized by comprising: The device includes: The pre-fill status acquisition module is used to acquire the status of the pre-fill server, the status of the pre-fill server including the status of multiple inference requests that have been pre-filled; The decoding status acquisition module is used to acquire the load status of multiple decoding servers; The sending module is used to send a transmission instruction to the pre-filled server based on the status of the pre-filled server and the load status of the plurality of decoding servers; The transmission instruction instructs the pre-filled server to transmit the key-value data of the target inference request to the target decoding server among the plurality of decoding servers.
7. A computer device, comprising: The computer device includes a processor and a memory, the processor being connected to the memory, and the computer device is used to implement the data scheduling method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one piece of program code, which is used to implement the data scheduling method as described in any one of claims 1 to 5.
9. A computer program product, characterised in that, The computer program product is used to implement the data scheduling method as described in any one of claims 1 to 5.
10. A computing system, comprising: The computing system includes multiple pre-fill servers, multiple decoding servers, and a scheduling controller. The multiple pre-fill servers are used to pre-fill inference requests, the multiple decoding servers are used to perform decoding tasks, and the scheduling controller is used to execute the data scheduling method as described in any one of claims 1 to 5.