Inference task scheduling method and computing device
By caching and parallel issuing inference tasks, and optimizing task scheduling combined with task waiting time and complexity information, the problem of underutilizing computing power resources in the existing technology is solved, and the inference efficiency and system throughput are improved.
Patent Information
- Application Number
- CN202510369821.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
When existing inference frameworks such as Transformers deal with low complexity inference tasks, the computing power resources of computing devices are not fully utilized, resulting in low inference efficiency and tasks that need to be queued up.
By caching multiple inference tasks, designing aggregation tasks are issued in parallel, and task waiting time and complexity information are combined to optimize task scheduling, ensuring that tasks with longer waiting time are aggregated first, and the computing power resources of computing devices are reasonably allocated.
It improves the computing power resource utilization rate of computing devices, shortens the completion time of inference tasks, improves the throughput and response speed of the inference system, and improves the user experience.
Smart Images

Figure CN120276822A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method for scheduling inference tasks and a computing device. Background Art
[0002] A large model refers to a neural network model with a very large number of parameters, usually including billions or even hundreds of billions of parameters, having powerful expression and learning capabilities, and being able to efficiently execute inference tasks. Specifically, large models usually rely on inference frameworks such as Transformers to execute inference tasks. An inference framework is a tool for efficiently executing model inferences, which can analyze input data and generate corresponding outputs or decisions.
[0003] In current inference frameworks, such as Transformers, when dealing with inference tasks with relatively low complexity, due to the limitations of their scheduling algorithms, the computing power resources of computing devices are often not fully utilized; in addition, inference tasks usually need to queue up and wait, so the inference efficiency is relatively low. Summary of the Invention
[0004] The embodiments of the present application provide a method for scheduling inference tasks and a computing device, which can issue inference tasks to an inference model in parallel to fully utilize the computing power resources of the computing device and improve the execution efficiency of inference tasks.
[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0006] In a first aspect, the embodiments of the present application provide a method for scheduling inference tasks, the method includes: caching a plurality of inference tasks; determining a first inference task from the plurality of inference tasks; wherein, the first inference task includes an inference task whose waiting duration is greater than or equal to a first time threshold; generating an aggregated task corresponding to the first inference task from the plurality of inference tasks based on the task complexity information of the inference tasks; wherein, the task complexity information is information related to the processing duration of the inference tasks; the aggregated task includes one or more inference tasks; issuing the inference tasks in the aggregated task in parallel.
[0007] Based on this solution, by designing a cache queue to cache multiple inference tasks, multiple inference tasks can be aggregated in the cache queue, so as to realize the parallel distribution of inference tasks in the aggregated task. In this way, the inference model can process multiple inference tasks simultaneously. Therefore, compared with the serial execution method that can only process a single inference task each time, this solution can improve the inference efficiency and complete multiple inference tasks in a shorter time. Moreover, this parallel execution method can consume more computing resources, so the utilization rate of the computing resources of the computing device can be improved, and the idle of computing resources caused by processing a single inference task each time can be avoided. On this basis, the first inference task is determined according to the waiting duration, so that the timeout tasks with longer waiting durations can be screened out, and then the timeout tasks with longer waiting durations are preferentially aggregated. This can prevent the inference tasks from being in the waiting aggregation state for a long time and unable to be executed, ensure that all inference tasks can obtain the processing opportunity, and reduce the response delay of the inference requests corresponding to the timeout tasks. In addition, during the task aggregation process, comprehensively considering the waiting duration and task processing duration of each inference task is beneficial to the reasonable allocation of the computing resources of the computing device and ensures efficient operation.
[0008] In a possible implementation, generating an aggregated task corresponding to the first inference task from multiple inference tasks based on the task complexity information of the inference tasks includes: determining a second inference task from multiple inference tasks based on the first task complexity information of the first inference task; wherein, the complexity distance between the second task complexity information of the second inference task and the first task complexity information of the first inference task is less than or equal to a distance threshold; aggregating the first inference task and the second inference task to form an aggregated task.
[0009] Based on this solution, determining other second inference tasks to be aggregated with the first inference task according to the task complexity information and then forming an aggregated task can ensure that the task complexities of the second inference tasks and the first inference task are similar, and their required task execution durations are also similar. Therefore, the completion times of these second inference tasks and the first inference task when executed in parallel are also similar. In this way, it can be avoided that after the second inference task with a shorter execution time is completed, the computing resources corresponding to the second inference task are idle for a long time. Therefore, the parallelism of each inference task in the aggregated task can be improved, which is beneficial to the reasonable allocation of the computing resources of the computing device, shortens the total duration from the generation to the execution completion of the inference task, and improves the throughput of the inference system.
[0010] In yet another possible implementation, the task complexity information of the inference task includes at least one of the token length and the task type.
[0011] Based on this solution, through the token length, task type, etc., the task complexity of the inference task can be simply estimated, providing a basis for the aggregation of the inference tasks.
[0012] In yet another possible implementation, the method further includes: when there is no inference task among the multiple inference tasks whose waiting duration is greater than or equal to the first time threshold, taking the inference task with the maximum waiting duration among the multiple inference tasks as the first inference task.
[0013] Based on this solution, if there is no inference task whose waiting duration reaches the first time threshold, that is, no timeout task, then determine the inference task with the longest waiting duration as the first task, so as to preferentially aggregate the inference task with the longest waiting duration, which can reduce the possibility of timeout of this inference task.
[0014] In yet another possible implementation, after generating the aggregation task corresponding to the first inference task from the multiple inference tasks based on the inference task complexity information, the method further includes: adding the aggregation task corresponding to the first inference task to the task scheduling queue; wherein, the task scheduling queue is used to store aggregation tasks.
[0015] Based on this solution, adding the aggregation task to the task scheduling queue is to schedule multiple aggregation tasks in the task scheduling queue, thereby further realizing the optimization of the computing power resources of the computing device and improving the overall inference efficiency.
[0016] In yet another possible implementation, after adding the aggregation task corresponding to the first inference task to the task scheduling queue, the method further includes: adjusting the order of each aggregation task in the task scheduling queue according to the waiting duration corresponding to each aggregation task.
[0017] Based on this solution, the arrangement order of each aggregation task in the task scheduling queue can be adjusted, so as to realize the dynamic adjustment of the execution order of the aggregation tasks, which is beneficial to dealing with the dynamically changing task situation. On this basis, adjusting the arrangement order based on the waiting duration can preferentially execute the inference tasks with longer waiting durations and improve the inference ability of the computing device under high load conditions.
[0018] In yet another possible implementation, the method further includes: adjusting the aggregation task corresponding to the first inference task to the head of the task scheduling queue.
[0019] Based on this solution, by arranging the first inference task in a relatively front position, it can ensure that the first inference task with the longest timeout or waiting duration is preferentially executed, thus preventing the first inference task from waiting too long. In this way, the response speed of the inference request corresponding to the timeout task can be improved, and the user experience can be improved.
[0020] In yet another possible implementation, after determining the first inference task from multiple inference tasks, the method further includes: determining whether the waiting duration of the first inference task is less than a second time threshold; in the case where the waiting duration of the first inference task is less than the second time threshold, generating an aggregated task corresponding to the first inference task from the multiple inference tasks based on the complexity information of the inference tasks; wherein the second time threshold is greater than the first time threshold; in the case where the waiting duration of the first inference task is greater than or equal to the second time threshold, issuing the first inference for execution.
[0021] Based on this solution, when the waiting duration of the first inference task does not reach the second duration threshold, a task aggregation operation is performed on the first inference task. If the waiting duration of the first inference task reaches the second duration threshold, no matter whether the aggregation is successful or not, the task aggregation operation is no longer performed. In this way, it is possible to avoid the first inference task waiting for a long time due to the task aggregation consuming too much time.
[0022] In yet another possible implementation, the method further includes: in response to an inference request, generating an inference task corresponding to the inference request; and putting the inference task into a cache queue.
[0023] Based on this solution, storing the inference task in the cache queue facilitates performing a multi-task aggregation operation in the cache queue.
[0024] In yet another possible implementation, there is no dependency relationship between the inference tasks in the aggregated task.
[0025] Based on this solution, it can be ensured that there is no dependency relationship between the second inference task and the first inference task, thereby avoiding the inference tasks in the aggregated task being unable to execute in parallel, resulting in blocking.
[0026] In yet another possible implementation, the method further includes: based on the dependency relationship between the aggregated tasks, adjusting the order of the aggregated tasks in the task scheduling queue so that each inference task in the aggregated task sorted later does not depend on each inference task in the aggregated task sorted earlier.
[0027] Based on this solution, it can be ensured that the aggregated tasks are executed in order based on the dependency relationship, avoiding logical errors or task execution blocking caused by the inference tasks sorted earlier requiring the inference results of the inference tasks sorted later when executing.
[0028] In a second aspect, an embodiment of the present application further provides an inference task scheduling apparatus, which includes: a caching module for caching multiple inference tasks; an aggregation module for determining a first inference task from the multiple inference tasks; where the first inference task includes inference tasks with a waiting duration greater than or equal to a first time threshold; and generating an aggregation task corresponding to the first inference task from the multiple inference tasks based on the task complexity information of the inference tasks; where the task complexity information is information related to the processing duration of the inference tasks; the aggregation task includes one or more inference tasks; a distribution module for distributing the inference tasks in the aggregation task in parallel.
[0029] In a possible implementation manner, the aggregation module includes a second inference task determination unit for determining a second inference task from the multiple inference tasks based on the first task complexity information of the first inference task; where the complexity distance between the second task complexity information of the second inference task and the first task complexity information of the first inference task is less than or equal to a distance threshold; an aggregation unit for aggregating the first inference task and the second inference task to form an aggregation task.
[0030] In a possible implementation manner, the task complexity information of the inference task includes at least one of the token length and the task type.
[0031] In a possible implementation manner, the aggregation module includes a first inference task determination unit for, when there is no inference task with a waiting duration greater than or equal to the first time threshold among the multiple inference tasks, using the inference task with the maximum waiting duration among the multiple inference tasks as the first inference task.
[0032] In a possible implementation manner, the inference task scheduling apparatus includes a scheduling module for: adding the aggregation task corresponding to the first inference task to a task scheduling queue; where the task scheduling queue is used to store aggregation tasks.
[0033] In a possible implementation manner, the scheduling module is used for: adjusting the order of each aggregation task in the task scheduling queue based on the waiting duration corresponding to each aggregation task.
[0034] In a possible implementation manner, the scheduling module is used for: adjusting the aggregation task corresponding to the first inference task to the head of the task scheduling queue.
[0035] In a possible implementation manner, the aggregation module is used for: determining whether the waiting duration of the first inference task is less than a second time threshold;
[0036] When the waiting duration of the first inference task is less than the second time threshold, an aggregation task corresponding to the first inference task is generated from multiple inference tasks based on the complexity information of the inference tasks; wherein, the second time threshold is greater than the first time threshold; the sending module is configured to, when the waiting duration of the first inference task is greater than or equal to the second time threshold, send the first inference task for execution.
[0037] In a possible implementation manner, the caching module is configured to, in response to an inference request, generate an inference task corresponding to the inference request; and put the inference task into a caching queue.
[0038] In a possible implementation manner, there is no dependency relationship between the inference tasks in the aggregation task.
[0039] In a possible implementation manner, the scheduling module is configured to: based on the dependency relationship between the aggregation tasks, adjust the order of the aggregation tasks in the task scheduling queue, so that each inference task in the aggregation task sorted later does not depend on each inference task in the aggregation task sorted earlier.
[0040] In a third aspect, an embodiment of the present application further provides a computing device, including: a processor and a memory; the processor and the memory are coupled; the memory is used to store program instructions; the processor is used to execute the program instructions to execute the method according to any one of the above first aspects.
[0041] In a fourth aspect, an embodiment of the present application provides a chip, and the chip is used to execute the method according to any one of the above first aspects.
[0042] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a computer, the method according to any one of the first aspects is implemented.
[0043] In a sixth aspect, an embodiment of the present application provides a program product, including a computer program, and when the computer program is executed by a processor, the method according to any one of the first aspects is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a system architecture diagram of an inference system provided by an embodiment of the present application;
[0045] Figure 2 is a flowchart of an inference task scheduling method provided by an embodiment of the present application;
[0046] Figure 3 is a schematic diagram of determining a first inference task provided by an embodiment of the present application;
[0047] Figure 4It is a schematic diagram for determining a first inference task provided by an embodiment of the present application;
[0048] Figure 5 It is a Gantt chart for parallel processing of a second inference task provided by an embodiment of the present application;
[0049] Figure 6 It is another Gantt chart for parallel processing of a second inference task provided by an embodiment of the present application;
[0050] Figure 7 It is a schematic diagram for determining a second inference task according to a second duration threshold provided by an embodiment of the present application;
[0051] Figure 8 It is another flowchart of the inference task scheduling method provided by an embodiment of the present application;
[0052] Figure 9 It is another flowchart of the inference task scheduling method provided by an embodiment of the present application;
[0053] Figure 10 It is a schematic diagram for adjusting an aggregation task to the head of the queue proposed by an embodiment of the present application;
[0054] Figure 11 It is a schematic diagram for sorting multiple aggregation tasks at the head of the queue proposed by an embodiment of the present application;
[0055] Figure 12 It is another schematic diagram for sorting multiple aggregation tasks at the head of the queue proposed by an embodiment of the present application;
[0056] Figure 13 It is a schematic diagram of an inference task scheduling device provided by an embodiment of the present application;
[0057] Figure 14 It is a schematic diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0058] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. For the convenience of clearly describing the technical solutions of the embodiments of the present application, the first, second, etc. descriptions that appear in the embodiments of the present application are only for the purpose of indicating and distinguishing the described objects, without an order, nor do they represent a special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.
[0059] Hereinafter, the professional terms mentioned in the embodiments of the present application will be explained to facilitate the understanding of those skilled in the art.
[0060] Inference framework: A system component or structure used to perform inference tasks, commonly used in the fields of machine learning and artificial intelligence. The main function of an inference framework is to analyze input data and generate corresponding outputs or decisions.
[0061] Token: A basic concept in natural language processing (NLP) and computer language processing, referring to a text element that is regarded as a single unit. In different contexts, a token can have different specific definitions.
[0062] Cache queue: A queue used to temporarily store and manage data or requests. The main function of a cache queue is to reduce latency during data processing and improve data access efficiency, thereby enhancing the performance and response speed of the system.
[0063] Computing power resources: The ability of a computing device to execute computing tasks, usually measured in floating-point operations per second (FLOPS), which can be used to process tasks such as machine learning, data analysis, and scientific computing, including central processing unit (CPU) resources, graphics processing unit (GPU) resources, tensor processing unit (TPU) resources, input / output (I / O) resources, etc.
[0064] The following combines Figures 1 to 13 , and exemplarily illustrates the process of the inference task scheduling method.
[0065] Figure 1 is the system architecture diagram of an inference system provided by an embodiment of the present application.
[0066] As Figure 1 shown, the inference system is a software system installed in a computing device. Among them, the computing device serves as a server and can be connected to a client through wired or wireless means, thereby realizing data interaction between the server and the client.
[0067] The inference system includes a task scheduler and an inference computing unit.
[0068] Among them, the task scheduler is provided with a task cache queue and a task scheduling queue. After receiving an inference request sent by the client, the task scheduler can generate a corresponding inference task and store it in the task cache queue. The task scheduler aggregates the inference tasks in the task cache queue to obtain an aggregated task, and stores the aggregated task in the task scheduling queue, waiting to be executed in the task scheduling queue.
[0069] Among them, the inference computing unit integrates an inference model and an inference framework. The inference model can be a large model, and the inference framework can be a Transformer inference framework. After the task scheduler obtains the aggregated tasks in the task scheduling queue and distributes them to the inference model, the inference computing unit can then call the computing power resources of the computing device, and based on the inference framework, use the inference model to parallelly process each inference task in the aggregated tasks, and finally return the corresponding inference results to the client. In this way, the inference system can utilize the task scheduler and the inference computing unit to achieve parallel processing of inference tasks.
[0070] Figure 2 It is a flowchart of the inference task scheduling method provided by an embodiment of this application.
[0071] As Figure 2 shown, the method includes the following steps S11 - S14:
[0072] S11: Cache multiple inference tasks.
[0073] In step S11, the inference system is provided with a cache queue. After receiving an inference request, the corresponding inference task can be stored in the cache queue so that the inference system can obtain the inference task from the cache queue and perform aggregation. Among them, the multiple inference tasks are tasks obtained based on multiple inference requests. The multiple inference requests can be sent by the same client or by different clients. The task type of the inference task can be determined based on the inference request, and the task types of the multiple inference tasks can be the same or different.
[0074] Optionally, the caching process of the inference task in the cache queue can be implemented through the following steps:
[0075] In response to the inference request, generate an inference task corresponding to the inference request; put the inference task into the cache queue.
[0076] In this step, the inference system generates a corresponding inference task in response to an inference request sent by the client. Exemplarily, when the inference system receives an inference request sent by the client and confirms that the request is to identify the object category in an image, it can generate a corresponding image classification task. After the inference task is generated, the inference system places it in the cache queue. Among them, the inference system can place the inference task in the cache queue according to the time when the inference request is received. The earlier the time, the more forward the inference task is sorted in the cache queue. It can also place the inference task in the cache queue according to the length of the input token of the inference task. The larger the input token length, the more forward the inference task is sorted in the cache queue. The task cache queue stores inference tasks as a preliminary task set. In this way, the inference system can monitor the inference tasks in the task cache queue, specifically analyze the waiting duration and task complexity information (such as the input token length, etc.) of each inference task, and prepare for subsequent task aggregation. It can be understood that in this embodiment of the present application, only caching inference tasks according to the request time and input token length is taken as an example. In other implementation manners, the inference system can also cache inference tasks in other orders, such as in the order of decreasing inference task priority, and the present application does not limit this.
[0077] S12: Determine a first inference task from multiple inference tasks.
[0078] In step S12, the inference system determines a first inference task from multiple inference tasks, so as to use the first inference task as an aggregation center and then perform task aggregation on it.
[0079] Among them, to avoid some inference requests from being delayed for a long time, the inference system can preferentially aggregate inference tasks with longer waiting durations. Specifically, the inference system can screen out inference tasks with longer waiting durations from multiple inference tasks as the first inference task, and then aggregate the first inference task and several other inference tasks to obtain an aggregated task.
[0080] Based on this, the first inference task can include inference tasks with a waiting duration greater than or equal to a first time threshold (hereinafter referred to as timeout tasks).
[0081] Among them, the waiting duration of the inference task can be determined based on the time difference between the generation time of the inference task and the current time, or based on the time difference between the reception time of the inference request and the current time.
[0082] Among them, the first time threshold is an aggregation parameter pre-loaded by the inference system. It can be understood that the value of the first time threshold can affect the inference performance of the inference system to a certain extent. Specifically, the larger the first time threshold is, the stricter the judgment on whether it times out will be, which is beneficial to timely processing of backlogged timeout tasks and avoiding further delays. However, if the first time threshold is set too large, and the inference system overly focuses on timeout tasks, non-timeout tasks will be ignored. The smaller the first time threshold is, the looser the judgment on whether it times out will be. Therefore, more timeout tasks can be obtained, which is beneficial to comprehensively considering the waiting duration and other factors, and determining the most prioritized timeout task to be aggregated among multiple timeout tasks from multiple dimensions. However, if the first time threshold is set too small, the inference system has a lower attention to the processing time limit of inference tasks, and there is a possibility of further delay. Therefore, the first time threshold can be set according to historical experience and actual application scenarios. For example, for application scenarios where the response time of inference requests is sensitive, a smaller first time threshold can be set. Exemplarily, in some implementation manners, the first time threshold can be set to 5 seconds.
[0083] The first inference task in the embodiment of the present application includes inference tasks with a waiting duration greater than or equal to the first time threshold, which is beneficial to screening out timeout tasks with a longer waiting duration and preferentially aggregating them, thereby reducing the waiting duration of timeout tasks, reducing the response delay of inference requests corresponding to timeout tasks, improving the overall performance of the inference system, and also improving the user experience.
[0084] Furthermore, if there is at least one inference task whose waiting duration reaches the first time threshold, that is, there is at least one timeout task, the inference system can select the first inference task to be preferentially aggregated among the timeout tasks. After the aggregation of the first inference task is completed, the first inference task can be re-selected and aggregated. For example, the timeout task with the longest waiting duration can be preferentially selected as the first inference task, or the timeout task with the lowest task complexity can be preferentially selected as the first inference task, or the waiting duration and task complexity information can be comprehensively analyzed to obtain the first inference task.
[0085] Figure 3 It is a schematic diagram for determining the first inference task provided by the embodiment of the present application.
[0086] As Figure 3 shown, in the left histogram, multiple inference tasks A to D are arranged horizontally, the vertical axis represents the waiting duration of each inference task, and T1 is the first time threshold. According to Figure 3 it can be known that the waiting durations of inference task A, inference task B, and inference task C are greater than T1. Therefore, the first inference task is determined among inference task A, inference task B, and inference task C.
[0087] For another example, the timeout task with the lowest task complexity can be preferentially selected as the first inference task. Herein, the task complexity is information describing the solving difficulty of the inference task, which can be used to measure the computing power resources and time required to complete the inference task. It can be understood that the higher the task complexity of the inference task, the greater the solving difficulty, the more computing power resources required, and the longer the required execution time.
[0088] As Figure 3 shown, in the right histogram, the inference tasks A, B, and C are arranged horizontally, and the input token lengths of the respective inference tasks are marked vertically. According to Figure 3 it can be known that the input token length of the inference task A is the largest. Therefore, it can be determined that the inference task A is the first inference task.
[0089] It can be understood that the embodiments of the present application only take the above method of determining the first inference task preferentially aggregated in the timeout task as an example. In other implementation manners, the inference system can also adopt other methods to determine the first inference task preferentially aggregated in the timeout task. For example, the waiting duration and task complexity information can be comprehensively analyzed to obtain the first inference task, or the waiting duration and task priority can be comprehensively analyzed to obtain the first inference task. The present application does not make any limitation thereto.
[0090] In one implementation manner, step S12 includes the following steps S121:
[0091] S121: In the case where there is no inference task with a waiting duration greater than or equal to the first time threshold among multiple inference tasks, the inference task with the largest waiting duration among the multiple inference tasks is used as the first inference task.
[0092] In step S121, if the waiting durations of all inference tasks are less than the first time threshold, that is, all inference tasks have not timed out, the inference system can preferentially aggregate the inference task with the longest waiting time. At this time, the inference system can use the inference task with the largest waiting duration as the first inference task.
[0093] Figure 4 is another schematic diagram for determining the first inference task provided by the embodiments of the present application.
[0094] As Figure 4 shown, multiple inference tasks A to D are arranged horizontally, the waiting durations of the respective inference tasks are marked vertically, and T1 is the first time threshold. According to Figure 4 it can be known that the waiting durations of the inference tasks A to D are all less than T1, and the waiting duration of the inference task A is the longest. Therefore, it can be determined that the inference task A is the first task.
[0095] In the embodiment of the present application, through step S121, when all inference tasks do not time out, the inference task with the longest waiting duration is determined as the first task, so as to preferentially aggregate the inference task with the longest waiting duration, which can reduce the possibility of timeout of this inference task.
[0096] S13: Generate an aggregated task corresponding to the first inference task from multiple inference tasks based on the task complexity information of the inference tasks.
[0097] In step S13, the inference system aggregates several inference tasks including the first inference task according to the task complexity information of each inference task to obtain an aggregated task. For example, the inference system can perform a similarity aggregation operation based on the task complexity information, and aggregate several inference tasks with similar task complexity information into one aggregated task. In this way, the execution durations of the inference tasks in the aggregated task are similar, which is beneficial to improving the parallelism during the execution of the aggregated task.
[0098] Among them, the task complexity information is information related to the processing duration of the inference task. It can be understood that the higher the task complexity, the greater the inference difficulty of the inference task and the longer its processing duration. In actual application scenarios, the processing duration of the inference task is affected by factors such as the Token length and the task type. For example, the larger the Token length, the more complex the inference task and the longer the processing duration of the inference task; for another example, the task solving difficulty of the image generation task is usually higher than that of the image classification task, so the task processing duration of the image generation task is usually greater than that of the image classification task.
[0099] Based on this, the task complexity information of the inference task may include at least one of the token length and the task type. For example, the task complexity information may include the Token length, and the larger the input Token length, the higher the complexity of the inference task. For another example, the task complexity information may include the input Token length and the task type. It can be understood that in the embodiment of the present application, only the task complexity information including the input Token length and the task type is taken as an example. In other implementation manners, the task complexity information may also include other information related to the processing duration of the inference task, and the present application does not make a limitation thereto.
[0100] In one implementation manner, step S13 includes the following steps S131 - S132:
[0101] S131: Determine a second inference task from multiple inference tasks based on the first task complexity information of the first inference task.
[0102] In step S131, the inference system determines at least one second inference task, and then can aggregate these second inference tasks with the first inference task to obtain an aggregated task. Exemplarily, after determining the first inference task A among multiple inference tasks, the second inference task B and the second inference task C can be determined among multiple inference tasks, and then the first inference task A, the second inference task B, and the second inference task C are aggregated into an aggregated task.
[0103] Considering that selecting different inference tasks as second inference tasks for aggregation and parallel processing will affect the overall inference efficiency, the difference in the total completion time when different inference tasks are selected as second inference tasks for parallel processing will be introduced below in conjunction with Figure 5 and Figure 6 .
[0104] Figure 5 Fig. is a Gantt chart for parallel processing of the second inference task provided by an embodiment of the present application.
[0105] Figure 6 Fig. is another Gantt chart for parallel processing of the second inference task provided by an embodiment of the present application.
[0106] As shown in Figure 5 and Figure 6 , there are four inference tasks to be executed, namely the first inference task A and the inference tasks B, C, and D. The time required for each inference task is positively correlated with the task complexity. Among them, as shown in Figure 5 , the inference system takes the inference tasks B and C as the second inference tasks. In this way, the inference system parallel processes the first inference task A and the second inference tasks B and C, and starts to process the inference task D after the above three inference tasks are all processed. As shown in Figure 6 , the inference system takes the inference tasks B and D as the second inference tasks. In this way, the inference system parallel processes the first inference task A and the second inference tasks B and D, and starts to process the inference task C after the above three inference tasks are all processed.
[0107] Comparing Figure 5 and Figure 6 shows that the difference in the time required between the inference task C and the first inference task A is small. Therefore, after the inference task C in Figure 5 is completed, the inference system only needs to wait for a short time to start executing the inference task D; while the difference in the time required between the inference task D and the first inference task A is large. Therefore, after the inference task D in Figure 6 is completed, the inference system needs to wait for a long time to start executing the inference task C. In this way, the total completion time t6 of the four inference tasks to be executed in Figure 6 is later than that in Figure 5The total completion time t5 of the four inference tasks to be executed.
[0108] Based on this, during the determination of the second inference task, the inference system can use the task complexity information as the determination basis. For example, the inference system can take several inference tasks with a task complexity similar to that of the first inference task as the second inference task. In this way, the time required for the inference system to process the first inference task and the second inference task is similar. When processing the first inference task and the second inference task in parallel, their completion times are similar, avoiding the situation where the inference task with a shorter processing time is completed earlier, resulting in the computing power resources corresponding to this inference task being idle for a long time. Through such a design, the parallel processing ability of the inference system can be improved, and the utilization rate of computing power resources can be increased.
[0109] Specifically, the inference system can limit the complexity distance between the second task complexity information of the second inference task and the first task complexity information of the first inference task to be less than or equal to a distance threshold. In this way, it can be ensured that the task execution durations of the first inference task and the second inference task are similar.
[0110] In the specific application process, the inference system can determine the difference between the complexity information of the first inference task (i.e., the first task complexity information) and the complexity information of each second inference task (i.e., the second task complexity information) respectively as the complexity distance between the two.
[0111] Exemplarily, if the task complexity information includes the input Token length of the inference task, and the inference system determines that the input Token length of the first inference task A is 100 and the input Token length of the second inference task B is 120, then the complexity distance between the first task complexity information of the first task A and the second task complexity information of the second inference task B can be determined to be 20.
[0112] Exemplarily, if the task complexity information includes the input Token length of the inference task and the task type, and the inference system determines that the first inference task A is an image classification task with an input Token length of 100, then the first task complexity information OA of the first inference task A can be calculated based on the input Token length (100) and the task type weight corresponding to the image classification task (such as 10); the second inference task B is an image generation task with an input Token length of 120, then the second task complexity information OB of the second inference task B can be calculated based on the input Token length (120) and the task type weight corresponding to the image generation task (such as 30). In this way, the complexity distance between the first task complexity information of the first inference task A and the second task complexity information of the second inference task B can be determined to be the difference between OB and OA.
[0113] Further, if the difference between the first task complexity information and the second task complexity information is negative, the absolute value of this difference can be used as the complexity gap.
[0114] After that, the inference system obtains a distance threshold and determines the second inference task according to the size relationship between the complexity distance and the distance threshold. Specifically, if the complexity distance between a certain inference task and the first inference task is small, specifically less than a preset distance threshold, it can be considered that the task complexity of this inference task is similar to that of the first inference task. Therefore, the task execution duration required for this inference task and the first inference task is similar. At this time, if the inference system processes this inference task and the first inference task in parallel at the same time, the completion times of this inference task and the first inference task are similar. Based on this, this inference task can be used as the second inference task and then aggregated with the first inference task.
[0115] Figure 7 It is a schematic diagram for determining the second inference task according to the second duration threshold provided by an embodiment of the present application.
[0116] As Figure 7 shown, taking the first task complexity OA of the first inference task A as a reference, and then according to the first inference task complexity OA and the distance threshold O, a range (OA - O, OA + O) with a complexity distance less than the distance threshold is defined. From Figure 7 it can be seen that the task complexities of inference task B and inference task C are within this range. Therefore, inference task B and inference task C can be used as the second inference tasks; the task complexity of inference task D is outside this range. Therefore, inference task D is not used as the second inference task.
[0117] Among them, the distance threshold is an aggregation parameter pre-loaded by the inference system, and the distance threshold can be set according to historical experience and actual application scenarios. It can be understood that the larger the distance threshold, the more second inference tasks with satisfied complexity distances can be screened out by the inference system. In this way, more second inference tasks can be executed in parallel at the same time, but the difference in the execution completion times of the second inference tasks also increases accordingly, resulting in a decrease in the task parallelism when the inference system executes multiple second inference tasks. Exemplarily, in some implementation manners, the input Token length can be used as the task complexity information, and the preset distance threshold can be set to 30.
[0118] In one implementation manner, there is no dependency relationship between the inference tasks in the aggregation task.
[0119] Specifically, considering that there may be data or logical dependencies between inference tasks, some inference tasks need to utilize the execution results of other inference tasks during the execution process. For example, if inference task A senses surrounding obstacles and inference task B performs path planning based on the sensing results, then there is a dependency relationship between inference task A and inference task B, and the two cannot be executed in parallel. Therefore, when performing the task aggregation operation, inference tasks with dependency relationships should be avoided from being aggregated into the same aggregation task.
[0120] Based on this, when the inference system performs the task aggregation operation, for each aggregated inference task, it determines whether there is a dependency relationship between this inference task and the inference tasks that already exist in the aggregation task. If there is a dependency relationship, it is determined that this inference task cannot be aggregated into this aggregation task.
[0121] Exemplarily, during the aggregation process for the first inference task C, if the second inference task A (sensing surrounding obstacles) has already been aggregated, then the inference task B (performing path planning based on the sensing results) cannot be aggregated into the aggregation task corresponding to the first inference task C.
[0122] In the embodiment of this application, through step S131, the second inference task is determined according to the task complexity information, which can ensure that the task complexity of the second inference task is similar to that of the first inference task, and the required consumption time is also similar. Therefore, the completion times of the second inference task and the first inference task when executed in parallel are also similar. In this way, the task parallelism when the inference system executes the second inference task and the first inference task in parallel can be improved, which is beneficial to improving the utilization rate of computing resources and achieving an increase in inference throughput.
[0123] S132: Aggregate the first inference task and the second inference task to form an aggregation task.
[0124] In step S132, the inference system constructs an aggregation task using the first inference task and the second inference task determined in the previous steps. The aggregation task can be a set formed by the first inference task and the second inference task. In this way, the inference tasks included in the aggregation task can be processed in parallel, thus improving the inference efficiency of the inference system.
[0125] S14: Distribute the inference tasks in the aggregation task in parallel.
[0126] In step S14, the inference system distributes the inference tasks in the aggregation task in parallel to the inference model so that the inference model can process the inference tasks in the aggregation task in parallel. In this way, the computing power utilization rate of the computing device where the inference system is located and the inference efficiency of the inference system can be improved by simultaneously processing multiple inference tasks in parallel.
[0127] In an actual application scenario, the inference system can adopt a cyclic manner. After constructing each aggregation task, while parallelly distributing the inference tasks in the aggregation task, it can also return the step of determining the first task among the foregoing multiple inference tasks to perform a new round of multi-task aggregation operation, thereby obtaining a new aggregation task. In this way, while the inference system performs a new round of multi-task aggregation operation, it can also perform the distribution and execution steps for the aggregation task obtained in the previous round, thus improving the inference efficiency.
[0128] Exemplarily, there are multiple inference tasks A - E. In the first round of multi-task aggregation operation, the inference task with the longest waiting duration is determined among the inference tasks A - E, that is, the inference task A is the first task. Thereafter, according to the second task determination method introduced in the foregoing steps, first, the first task A is determined as the second task, and then the complexity distances between the inference tasks B - E corresponding to the second task complexity information and the inference task A corresponding to the first task complexity information are respectively determined. Furthermore, the inference tasks B and C with complexity distances less than the preset distance threshold are determined as the second tasks. In this way, an aggregation task Agg1 can be constructed using the second task A, the second task B, and the second task C. Thus, the first round of multi-task aggregation operation is completed.
[0129] The inference system removes the second task A, the second task B, and the second task C determined in the first round from the multiple inference tasks A - E. In the second round of multi-task aggregation operation, the inference task D with the longest waiting duration is determined among the inference tasks D and E, and an aggregation task Agg2 is generated for the inference task D. The subsequent steps are similar to those of the first round of multi-task aggregation operation and will not be elaborated here.
[0130] Figure 8 It is another flowchart of the inference task scheduling method provided by the embodiments of the present application.
[0131] As Figure 8 shown, the method includes the following steps S21 - S26:
[0132] S21: Cache multiple inference tasks.
[0133] S22: Determine a first inference task from the multiple inference tasks.
[0134] Among them, the first inference task includes inference tasks with a waiting duration greater than or equal to a first time threshold.
[0135] For the descriptions of steps S21 - S22, please refer to the foregoing steps S11 - S12 and will not be elaborated here.
[0136] S23: Determine whether the waiting duration of the first inference task is less than a second time threshold. If so, jump to step S24; otherwise, jump to step S26.
[0137] S24: Generate an aggregation task corresponding to the first inference task from multiple inference tasks based on the complexity information of the inference tasks.
[0138] Among them, the task complexity information is information related to the processing duration of the inference task; the aggregation task includes one or more inference tasks.
[0139] S25: Parallelly distribute the inference tasks in the aggregation task.
[0140] For the descriptions of steps S24 - S25, please refer to the foregoing steps S13 - S14, which will not be elaborated here.
[0141] S26: Distribute the first inference task for execution.
[0142] In steps S23 - S26, considering that the waiting duration of the first inference task is too long, which will cause the inference request corresponding to the first inference task to be delayed for a long time. Therefore, a second duration threshold can be preset, and then the inference system can use the second duration threshold to limit the waiting duration of the first inference task. Specifically, if the waiting duration of the first inference task does not reach the second duration threshold, the inference system can execute steps S24 - S25, continuously judge the complexity information corresponding to each inference task to obtain a second inference task, and aggregate the first inference task and the second inference task; at the same time, the inference system can also continuously obtain the cached inference tasks, judge the complexity information corresponding to the new inference tasks, and then continuously screen new second inference tasks according to the judgment results for task aggregation operations. If the waiting duration of the first inference task reaches the second duration threshold, the inference system executes step S26. At this time, even if the inference system still has not screened out a second inference task with a complexity distance that meets the requirements and has not completed the task aggregation, it will no longer continue with the task aggregation operation, but directly distribute the first inference task to the inference model for execution to ensure that the waiting duration of the first inference task is not too long.
[0143] Exemplarily, if before the waiting duration T of the first inference task A reaches the second waiting duration threshold T2, the inference system never screens out an inference task with an input Token length similar to that of the first inference task; then when the waiting duration T of the first inference task A reaches the second waiting duration threshold T2, the inference system no longer screens for the second inference task, but directly distributes the first inference task.
[0144] Among them, the second time threshold can be set according to historical experience or according to the actual application scenario. It can be understood that the larger the second time threshold, the longer the possible waiting duration of the first inference task, but the higher the possibility of screening out the second inference task to complete the task aggregation.
[0145] Generally speaking, the second time threshold is greater than the first time threshold. In this way, the timeout tasks whose waiting duration reaches the first time threshold can be preferentially used as the first inference tasks, and task aggregation is performed on the first inference tasks until their waiting duration reaches the second time threshold. Whether the aggregation is successful or not, they are all sent to the inference model.
[0146] In some implementation manners, the second time threshold can also be set to be equal to the first time threshold. In this way, if the waiting duration of the first inference task does not reach the first time threshold, that is, the first inference task does not time out, the inference system continuously screens the second inference tasks whose complexity distance meets the requirements for task aggregation; if the waiting duration of the first inference task reaches the first time threshold, that is, the first inference task times out, the inference system no longer continues to screen other second inference tasks. This can reduce the waiting duration of the timeout tasks and improve the response speed of the timeout tasks.
[0147] Further, in step 24, considering the limitations of the computing device, if the inference system concurrently runs multiple inference tasks in parallel, it will exceed the load of the computing device. Therefore, the inference system can limit the number of the second inference tasks to be less than or equal to a number threshold, where the number threshold is an aggregation parameter pre-loaded by the inference system. In this way, during the task aggregation process, if the number of the second inference tasks that have been screened out does not reach the number threshold, new inference tasks are continuously screened and aggregated into the aggregation task; if the number of the second inference tasks that have been screened out has reached the number threshold, other second inference tasks are no longer continuously screened. In this way, it can be avoided that the aggregation task contains too many inference tasks, resulting in the computing power resources of the computing device being unable to handle the parallel processing of these inference tasks.
[0148] By performing the task aggregation operation before the waiting duration of the first inference task reaches the second duration threshold in the embodiments of the present application, it can be avoided that the first inference task waits for a long time due to the long time consumed by performing the task aggregation operation. Therefore, the response speed of the inference request corresponding to the first inference task can be improved.
[0149] Figure 9 It is another flowchart of the inference task scheduling method provided by the embodiments of the present application.
[0150] As Figure 9 shown, the method includes the following steps S31 - S36:
[0151] S31: Cache multiple inference tasks.
[0152] S32: Determine the first inference task from multiple inference tasks.
[0153] Among them, the first inference task includes the inference tasks whose waiting duration is greater than or equal to the first time threshold.
[0154] S33: Generate an aggregation task corresponding to the first inference task from multiple inference tasks based on the task complexity information of the inference tasks.
[0155] Among them, the task complexity information is information related to the processing duration of the inference task; the aggregation task includes one or more inference tasks.
[0156] For the descriptions of steps S31 - S33, please refer to the foregoing steps S11 - S12, which will not be elaborated here.
[0157] S34: Add the aggregation task corresponding to the first inference task to the task scheduling queue.
[0158] Among them, the task scheduling queue is used to store the aggregation tasks.
[0159] In step S34, the inference system adds the aggregation task to the task scheduling queue, and the aggregation task waits in the task scheduling queue to be allocated computing power resources and executed.
[0160] In one implementation, step S36 can be directly executed after step S34.
[0161] In one implementation, step S34 further includes the following step S35:
[0162] S35: Adjust the order of each aggregation task in the task scheduling queue.
[0163] In step S35, the inference system can flexibly adjust the order of each aggregation task in the task scheduling queue according to the actual application scenario and requirements, realizing the adjustment of the execution order of the aggregation tasks, so that high-priority inference tasks and timeout tasks can be executed first, thereby optimizing the response time of the inference system.
[0164] Among them, steps S34 and S35 can be carried out simultaneously. For example, the inference system extracts the aggregation task from the task cache queue, determines the position of the aggregation task in the task scheduling queue based on the Token length and waiting duration of the inference tasks in the aggregation task, and then inserts the aggregation task into this position, without having to adjust the position of the aggregation task after putting it into the task scheduling queue.
[0165] In one implementation, step S35 includes at least one of the following steps S351 - S353:
[0166] S351: Adjust the order of each aggregation task in the task scheduling queue based on the waiting duration corresponding to each aggregation task.
[0167] In step S351, to avoid excessive waiting time for certain inference tasks, the inference system can adjust the order of each aggregation task in the task scheduling queue according to the waiting duration corresponding to each aggregation task, so as to adjust the execution order of each aggregation task.
[0168] Exemplarily, the inference system can compare the inference tasks with the longest waiting duration in each aggregation task and adjust the sorting of each aggregation task according to the comparison result; it can also calculate the average waiting duration of all inference tasks in each aggregation task, compare the average waiting durations corresponding to each aggregation task, and adjust the sorting of each aggregation task according to the comparison result. It can be understood that in the embodiments of the present application, only adjusting the sorting according to the maximum waiting duration and the average waiting duration is taken as an example. In other implementation manners, the inference system can also adopt other ways to adjust the sorting of each aggregation task. For example, it can comprehensively analyze the waiting duration of the inference tasks in each aggregation task and the corresponding priorities of the inference tasks. The longer the waiting duration and / or the higher the priority, the more forward the sorting. The present application does not make any limitations in this regard.
[0169] Through step S351 in the embodiments of the present application, the arrangement order of each aggregation task in the task scheduling queue can be adjusted, so as to realize the dynamic adjustment of the execution order of the aggregation tasks, enhance the flexibility of the inference system, and is beneficial to coping with the dynamically changing task situation. On this basis, the inference system adjusts the arrangement order based on the waiting duration, and can preferentially execute the inference tasks with longer waiting durations, thereby improving the processing ability of the system under high-load conditions.
[0170] S352: Adjust the aggregation task corresponding to the first inference task to the head of the task scheduling queue.
[0171] In step S352, considering that the waiting duration of the first inference task is greater than the first time threshold, or the waiting duration of the first inference task is the longest among all inference tasks, therefore, the inference task adjusts the aggregation task corresponding to the first inference task to the head of the task scheduling queue. In this way, the aggregation task corresponding to the first inference task can be preferentially dispatched and executed, thereby improving the response speed of the inference request corresponding to the first inference task.
[0172] In one implementation manner, step S352 includes the following steps S3521 - S3522:
[0173] S3521: When the waiting duration of the first inference task is greater than the first time threshold, adjust the aggregation task corresponding to the first inference task to the head of the task scheduling queue.
[0174] In step S3521, if the waiting duration of the first inference task is greater than the first time threshold, that is, the first inference task times out, it can be considered that the first inference task needs to be processed as soon as possible. At this time, the inference system can adjust the aggregation task corresponding to the first inference task to the head of the task scheduling queue. In this way, when the aggregation tasks are issued and executed one by one in order, this aggregation task can be preferentially executed.
[0175] In addition, if there are multiple aggregation tasks and the waiting durations of their first inference tasks are all greater than the first time threshold, the inference system can arrange the multiple aggregation tasks in several positions at the head of the task scheduling queue according to the waiting durations of the first inference tasks.
[0176] Figure 10 It is a schematic diagram of adjusting the aggregation task to the head of the queue proposed by the embodiment of the present application.
[0177] As Figure 10 shown, there are 5 aggregation tasks, denoted as Agg1 - 5. Figure 11 The horizontal axis of the histogram of Figure 10 arranges the aggregation tasks Agg1 - 5, and the vertical axis identifies the waiting durations of the first inference tasks in each aggregation task. T1 is the first time threshold. According to
[0178] Next, in combination with Figure 11 and Figure 12 it is explained how to determine Figure 11 the order of the aggregation tasks Agg1, Agg2, and Agg3 at the head of the task scheduling queue.
[0179] Figure 11 It is a schematic diagram of sorting multiple aggregation tasks at the head of the queue proposed by the embodiment of the present application.
[0180] Figure 12 It is another schematic diagram of sorting multiple aggregation tasks at the head of the queue proposed by the embodiment of the present application.
[0181] As Figure 11 shown, the inference system sorts the three aggregation tasks at the head of the queue according to the waiting durations of the first inference tasks in the foregoing aggregation tasks Agg1, Agg2, and Agg3. Figure 11 The horizontal axis of the histogram of Figure 11It can be seen that the waiting duration of the first inference task in the aggregation task Agg1 is the longest. Therefore, the aggregation task Agg1 is sorted first in the task scheduling queue and is located at the first position on the left (i.e., the head of the queue). Similarly, the aggregation tasks Agg2 and Agg3 are ranked at the 2nd and 3rd positions respectively. In this way, it is possible to avoid the first inference task in the aggregation task Agg1 waiting for execution for a long time, resulting in an overly slow response to the inference request corresponding to the first inference task.
[0182] As Figure 12 shown, the inference system sorts the three aggregation tasks at the head of the queue according to the task complexity information (such as the average input token length) of each inference task in the foregoing aggregation tasks Agg1, Agg2, and Agg3. Figure 12 The horizontal axis of the histogram in shows the aggregation tasks Agg1 to 3, and the vertical axis indicates the average input token length of each inference task in each aggregation task. According to Figure 12 it can be seen that the average input token length of each inference task in the aggregation task Agg2 is the smallest. Therefore, the aggregation task Agg2 is sorted first in the task scheduling queue and is located at the first position on the left (i.e., the head of the queue). Similarly, the aggregation tasks Agg1 and Agg3 are ranked at the 2nd and 3rd positions respectively. In this way, the inference system can process and complete more aggregation tasks faster, so the number of inference tasks waiting for execution can be reduced.
[0183] It can be understood that the embodiments of the present application only take the head sorting of the task scheduling sequence determined according to the waiting duration of the first inference task and the task complexity information as an example. In other implementation manners, the inference system can also use other methods to determine the head sorting. For example, it can comprehensively analyze the waiting duration of the first inference task in each aggregation task and the priority corresponding to the first task. The longer the waiting duration and / or the higher the priority, the more forward the sorting. The present application does not make any limitations in this regard.
[0184] S3522: Otherwise, place the aggregation task at the end of the task scheduling queue.
[0185] In step S3522, contrary to the foregoing step S3521, if the waiting duration of the first inference task in a certain aggregation task is less than or equal to the first time threshold, the inference system can place the aggregation task at the end of the task scheduling queue. The specific method, principle, and beneficial effects can be analogized to the foregoing step S3521 and will not be elaborated here.
[0186] In the embodiment of the present application, through steps S3521 and S3522, the first inference tasks that time out are placed at the head of the queue, and the first inference tasks that do not time out are placed at the tail of the queue, which can ensure that the first inference tasks that time out are preferentially executed, thereby avoiding the first inference tasks from timing out for too long. In this way, the response speed of the inference requests corresponding to the timeout tasks can be improved, the overall performance of the inference system can be enhanced, and the user experience can be improved.
[0187] S353: Based on the dependency relationships between the aggregation tasks, adjust the order of the aggregation tasks in the task scheduling queue so that each inference task in the aggregation tasks sorted later does not depend on each inference task in the aggregation tasks sorted earlier.
[0188] In step S353, to prevent some inference tasks sorted earlier from requiring the inference results of inference tasks sorted later during execution, the inference system can adjust the order of the aggregation tasks in the task scheduling queue according to the dependency relationships between the aggregation tasks, specifically for the dependency relationships between the inference tasks in different aggregation tasks, so as to adjust the execution order of the aggregation tasks.
[0189] Exemplarily, there is an inference task A in the aggregation task Agg1 for perceiving surrounding obstacles; there is an inference task B in the aggregation task Agg2 for path planning based on the perception result of the inference task A. Then the inference system can adjust the order of the aggregation task Agg1 and the aggregation task Agg2 in the task scheduling queue according to the dependency relationship between the inference task A and the inference task B, and arrange the aggregation task Agg1 before the aggregation task Agg2. In this way, the aggregation task Agg1 can be issued and executed first, and then the aggregation task Agg2 can be issued and executed. Therefore, the perception result of the inference task A can be used during the execution of the inference task B.
[0190] In the embodiment of the present application, through step S353, it can be ensured that the aggregation tasks are executed in order based on the dependency relationships, and it is avoided that the inference tasks sorted earlier require the inference results of the inference tasks sorted later during execution, resulting in logical errors or task execution blocking.
[0191] S36: Issue the inference tasks in the aggregation task in parallel.
[0192] In step S36, the inference system sequentially takes out the aggregation tasks from the head of the task scheduling queue in the first-in-first-out order and issues them to the inference model, and then uses the inference model to process the aggregation tasks. Among them, the inference model can be a large model, and the present application does not limit the model structure of the inference model.
[0193] After that, the inference system can use the inference model to perform each inference task in the parallel execution aggregation task, output the inference results corresponding to each inference task, and return the inference results to the corresponding client, thereby completing the response based on the inference request. In this way, the inference model can process multiple inference tasks simultaneously each time, effectively improving the computing power utilization rate of the computing device where the inference system is located and the inference efficiency of the inference system.
[0194] By designing a sequence adjustment strategy for the task scheduling queue in this embodiment, the parallelism of the inference system when performing inference tasks and the resource utilization rate of the computing device where the inference system is located can be improved, thereby improving the response speed and throughput of the inference system.
[0195] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. For example, the steps recorded in the above embodiments can be executed in parallel, sequentially, or in different orders, as long as the results expected by the technical solutions disclosed in the present invention can be achieved. The present invention makes no limitation thereto.
[0196] Corresponding to the embodiment of the foregoing inference task scheduling method, the present application also provides an embodiment of an inference task scheduling device.
[0197] Figure 13 It is a schematic diagram of an inference task scheduling device provided by an embodiment of the present application.
[0198] As Figure 13 shown, the inference task scheduling device 1300 may include a cache module 1310, an aggregation module 1320, and a distribution module 1330.
[0199] The cache module 1310 is used to cache multiple inference tasks; the aggregation module 1320 is used to determine a first inference task from multiple inference tasks; wherein, the first inference task includes inference tasks with a waiting duration greater than or equal to a first time threshold; and, generate an aggregation task corresponding to the first inference task from multiple inference tasks based on the task complexity information of the inference tasks; wherein, the task complexity information is information related to the processing duration of the inference tasks; the aggregation task includes one or more inference tasks; the distribution module 1330 is used to distribute the inference tasks in the aggregation task in parallel.
[0200] In a possible implementation, the aggregation module 1320 includes a second inference task determination unit configured to determine a second inference task from multiple inference tasks based on the first task complexity information of the first inference task; wherein, the complexity distance between the second task complexity information of the second inference task and the first task complexity information of the first inference task is less than or equal to a distance threshold; and an aggregation unit configured to aggregate the first inference task and the second inference task to form an aggregated task.
[0201] In a possible implementation, the task complexity information of an inference task includes at least one of a token length and a task type.
[0202] In a possible implementation, the aggregation module 1320 includes a first inference task determination unit configured to, when there is no inference task with a waiting duration greater than or equal to a first time threshold among multiple inference tasks, use the inference task with the maximum waiting duration among the multiple inference tasks as the first inference task.
[0203] In a possible implementation, the inference task scheduling device 1300 includes a scheduling module configured to: add the aggregated task corresponding to the first inference task to a task scheduling queue; wherein, the task scheduling queue is used to store aggregated tasks.
[0204] In a possible implementation, the scheduling module is configured to: adjust the order of each aggregated task in the task scheduling queue based on the waiting duration corresponding to each aggregated task.
[0205] In a possible implementation, the scheduling module is configured to: adjust the aggregated task corresponding to the first inference task to the head of the task scheduling queue.
[0206] In a possible implementation, the aggregation module 1320 is configured to: determine whether the waiting duration of the first inference task is less than a second time threshold;
[0207] When the waiting duration of the first inference task is less than the second time threshold, generate an aggregated task corresponding to the first inference task from multiple inference tasks based on the complexity information of the inference tasks; wherein, the second time threshold is greater than the first time threshold; the issuing module 1330 is configured to: when the waiting duration of the first inference task is greater than or equal to the second time threshold, issue the first inference task for execution.
[0208] In a possible implementation, the caching module 1310 is configured to, in response to an inference request, generate an inference task corresponding to the inference request; and place the inference task in a caching queue.
[0209] In a possible implementation, there is no dependency relationship between the inference tasks in an aggregated task.
[0210] In a possible implementation, the scheduling module is used for:
[0211] Based on the dependency relationships among the aggregation tasks, adjust the order of the aggregation tasks in the task scheduling queue, so that each inference task in the aggregation tasks sorted later does not depend on each inference task in the aggregation tasks sorted earlier.
[0212] Figure 14 It is a schematic diagram of a computing device provided by an embodiment of the present application.
[0213] As Figure 14 shown, the computing device 1400 includes: a processor 1401 and a memory 1402. Exemplarily, the computing device 1400 may further include: a communications interface 1403 and a communication bus 1404.
[0214] Wherein, the processor 1401, the memory 1402, and the communications interface 1403 complete mutual communication through the communication bus 1404. The communications interface 1403 may include a type of device such as a transmitter and a receiver, and is used to communicate with other devices or communication networks, and may be a wired interface (port), such as a fiber distributed data interface (FDDI), a gigabit ethernet (GE).
[0215] In some embodiments, the processor 1401 is used to execute the program 1405, and specifically may execute the relevant steps in the above-mentioned embodiment of the inference task execution method. Specifically, the program 1405 may include program code, and the program code includes computer-executable instructions.
[0216] Exemplarily, the processor 1401 may be a central processing unit CPU, or a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement some embodiments of the present application. The computing device 1400 may include one or more processors, which may be the same type of processor, such as one or more CPUs; or may be different types of processors, such as one or more CPUs and one or more ASICs. The CPU may be a single-core CPU or a multi-core CPU.
[0217] In some embodiments, the memory 1402 is used to store the program 1405. The memory 1402 may include high-speed random access memory (RAM), or may include non-volatile memory (NVM), such as at least one disk memory.
[0218] The program 1405 can be specifically called by the processor 1401 to enable the computing device 1400 to execute the operations of the inference task execution method.
[0219] Some embodiments of the present application provide a computer-readable storage medium storing at least one executable instruction. When the executable instruction runs on the computing device 1400, it causes the computing device 1400 to execute the inference task scheduling method in the above embodiments.
[0220] The executable instruction can be specifically used to cause the computing device 1400 to execute the operations of the inference task scheduling method.
[0221] For example, the computer-readable storage medium can be a read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0222] Some embodiments of the present application provide a chip system applied to a server. The chip system includes one or more interface circuits and one or more processors. The interface circuits and the processors are interconnected by lines. The interface circuit is used to receive signals from the memory of the server and send signals to the processor. The signals include computer instructions stored in the memory. When the server processor executes the computer instructions, the server executes each step in the inference task scheduling method shown in the above method embodiments.
[0223] The beneficial effects that can be achieved by the readable storage medium provided in some embodiments of the present application can refer to the beneficial effects in the corresponding inference task scheduling method provided above, and will not be elaborated here.
[0224] It should be noted that in the application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0225] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments.
[0226] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus or device), or in conjunction with these instruction execution systems, apparatuses or devices.
[0227] For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device.
[0228] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory, an optical fiber device, and a portable compact disc read-only memory (CDROM).
[0229] In addition, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation or, if necessary, other suitable processing, and then stored in a computer memory. It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof.
[0230] In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, etc. The above-described embodiments are only specific embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present application shall be included in the protection scope of the present application.
Claims
1. A method for scheduling inference tasks, characterized in that, The method includes: Caching multiple inference tasks; Determining a first inference task from the multiple inference tasks; wherein, the first inference task includes an inference task with a waiting duration greater than or equal to a first time threshold; Generating an aggregated task corresponding to the first inference task from the multiple inference tasks based on the task complexity information of the inference tasks; wherein, the task complexity information is information related to the processing duration of the inference tasks; the aggregated task includes one or more inference tasks; Parallelly distributing the inference tasks in the aggregated task.
2. The method according to claim 1, characterized in that The generating an aggregated task corresponding to the first inference task from the multiple inference tasks based on the task complexity information of the inference tasks includes: Determining a second inference task from the multiple inference tasks based on the first task complexity information of the first inference task; wherein, the complexity distance between the second task complexity information of the second inference task and the first task complexity information of the first inference task is less than or equal to a distance threshold; the first task complexity information is the task complexity information of the first inference task; the second task complexity information is the task complexity information of the second inference task; Aggregating the first inference task and the second inference task to form the aggregated task.
3. The method according to claim 1 or 2, characterized in that, The task complexity information of the inference tasks includes at least one of the token length and the task type.
4. The method according to any one of claims 1 to 3, characterized in that The method further includes: In the case that there is no inference task with a waiting duration greater than or equal to the first time threshold among the multiple inference tasks, using the inference task with the maximum waiting duration among the multiple inference tasks as the first inference task.
5. The method according to any one of claims 1-4, characterized in that, After generating the aggregated task corresponding to the first inference task from the multiple inference tasks based on the inference task complexity information, the method further includes: Adding the aggregated task corresponding to the first inference task to a task scheduling queue; wherein, the task scheduling queue is used to store the aggregated task.
6. The method according to claim 5, wherein After adding the aggregated task corresponding to the first inference task to the task scheduling queue, the method further includes: Adjusting the order of the aggregated tasks in the task scheduling queue based on the waiting duration corresponding to each aggregated task; Adjusting the aggregated task corresponding to the first inference task to the head of the task scheduling queue.
7. The method according to any one of claims 1-6, characterized in that, After determining the first inference task from the multiple inference tasks, the method further includes: Determining whether the waiting duration of the first inference task is less than a second time threshold; In the case that the waiting duration of the first inference task is less than the second time threshold, generating an aggregated task corresponding to the first inference task from the multiple inference tasks based on the complexity information of the inference tasks; wherein, the second time threshold is greater than the first time threshold; In the case that the waiting duration of the first inference task is greater than or equal to the second time threshold, distributing the first inference task for execution.
8. The method according to any one of claims 1-7, characterized in that There is no dependency relationship between the inference tasks in the aggregated task.
9. The method according to claim 8, wherein The method further includes: Based on the dependencies between the aggregation tasks, adjust the order of the aggregation tasks in the task scheduling queue so that each inference task in the aggregation tasks sorted later does not depend on each inference task in the aggregation tasks sorted earlier.
10. A computing device, characterized in that, The computing device includes a memory and a processor; the memory is coupled to the processor; The memory is used to store computer instructions; The processor is used to execute the computer instructions so that the computing device executes the inference task scheduling method according to any one of claims 1 to 9.