Task processing method and device
By identifying and utilizing the future idle time of idle processing units to process low-priority tasks, the problem of uneven resource utilization in AI model inference clusters is solved, achieving efficient resource utilization and task processing.
Patent Information
- Application Number
- CN202410799780.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-21
- Filing Date
- 2024-06-19
- Publication Date
- 2025-11-21
AI Technical Summary
During peak and off-peak periods, the resource utilization of the AI model inference cluster is uneven, resulting in delays in task processing during peak periods and waste of resources during off-peak periods, which affects the guarantee of service level agreements.
By identifying idle processing units in the cluster and predicting their future idle time, and if the future idle time is greater than or equal to the time required for low-priority tasks, the idle processing units are instructed to process low-priority tasks, thereby completing the processing of low-priority tasks before high-priority tasks arrive and improving resource utilization.
While ensuring timely processing of high-priority tasks, it improved resource utilization, reduced task processing latency, and optimized cluster resource allocation.
Smart Images

Figure CN120994357A_ABST
Abstract
Description
[0001] This application claims priority to Chinese Patent Application No. 202410634507.6, filed on May 21, 2024, entitled “A Method and Apparatus for Utilizing Resources”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer technology, and in particular to a task processing method and apparatus. Background Technology
[0003] With the development of artificial intelligence (AI), AI models are becoming increasingly intelligent, enabling them to save manpower and improve business processing efficiency. Therefore, more and more industries are beginning to apply AI models to their relevant operations.
[0004] The application of AI models involves model inference and model training. Model inference, which utilizes AI models to process business logic, directly interacts with users and requires a high service level agreement (SLA). To ensure the SLA of model inference, significant resources are allocated to the cluster used for its execution. Typically, business operations have peak and off-peak periods. During peak periods, the workload for model inference is large, fully utilizing cluster resources. During off-peak periods, the workload for model inference is small, resulting in low resource utilization and wasted resources. Summary of the Invention
[0005] This application provides a task processing method and apparatus that can improve the resource utilization of a cluster.
[0006] In a first aspect, a task processing method is provided, which is applied to a control device in a computing system. The computing system further includes a cluster of processing units, wherein, under the instruction of the control device, processing units in the cluster of processing units can process different tasks at different times. The method includes: identifying at least one idle processing unit in the cluster of processing units; predicting the future idle duration of at least one idle processing unit based on historical information of the cluster of processing units processing a first task, wherein the future idle duration is the duration during which at least one idle processing unit does not need to process the first task; and, if the future idle duration is greater than or equal to the required duration of a second task, instructing at least one idle processing unit to process the second task; wherein the priority of the second task is lower than the priority of the first task, and the required duration of the second task is the duration required for at least one idle processing unit to process the second task.
[0007] In this task, the first task is a high-priority task, and the second task is a low-priority task. For example, the first task can be a model inference task, and the second task can be a model training task.
[0008] An idle processing unit is a processing unit in the cluster that is currently idle, meaning it is idle when the cluster is not processing high-priority tasks. In other words, an idle processing unit is a processing unit in the cluster that is idle when there are no high-priority tasks to process or when it is in a waiting state.
[0009] This method identifies idle processing units in the cluster that are not currently handling high-priority tasks, and predicts the duration during which these idle units will not need to process high-priority tasks in the future. Then, if this duration is greater than or equal to the duration required to process low-priority tasks, the idle unit is instructed to handle the low-priority tasks. This allows low-priority tasks to be processed before the idle unit is needed to handle high-priority tasks, ensuring that high-priority tasks can be processed promptly when they arrive. In this way, resource utilization is improved while ensuring that future high-priority tasks are processed in a timely manner.
[0010] In one possible implementation, the method further includes: identifying the idle processing capability of at least one idle processing unit; and instructing at least one idle processing unit to process the second task if the future idle duration is greater than or equal to the required duration of the second task, including: instructing at least one idle processing unit to process the second task if the future idle duration is greater than or equal to the required duration of the second task and the idle processing capability is greater than or equal to the required processing capability of the second task, wherein the required processing capability of the second task is the capability required to process the second task.
[0011] The processing capacity of the idle processing unit is its processing capacity in an idle state, and it is also the maximum available processing capacity of the processing unit. If the processing capacity required for the second task is less than or equal to the idle processing capacity of the idle processing unit, the idle processing unit is instructed to process the second task, thereby ensuring that the processing of the second task is completed before the end of the future idle period.
[0012] In one possible implementation, the second task is the task that requires the most processing power among the multiple tasks, wherein the priority of each task among the multiple tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each task among the multiple tasks.
[0013] If there are multiple low-priority tasks currently waiting to be processed by the processing unit cluster, and the required duration of each task is less than or equal to the future idle duration of the idle processing unit, and the required processing capacity of each task is less than or equal to the processing capacity of the idle processing unit, then the idle processing unit is instructed to process the task with the highest required processing capacity among these multiple tasks. In this way, the computing resources of the idle processing unit can be fully utilized, improving resource utilization.
[0014] In one possible implementation, the second task is the task with the longest processing time among the multiple tasks; wherein the priority of each task among the multiple tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each task among the multiple tasks.
[0015] If there are multiple low-priority tasks currently waiting to be processed by the processing unit cluster, and the required processing time for each of these tasks is less than or equal to the future idle time of the idle processing unit, then the idle processing unit is instructed to process the task with the longest processing time among these tasks. This fully utilizes the computing resources of the idle processing unit, improving resource utilization.
[0016] In one possible implementation, the second task is the highest priority task among the multiple tasks; wherein the priority of each task among the multiple tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each task among the multiple tasks.
[0017] If there are multiple low-priority tasks currently waiting to be processed by the processing unit cluster, and the required time for each of these tasks is less than or equal to the future idle time of the idle processing unit, then if these tasks include tasks of different priorities, the idle processing unit is instructed to process the highest-priority task among them. This ensures that the highest-priority task is processed first.
[0018] In one possible implementation, the method further includes: when there is a third task to be processed and the processing unit cluster does not have enough idle processing units to process the third task, removing the fourth task currently being processed by the processing unit cluster from the processing unit cluster, so that the processing unit cluster has enough idle processing units to process the third task; wherein the priority of the fourth task is lower than the priority of the third task.
[0019] In this implementation, when a high-priority task arrives, if the processing unit cluster does not have enough idle processing units to handle the high-priority task, it can evict the currently processing low-priority task and free up processing units, thereby ensuring that the high-priority task is processed in a timely manner.
[0020] In one possible implementation, the fourth task is the lowest priority task among the tasks currently being processed by the processing unit cluster.
[0021] In this implementation, the lowest priority task is evicted first, thereby ensuring the rational use of resources in the processing unit cluster.
[0022] In one possible implementation, the fourth task is the task with the longest remaining processing time among the tasks currently being processed by the processing unit cluster.
[0023] In this implementation, tasks with the longest remaining processing time are evicted first, thus avoiding the evicting of tasks that are about to be completed and reducing the cost of task eviction.
[0024] In one possible implementation, the fourth task is the task with the shortest processing time among the tasks currently being processed by the processing unit cluster.
[0025] In this implementation, tasks with the shortest processing time are evicted first, thus avoiding the eviction of tasks that are about to be completed and reducing the cost of task eviction.
[0026] In one possible implementation, the fourth task is the task with the fewest interruptions among the tasks currently being processed by the processing unit cluster.
[0027] In this implementation, the task with the fewest interruptions is prioritized for eviction, thereby avoiding multiple interruptions and thus preventing impact on user experience.
[0028] In one possible implementation, the fourth task is the task that occupies the best-performing processing unit among the tasks currently being processed by the processing unit cluster.
[0029] In this implementation, the processing unit with the best performance can be freed up, so that it can be used to process high-priority tasks, thus ensuring the processing effect and efficiency of high-priority tasks.
[0030] In one possible implementation, the processing units in the processing unit cluster access different networks when processing different tasks.
[0031] In this way, network isolation between different tasks is achieved, ensuring the security of task processing.
[0032] Secondly, a control device is provided, wherein the computing system in which the device is located further includes a cluster of processing units, wherein, under the instruction of the device, the processing units in the cluster of processing units can process different tasks at different times; the device includes: an identification module for identifying at least one idle processing unit in the cluster of processing units; a prediction module for predicting the future idle duration of at least one idle processing unit based on historical information of the cluster of processing units processing a first task, wherein the future idle duration is the duration during which at least one idle processing unit does not need to process the first task; and an instruction module for instructing at least one idle processing unit to process a second task if the future idle duration is greater than or equal to the required duration of a second task; wherein the priority of the second task is lower than the priority of the first task, and the required duration of the second task is the duration required for at least one idle processing unit to process the second task.
[0033] In one possible implementation, the identification module is further configured to: identify the idle processing capability of at least one idle processing unit; the instruction module is configured to: instruct at least one idle processing unit to process the second task if the future idle duration is greater than or equal to the required duration of the second task and the idle processing capability is greater than or equal to the required processing capability of the second task, wherein the required processing capability of the second task is the capability required to process the second task.
[0034] In one possible implementation, the second task is the task that requires the most processing power among the multiple tasks, wherein the priority of each task among the multiple tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each task among the multiple tasks.
[0035] In one possible implementation, the second task is the task with the longest processing time or the highest priority among the multiple tasks; wherein the priority of each task among the multiple tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each task among the multiple tasks.
[0036] In one possible implementation, the instruction module is used to: remove the fourth task currently being processed by the processing unit cluster from the processing unit cluster when there is a third task to be processed and the processing unit cluster does not have enough idle processing units to process the third task, so that the processing unit cluster has enough idle processing units to process the third task; wherein the priority of the fourth task is lower than the priority of the third task.
[0037] In one possible implementation, the fourth task is the lowest priority task among the tasks currently being processed by the processing unit cluster; or,
[0038] The fourth task is to process the task with the longest remaining processing time among the tasks currently being processed by the cluster; or,
[0039] The fourth task is to process the task with the fewest interruptions among the tasks currently being processed by the unit cluster; or,
[0040] The fourth task is to process the task currently being processed by the unit cluster that occupies the best performance processing unit.
[0041] In one possible implementation, the processing units in the processing unit cluster access different networks when processing different tasks.
[0042] In one possible implementation, the first task is model inference, and the second task is model training.
[0043] Thirdly, a computing device cluster is provided, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method provided in the first aspect.
[0044] Fourthly, a computer-readable storage medium is provided, including computer program instructions that, when executed by a cluster of computing devices, execute the method provided in the first aspect.
[0045] Fifthly, a computer program product containing instructions is provided, which, when executed by a cluster of computer devices, causes the cluster of computer devices to perform the method provided in the first aspect.
[0046] The beneficial effects of the second to fifth aspects can be referred to the introduction of the beneficial effects of the first aspect above, and will not be repeated here. Attached Figure Description
[0047] Figure 1 A schematic diagram of a computing system provided in an embodiment of this application;
[0048] Figure 2 A schematic diagram of a computing system provided in an embodiment of this application;
[0049] Figure 3 A schematic diagram of a control device provided in an embodiment of this application;
[0050] Figure 4 A schematic diagram of a control device provided in an embodiment of this application;
[0051] Figure 5 A schematic diagram illustrating a task processing method provided in an embodiment of this application;
[0052] Figure 6 A schematic diagram illustrating a task processing method provided in an embodiment of this application;
[0053] Figure 7 A flowchart illustrating a task processing method provided in an embodiment of this application;
[0054] Figure 8 This is a schematic diagram of the structure of a control device provided in an embodiment of this application;
[0055] Figure 9 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0056] Figure 10 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;
[0057] Figure 11 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application. Detailed Implementation
[0058] The solutions provided in the embodiments of this application will now be described with reference to the accompanying drawings. In the embodiments of this application, "multiple" refers to two or more objects, and "various types" refers to two or more types. Terms such as "first," "second," etc., are only used to distinguish similar objects and are not necessarily used to describe a specific order or number of objects.
[0059] To facilitate understanding of the solutions provided in the embodiments of this application, the technical terms that may be involved in the embodiments of this application will be introduced first.
[0060] Artificial intelligence (AI) is a branch of computer science that attempts to understand the nature of intelligence and to produce new intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems.
[0061] Machine learning (ML) is a technique that specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structures to continuously improve the computer's performance. It is the core of artificial intelligence and the fundamental way to enable computers to gain intelligence. Machine learning involves inputting data into machine learning algorithms and allowing them to learn from that data, thereby making accurate predictions or classifications of new data.
[0062] Checkpoint: refers to the periodic saving of model weights or other information during model training.
[0063] A Service Level Agreement (SLA) is a mutually agreed-upon agreement or contract between a service provider and its customer regarding the quality, level, and performance of the service. An SLA may include multiple metrics, such as timeliness, reliability, response time, and repair time.
[0064] Model training refers to the task of optimizing a model using machine learning algorithms and large amounts of data to enable it to make better predictions and decisions. Model training is an offline task with relatively low SLA requirements.
[0065] Model inference refers to the task of using a trained model to infer and judge unknown information under given conditions, leveraging the intelligence gained through model training. Model inference is an online task; the model directly faces the customer and interacts with the user. Therefore, model inference has high SLA requirements.
[0066] A cluster refers to a group of processing units used to execute tasks. A common example is a Kubernetes cluster. Different clusters can be used to execute different tasks. For instance, a cluster executing model inference tasks is called an inference cluster, and a cluster executing model training tasks is called a training cluster. A cluster can consist of multiple processing units. For example, a cluster can consist of one or more compute nodes. A compute node can act as a processing unit, or a single compute node can include multiple processing units.
[0067] A processing unit is the smallest computational unit that executes a task. A processing unit processes one task at a time, not two tasks simultaneously. A processing unit can process different tasks at different times. A processing unit can be hardware such as a graphics processing unit (GPU) or a neural processing unit (NPU), or it can be a virtual computing instance such as a container.
[0068] Compute nodes are worker nodes in a cluster, providing the runtime environment for tasks and responsible for running and managing containerized applications. Compute nodes can be hardware such as servers or virtual computing devices such as virtual machines (VMs).
[0069] A Kubernetes cluster refers to a cluster managed by Kubernetes. Kubernetes is a container orchestration tool. It facilitates the scheduling and orchestration of containers. Kubernetes treats a large number of servers as a single giant server, running applications on that large server. Regardless of the number of servers in a Kubernetes cluster, the method for deploying applications on Kubernetes is the same.
[0070] Compared to model training, model inference has higher SLA requirements and exhibits peak and trough phenomena. To ensure SLA during peak periods, inference clusters are configured with a larger number of processing units. During peak periods, these processing units in the inference cluster may be fully utilized. However, during trough periods, a large number of processing units in the inference cluster remain idle, leading to resource waste.
[0071] In one approach, during the trough of model inference, some processing units in the inference cluster are migrated to the training cluster, allowing them to participate in model training and improving resource utilization. During the peak of model inference, these processing units are migrated back from the training cluster to the inference cluster to ensure the SLA of model inference. Migrating processing units between different clusters involves operations such as removing and adding processing units, which is time-consuming. Since processing units do not process tasks during migration, this approach offers limited improvement in resource utilization. Furthermore, when there is a sudden surge in model inference tasks, the processing units cannot be migrated back to the inference cluster in a timely manner, making it difficult for the inference cluster to handle such surges.
[0072] In one approach, the priority of model inference tasks is set higher than that of model training tasks. Tasks are then added to a scheduling queue based on their priority, and processed sequentially. While this method allows model inference tasks to be processed before model training tasks, if a model inference task arrives during the processing of model training tasks, it must wait for the model training tasks to finish before it can begin processing. This increases the latency of the model inference task and negatively impacts its performance level (SLA).
[0073] This application provides a task processing method. In this method, idle processing units in the cluster can be identified after meeting the processing demands of high-priority tasks, and the duration for which these idle processing units will not need to process high-priority tasks in the future can be predicted. Then, if this duration is greater than or equal to the duration required to process low-priority tasks, the idle processing unit is instructed to process the low-priority tasks. In this way, while improving resource utilization, it ensures that future high-priority tasks are processed in a timely manner.
[0074] Next, the task processing method provided in the embodiments of this application will be described in detail.
[0075] Figure 1 A computing system 100 for implementing this method is shown. The computing system 100 includes a control device 110 and a cluster of processing units 120. The cluster of processing units 120 may be simply referred to as cluster 120. Cluster 120 may include multiple processing units such as processing unit 121 and processing unit 122. The control device 110 is used to instruct the processing units in cluster 120 to perform tasks. Under the instruction of the control device 110, the processing units can perform different tasks. Specifically, a processing unit may perform the same task at the same time, or it may perform different tasks at different times.
[0076] In some embodiments, such as Figure 2 As shown, the control device 110 can receive various tasks initiated by the user through the console, such as type A tasks and type B tasks. Type A tasks have a higher priority than type B tasks. For example, type A tasks can be online tasks, such as model inference tasks. Type B tasks can be offline tasks, such as model training tasks.
[0077] The control device 110 can instruct the processing units in the processing cluster 120 to process different tasks according to the task processing method provided in the embodiments of this application.
[0078] In some embodiments, such as Figure 2 As shown, in order to achieve the aforementioned objective, the control device 110 may include a resource manager 111 and an exterminator 112.
[0079] Resource Manager 111 can calculate, based on resource allocation policies, the processing units available for handling Class B tasks, and the duration for which each processing unit can handle Class B tasks. These processing units, also known as idle processing units, are specifically the idle processing units available to cluster 120 when it is not processing Class A tasks; that is, idle processing units that meet the processing requirements of Class A tasks. The duration for which these processing units can handle Class B tasks is also called the future idle duration of the idle processing unit; specifically, it is the duration for which the idle processing unit will not be needed to handle Class A tasks in the future.
[0080] The resource allocation strategy can involve identifying idle processing units and then calculating their future idle duration. This can be achieved by identifying currently idle processing units based on the operational status of each processing unit in cluster 120. Specifically, an idle processing unit refers to a currently idle processing unit. Then, based on the historical information of cluster 120's processing of type A tasks, the future idle duration of each idle processing unit is identified. This historical information can include the historical arrival time of type A tasks at cluster 120 (i.e., the historical time cluster 120 began processing type A tasks), and the load information of cluster 120 when processing type A tasks. Based on the load information of cluster 120 when processing type A tasks, it can be determined whether an idle processing unit needs to process type A tasks. If an idle processing unit does not need to process type A tasks, the future start time of processing type A tasks by the idle processing unit is identified based on the historical arrival time of type A tasks at cluster 120. Therefore, the future idle duration of each idle processing unit can be calculated.
[0081] Resource allocation strategies can also identify task priorities and determine whether a task is a Class A task. In one example, each task corresponds to relevant user information for that task. In one example, the relevant information for a task could be its priority. Thus, the task's priority can be directly used to determine whether it is a Class A task. In another example, the relevant task information could include real-time requirements, response latency requirements, task tag classification, SLA, etc. The task's priority can be calculated based on its custom information. Then, based on the task's priority, it can be determined whether the task is a Class A task. For example, the relevant task information is custom information, such as information defined by the user for that task.
[0082] Resource Manager 111 can instruct the idle processing unit to process type B tasks within a future idle period based on a task filling strategy. The task filling strategy instructs the idle processing unit to process tasks whose required execution time is less than or equal to the future idle period of the idle processing unit. Here, the required execution time of the task refers to the time needed for the idle processing unit to process the task. This ensures that the idle processing unit completes processing type B tasks before or when type A tasks are required to be processed, so that type A tasks do not need to wait for execution, and type B tasks do not need to be frequently interrupted.
[0083] In some embodiments, there are multiple Class B tasks currently pending processing, and the required duration of each of these tasks is less than or equal to the future idle time of the idle processing unit. In this case, the task filling strategy instructs the idle processing unit to process the task with the longest required duration among the multiple tasks. This fully utilizes the computing resources of the idle processing unit, thereby improving resource utilization.
[0084] In some embodiments, the task filling strategy instructs idle processing units to process tasks whose required processing power is less than or equal to the processing power of the idle processing unit. Here, the processing power of the idle processing unit refers to its processing power in an idle state, and can be called its idle processing capacity. The idle processing capacity of a processing unit is also its maximum available processing power. In one example of this embodiment, there are multiple Class B tasks currently pending processing, and the required processing power of each of these tasks is less than or equal to the processing power of the idle processing unit. In this case, the task filling strategy instructs the idle processing unit to process the task with the highest required processing power among these multiple tasks. This fully utilizes the computing resources of the idle processing unit, thereby improving resource utilization.
[0085] Resource Manager 111 can guarantee the quality of service (QoS) of tasks based on service assurance policies, such as Service Level Agreements (SLAs). Specifically, tasks of type A and type B are different categories of tasks; for example, type A tasks are model inference tasks, and type B tasks are model training tasks. Type B tasks can be further prioritized, meaning different types of type B tasks may have different priorities. There are currently multiple type B tasks to be processed, and each of these tasks has a lower priority than the aforementioned type A tasks, and the required duration of each of these tasks is less than the aforementioned idle time. In this case, the service assurance policy instructs the idle processing unit to process the highest-priority task among these multiple tasks, thereby ensuring that the highest-priority task is processed first and guaranteeing its QoS. For example, Resource Manager 111 can calculate task priorities based on relevant task information. This information may include task real-time requirements, response latency requirements, task label classification, SLAs, etc. It can calculate the priorities of multiple tasks based on relevant task information and sort the priorities of these multiple tasks. The relevant task information may include real-time requirements, response latency requirements, task label classification, SLAs, etc.
[0086] Resource Manager 111 can provide network plane isolation capabilities for tasks based on network isolation policies to ensure isolation between different tasks. When instructing a processing unit to handle different tasks, the processing unit is instructed to access different networks. That is, the processing unit uses different networks when handling different tasks. For example, the processing unit cluster 120 may include network C1 and network C2. Network C1 corresponds to task type A, and network C2 corresponds to task type B. When instructing a processing unit to handle task type A, the processing unit is instructed to access network C1. When instructing a processing unit to handle task type B, the processing unit is instructed to access network C2. For example, each processing unit has multiple network interfaces, and different network interfaces access different networks. When instructing a processing unit to handle a task, the processing unit is instructed to enable the network interface corresponding to that task and disable other network interfaces, thereby enabling the processing unit to access the network corresponding to that task.
[0087] In some embodiments, the processing unit is connected to network C1 by default, so that it does not need to switch networks when a type A task arrives, thus improving the timeliness of processing type A tasks. When the processing unit is instructed to process a type B task, the network connected to the processing unit is switched from network C1 to network C2. When the type B task is completed, the network connected to the processing unit is switched from network C2 to network C1.
[0088] In some embodiments, such as Figure 3 As shown, the control device 110 also includes a control plane 113. The resource manager 111 implements related functions through the control plane 113. For example, the resource manager 111 can formulate specific task processing decisions based on resource allocation policies, task filling policies, service guarantee policies, or network isolation policies, and then send the task processing decisions to the control plane 113. Based on the task processing decisions, the control plane 113 controls the processing units in the cluster 120 to process the relevant tasks.
[0089] The required processing time for Category B tasks is calculated based on experience or experiments and may not be entirely accurate. Furthermore, unexpected events such as malfunctions may occur when the processing unit handles Category B tasks, potentially resulting in unfinished tasks by the end of the aforementioned future idle period. To address this, when the future idle period ends, the expulsion device 112 will remove Category B tasks from the processing unit based on a resource reclamation strategy. Here, the processing unit refers to the one that processes Category B tasks based on the resource relinquishment strategy.
[0090] Resource reclamation strategies can instruct processing units to evict tasks according to specified eviction methods. These eviction methods can be specified by the user. For example, there are two eviction methods: graceful eviction and direct eviction. Graceful eviction means that the processing unit performs at least one checkpoint on the task before eviction to preserve its intermediate state. When the task is processed again, processing can begin from this intermediate state, improving processing efficiency. Direct eviction means directly deleting the task to quickly reclaim resources.
[0091] As mentioned above, the future idle time of idle processing units is predicted based on historical information of type A tasks. However, unexpected events may occur, causing a sudden surge in the workload of type A tasks, and cluster 120 may not have enough idle processing units. In this case, it is necessary to remove currently processing type B tasks from the processing units so that cluster 120 has enough idle processing units to handle the sudden surge in type A tasks.
[0092] The expulsion device 112 can select a processing unit based on a conflict resolution strategy and instruct the selected processing unit to expel the task being processed.
[0093] In some embodiments, such as Figure 4 As shown, the conflict resolution strategy may include a service guarantee eviction strategy. As mentioned above, different B-type tasks may have different priorities. There are multiple B-type tasks currently being processed in cluster 120. The service guarantee strategy is to evict the lowest-priority task among these multiple tasks. That is, based on the service guarantee strategy, the eviction device 112 identifies the lowest-priority task among the multiple tasks and then instructs the processing unit handling the lowest-priority task to evict that task.
[0094] In one instance of this embodiment, if the multiple tasks have the same priority, the task with the fewest interruptions among the multiple tasks will be evictped.
[0095] In some embodiments, such as Figure 4 As shown, conflict resolution strategies can include cost-based eviction policies. There are multiple Class B tasks currently being processed in cluster 120. The cost-based eviction policy evicts the task with the fewest interruptions, the task with the longest remaining processing time, or the task with the shortest start time. The remaining processing time of a task can be estimated based on information from similar tasks that have historically processed that task.
[0096] In some embodiments, such as Figure 4As shown, conflict resolution strategies can include performance-guaranteed eviction policies. There are multiple Class B tasks currently being processed in cluster 120. Different tasks may occupy different processing unit performances. A cost-based eviction policy evicts the task occupying the processing unit with the best performance among these multiple tasks. For example, performance can be network performance. Specifically, network performance can be the connectivity performance of processing units. The connectivity performance of each processing unit with other processing units can be identified based on the network topology information between processing units in cluster 120. There are two connection methods between processing units in cluster 120: full mesh and partial mesh. Full mesh has higher connectivity performance than partial mesh. Full mesh means that any two processing units in the cluster are directly connected. Partial mesh means that some processing units in the cluster are not directly connected, and the information they exchange needs to be forwarded by other processing units in the cluster.
[0097] By employing a performance-guaranteed eviction strategy, Class B tasks that occupy high-performance processing units can be removed from these units, allowing them to be used for Class A tasks, thereby ensuring the SLA of Class A tasks.
[0098] In some embodiments, the expulsion device 112 may obtain a specific expulsion decision based on the aforementioned conflict resolution strategy, and then send the expulsion decision to the control plane 113. Based on the expulsion decision, the control plane 113 instructs the relevant processing units to perform the expulsion task.
[0099] Using the solution described above, users can uniformly distribute Class A and Class B tasks to cluster 120.
[0100] When the cluster 120 has sufficient resources to process all tasks issued by the user simultaneously, the control device 110 can allocate processing units to all tasks to handle them. In some embodiments, the control device 110 can allocate processing units to tasks in descending order of priority, thereby ensuring that high-priority tasks receive processing units first and that high-priority tasks are processed in a timely manner.
[0101] When cluster 120 is unable to process all tasks simultaneously and a task waiting queue appears, control device 110 instructs cluster 120 to prioritize processing type A tasks. Specifically, this can be divided into the following two situations.
[0102] Scenario 1: The queue for Class A tasks is not empty (i.e., there are Class A tasks waiting to be processed), while cluster 120 has a processing unit (e.g., processing unit 123) currently processing a Class B task. Regarding Scenario 1, if... Figure 5As shown, in step 51, the eviction device 112 in the control unit 110 instructs the processing unit 123 to evict the Class B task, that is, to remove the Class B task currently being processed by the processing unit 123 from the processing unit 123. And, in step 52, the control plane 123 switches the network to which the processing unit 123 is connected from network C2 to network C1. Then, in step 53, the resource manager 111 instructs the processing unit 123 to process the Class A task. Specifically, under the instruction of the resource manager 111, the processing unit 123 retrieves a Class A task from the Class A task waiting queue and processes the retrieved Class A task.
[0103] Scenario 2: The A-type task waiting queue is empty (i.e., there are no A tasks waiting to be processed), the B-type task waiting queue is not empty (i.e., there are B-type tasks waiting to be processed), and cluster 120 has an idle processing unit (e.g., processing unit 123). For request 2, such as... Figure 6 As shown, in step 61, resource manager 111 calculates the future idle time of processing unit 123. And, in step 62, control plane 113 switches the network to which processing unit 123 is connected from network C1 to network C2. Then, in step 63, resource manager 111 instructs processing unit 123 to process class B tasks. Specifically, under the instruction of resource manager 111, processing unit 123 retrieves class B tasks from the class B task waiting queue and processes the retrieved class B tasks. Finally, in step 64, when the future idle time of processing unit 123 ends, the resources of processing unit 123 are reclaimed.
[0104] Thus, the above solution allows low-priority tasks to be processed using the cluster's idle resources without affecting the processing of high-priority tasks, ensuring the SLA of high-priority tasks while improving the cluster's resource utilization.
[0105] Based on the above description, this application provides a task processing method. This method can be executed by the control device 110 in the computing system 100. For example... Figure 7 As shown, the method includes the following steps.
[0106] Step 701: Identify at least one idle processing unit in the processing unit cluster 120. An idle processing unit is a processing unit in the processing unit cluster 120 that is not processing type A tasks. Specifically, an idle processing unit is a processing unit in the processing unit cluster 120 that is not currently processing any type A tasks. In step 701, the control device 110 identifies the at least one idle processing unit when no type A tasks are waiting.
[0107] Step 702: Based on the historical information of the processing unit cluster 120 in processing the first task, predict the future idle time of the at least one idle processing unit. The future idle time is the time during which the at least one idle processing unit does not need to process the first task. The first task is a type A task. The future idle time is the time it takes for the at least one idle processing unit to process a single type A task (i.e., one instance of a type A task). The control device 110 can calculate the future idle time based on the resource allocation strategy described above. See the above description for details, which will not be repeated here.
[0108] Step 703: If the future idle time is greater than or equal to the required time of the second task, instruct the at least one idle processing unit to process the second task; wherein the priority of the second task is lower than the priority of the first task, and the required time of the second task is the time required for the at least one idle processing unit to process the second task. The second task belongs to class B tasks. For example, the second task is one or more instances of class B tasks.
[0109] In some embodiments, in step 703, the control device 110 may calculate the time required for the at least one idle processing unit to process the second task, i.e., the required time for the second task, based on information about the historical processing of type B tasks by the processing unit cluster 120.
[0110] In some embodiments, the method further includes: identifying the idle processing capability of the at least one idle processing unit. As described above, the processing capability of the idle processing unit here refers to the processing capability of the processing unit in an idle state, and is also the maximum available processing capability of the processing unit. The idle processing capability of the processing unit can be obtained by querying information such as the specifications of the processing unit.
[0111] In this embodiment, step 703 specifically involves: if the future idle duration is greater than or equal to the required duration of the second task, and the idle processing capacity is greater than or equal to the required processing capacity of the second task, instructing at least one idle processing unit to process the second task, wherein the required processing capacity of the second task is the capacity needed to process the second task. For example, the control device 110 can calculate the required processing capacity of the second task based on information about the historical processing of type B tasks by the processing unit cluster 120.
[0112] In one example of this embodiment, the second task is the task that requires the most processing power among a plurality of tasks, wherein the priority of each of the plurality of tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each of the plurality of tasks.
[0113] In this example, there are multiple Class B tasks currently waiting to be processed by the cluster of processing units 120. The required duration of each task is less than or equal to the future idle time of at least one processing unit, and the required processing capacity of each task is less than or equal to the processing capacity of at least one processing unit. The second task requires the largest processing capacity among these multiple tasks. This fully utilizes the computing resources of idle processing units, improving resource utilization. For more details, please refer to the above introduction to task filling strategies.
[0114] In some embodiments, the second task is the task with the longest processing time or the highest priority among a plurality of tasks; wherein the priority of each of the plurality of tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each of the plurality of tasks. For details, please refer to the above description of service assurance strategies and task filling strategies.
[0115] In some embodiments, the method further includes: upon obtaining a third task to be processed, and if the processing unit cluster does not have enough idle processing units to process the third task, removing the fourth task currently being processed by the processing unit cluster from the processing unit cluster, so that the processing unit cluster has enough idle processing units to process the third task; wherein the priority of the fourth task is lower than the priority of the third task. The third task belongs to class A tasks and may be one or more instances of class A tasks. The fourth task belongs to class B tasks and may be one or more instances of class B tasks.
[0116] In one example of this embodiment, the fourth task is the lowest priority task currently being processed by the processing unit cluster 120. For details, please refer to the above description of the service guarantee eviction policy.
[0117] In another example of this embodiment, the fourth task is the task with the longest remaining processing time among the tasks currently being processed by the processing unit cluster 120. For details, please refer to the above description of the cost-based eviction strategy.
[0118] In another example of this embodiment, the fourth task is the task with the fewest interruptions among the tasks currently being processed by the processing unit cluster 120. For details, please refer to the above description of the cost-based eviction policy.
[0119] In another example of this embodiment, the fourth task is the task currently being processed by the processing unit cluster 120 that occupies the best-performing processing unit. For details, please refer to the above description of the performance-guaranteed eviction policy.
[0120] In some embodiments, the processing units in the processing unit cluster 120 access different networks when processing different tasks. For details, please refer to the above description of the network isolation strategy.
[0121] In some embodiments, the first task is a model inference task, and the second task is a model training task.
[0122] In summary, the task processing method provided in this application utilizes idle processing units while processing high-priority tasks to process low-priority tasks. Furthermore, the time required for the processing unit to process low-priority tasks is less than or equal to the time when the processing unit does not need to process high-priority tasks. This improves resource utilization while avoiding the impact of low-priority task processing on the SLA of high-priority tasks.
[0123] Based on the above description, this application provides a control device 800. The computing system containing device 800 further includes a processing unit cluster, wherein, under the instruction of device 800, the processing units in the processing unit cluster can process different tasks at different times. For example... Figure 8 As shown, the device 800 includes:
[0124] The identification module 810 is used to identify at least one idle processing unit in the processing unit cluster;
[0125] The prediction module 820 is used to predict the future idle time of the at least one idle processing unit based on the historical information of the processing unit cluster processing the first task. The future idle time is the time during which the at least one idle processing unit does not need to process the first task.
[0126] The instruction module 830 is configured to instruct the at least one idle processing unit to process the second task when the future idle time is greater than or equal to the required time of the second task; wherein the priority of the second task is lower than the priority of the first task, and the required time of the second task is the time required for the at least one idle processing unit to process the second task.
[0127] In some embodiments, the identification module 810 is further configured to: identify the idle processing capability of the at least one idle processing unit; the indication module 830 is configured to: instruct the at least one idle processing unit to process the second task when the future idle duration is greater than or equal to the required duration of the second task, and the idle processing capability is greater than or equal to the required processing capability of the second task, wherein the required processing capability of the second task is the capability required to process the second task.
[0128] In one example of this embodiment, the second task is the task that requires the most processing power among a plurality of tasks, wherein the priority of each of the plurality of tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each of the plurality of tasks.
[0129] In some embodiments, the second task is the task with the longest processing time or the highest priority among a plurality of tasks; wherein the priority of each of the plurality of tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each of the plurality of tasks.
[0130] In some embodiments, the indicating module 830 is configured to: remove the fourth task currently being processed by the processing unit cluster from the processing unit cluster when a third task needs to be processed and the processing unit cluster does not have enough idle processing units to process the third task, so that the processing unit cluster has enough idle processing units to process the third task; wherein the priority of the fourth task is lower than the priority of the third task.
[0131] In one example of this embodiment, the fourth task is the lowest priority task among the tasks currently being processed by the processing unit cluster; or, the fourth task is the task with the longest remaining processing time among the tasks currently being processed by the processing unit cluster; or, the fourth task is the task with the fewest interruptions among the tasks currently being processed by the processing unit cluster; or, the fourth task is the task occupying the optimal performance processing unit among the tasks currently being processed by the processing unit cluster.
[0132] In some embodiments, the processing units in the processing unit cluster access different networks when processing different tasks.
[0133] In some embodiments, the first task is a model inference task, and the second task is a model training task.
[0134] The identification module 810, prediction module 820, and indication module 830 can all be implemented in software or in hardware. For example, the implementation of the identification module 810 will be described below. Similarly, the implementation of the prediction module 820 and indication module 830 can refer to the implementation of the identification module 810.
[0135] As an example of a software functional unit, the identification module 810 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the aforementioned computing instance may be one or more. For example, the identification module 810 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same Availability Zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0136] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0137] As an example of a hardware functional unit, the identification module 810 may include at least one computing device, such as a server. Alternatively, the identification module 810 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a general-purpose array logic (GAL), or any combination thereof.
[0138] The multiple computing devices included in the identification module 810 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the identification module 810 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the identification module 810 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0139] It should be noted that, in other embodiments, the identification module 810 can be used to perform... Figure 7 The prediction module 820 can be used to perform any step in the method shown. Figure 7 Any step in the method shown can be executed by the instruction module 830. Figure 7 Any step in the method shown. The steps implemented by the identification module 810, prediction module 820, and indication module 830 can be specified as needed, and implemented by the identification module 810, prediction module 820, and indication module 830 respectively. Figure 7 The different steps in the method shown enable the full functionality of device 800.
[0140] This application also provides a computing device 900. For example... Figure 9 As shown, the computing device 900 includes a bus 902, a processor 904, a memory 906, and a communication interface 908. The processor 904, the memory 906, and the communication interface 908 communicate with each other via the bus 902. The computing device 900 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 900.
[0141] The 902 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 9 The bus 902 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 902 may include a path for transmitting information between various components of the computing device 900 (e.g., memory 906, processor 904, communication interface 908).
[0142] Processor 904 may include a central processing unit (CPU) and a graphics processing unit (GPU).
[0143] Processing unit (GPU), microprocessor (MP), or digital signal processor (DSP) are any one or more of the following:
[0144] Memory 906 may include volatile memory, such as random access memory (RAM). Memory 906 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0145] The memory 906 stores executable program code, and the processor 904 executes the executable program code to implement the functions of the aforementioned identification module 810, prediction module 820, and indication module 830, thereby achieving... Figure 7 The method shown. That is, the memory 906 stores the method for execution. Figure 7 The instructions for the method shown.
[0146] The communication interface 908 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 900 and other devices or communication networks.
[0147] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0148] like Figure 10 As shown, the computing device cluster includes at least one computing device 900. The memory 906 in one or more computing devices 900 within the computing device cluster may store the same memory for executing... Figure 7 The instructions for the method shown.
[0149] In some possible implementations, the memory 906 of one or more computing devices 900 in the computing device cluster may also store data for execution. Figure 7 The instructions of the method shown are partial. In other words, a combination of one or more computing devices 900 can jointly execute instructions for performing... Figure 7 The instructions for the method shown.
[0150] It should be noted that the memory 906 in different computing devices 900 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the device 800. That is, the instructions stored in the memory 906 of different computing devices 900 can implement the functions of one or more modules among the identification module 810, prediction module 820, and indication module 830.
[0151] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 11 One possible implementation is shown. For example... Figure 11 As shown, two computing devices 900A and 900B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 906 in computing device 900A stores instructions for executing the functions of the identification module 810. Simultaneously, the memory 906 in computing device 900B stores instructions for executing the functions of the prediction module 820 and the indication module 830.
[0152] It should be understood that Figure 11 The functions of the computing device 900A shown can also be performed by multiple computing devices 900. Similarly, the functions of the computing device 900B can also be performed by multiple computing devices 900.
[0153] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 10 and Figure 11 The connection method of the computing device cluster. The difference is that the memory 906 in one or more computing devices 900 within this computing device cluster can store the same information for execution. Figure 7 The instructions for the method shown.
[0154] In some possible implementations, the memory 906 of one or more computing devices 900 in the computing device cluster may also store data for execution. Figure 7 The instructions of the method shown are partial. In other words, a combination of one or more computing devices 900 can jointly execute instructions for performing... Figure 7 The instructions for the method shown.
[0155] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform... Figure 7 The method shown.
[0156] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a host migration device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute... Figure 7 The method shown.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
Claims
1. A task processing method, characterized in that, The method is applied to a control device in a computing system, the computing system further comprising a cluster of processing units, wherein, under the instruction of the control device, the processing units in the cluster of processing units can process different tasks at different times; the method includes: Identify at least one idle processing unit in the processing unit cluster; Based on the historical information of the processing unit cluster in processing the first task, the future idle time of the at least one idle processing unit is predicted, wherein the future idle time is the time during which the at least one idle processing unit does not need to process the first task. If the future idle time is greater than or equal to the required time of the second task, the at least one idle processing unit is instructed to process the second task; wherein the priority of the second task is lower than the priority of the first task, and the required time of the second task is the time required for the at least one idle processing unit to process the second task.
2. The method according to claim 1, characterized in that, The method further includes: identifying the idle processing capability of the at least one idle processing unit; The step of instructing at least one idle processing unit to process the second task when the future idle time is greater than or equal to the required time of the second task includes: instructing at least one idle processing unit to process the second task when the future idle time is greater than or equal to the required time of the second task and the idle processing capacity is greater than or equal to the required processing capacity of the second task, wherein the required processing capacity of the second task is the capacity required to process the second task.
3. The method according to claim 2, characterized in that, The second task is the task that requires the most processing power among the multiple tasks, wherein the priority of each of the multiple tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each of the multiple tasks.
4. The method according to any one of claims 1-3, characterized in that, The second task is the one that takes the longest to process or has the highest priority among the multiple tasks; In this context, the priority of each of the plurality of tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each of the plurality of tasks.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: If a third task needs to be processed and the processing unit cluster does not have enough idle processing units to process the third task, the fourth task currently being processed by the processing unit cluster is removed from the processing unit cluster so that the processing unit cluster has enough idle processing units to process the third task; wherein the priority of the fourth task is lower than the priority of the third task.
6. The method according to claim 5, characterized in that, The fourth task is the lowest priority task among the tasks currently being processed by the processing unit cluster; or... The fourth task is the task with the longest remaining processing time among the tasks currently being processed by the processing unit cluster; or... The fourth task is the task with the fewest interruptions among the tasks currently being processed by the processing unit cluster; or... The fourth task is the task currently being processed by the processing unit cluster that occupies the best-performing processing unit.
7. The method according to any one of claims 1-6, characterized in that, The processing units in the processing unit cluster access different networks when processing different tasks.
8. The method according to any one of claims 1-7, characterized in that, The first task is the model inference task, and the second task is the model training task.
9. A control device, characterized in that, The computing system containing the device also includes a cluster of processing units, wherein, under the instruction of the device, the processing units in the cluster can process different tasks at different times; the device includes: The identification module is used to identify at least one idle processing unit in the processing unit cluster; The prediction module is used to predict the future idle time of the at least one idle processing unit based on the historical information of the processing unit cluster processing the first task. The future idle time is the time during which the at least one idle processing unit does not need to process the first task. An instruction module is configured to instruct at least one idle processing unit to process the second task when the future idle time is greater than or equal to the required time of the second task; wherein the priority of the second task is lower than the priority of the first task, and the required time of the second task is the time required for the at least one idle processing unit to process the second task.
10. The apparatus according to claim 9, characterized in that, The identification module is also used to: identify the idle processing capability of the at least one idle processing unit; The instruction module is used to instruct the at least one idle processing unit to process the second task when the future idle duration is greater than or equal to the required duration of the second task and the idle processing capacity is greater than or equal to the required processing capacity of the second task, wherein the required processing capacity of the second task is the capacity required to process the second task.
11. The apparatus according to claim 10, characterized in that, The second task is the task that requires the most processing power among the multiple tasks, wherein the priority of each of the multiple tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each of the multiple tasks.
12. The apparatus according to any one of claims 9-11, characterized in that, The second task is the one that takes the longest to process or has the highest priority among the multiple tasks; In this context, the priority of each of the plurality of tasks is lower than the priority of the first task, and the future idle time is greater than or equal to the required time of each of the plurality of tasks.
13. The apparatus according to any one of claims 9-12, characterized in that, The instruction module is used to: remove the fourth task currently being processed by the processing unit cluster from the processing unit cluster when a third task needs to be processed and the processing unit cluster does not have enough idle processing units to process the third task, so that the processing unit cluster has enough idle processing units to process the third task; wherein the priority of the fourth task is lower than the priority of the third task.
14. The apparatus according to claim 13, characterized in that, The fourth task is the lowest priority task among the tasks currently being processed by the processing unit cluster; or... The fourth task is the task with the longest remaining processing time among the tasks currently being processed by the processing unit cluster; or... The fourth task is the task with the fewest interruptions among the tasks currently being processed by the processing unit cluster; or... The fourth task is the task currently being processed by the processing unit cluster that occupies the best-performing processing unit.
15. The apparatus according to any one of claims 9-14, characterized in that, The processing units in the processing unit cluster access different networks when processing different tasks.
16. The apparatus according to any one of claims 9-15, characterized in that, The first task is the model inference task, and the second task is the model training task.
17. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-8.
18. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1-8.
19. A computer program product containing instructions, characterized in that, When the instruction is executed by a cluster of computer devices, the cluster of computer devices causes the cluster of computer devices to perform the method as described in any one of claims 1-8.