Task scheduling method and device, equipment and storage medium

By migrating tasks in the task scheduling cluster, using fragmented resources of the second node to perform the second task, and releasing the resources of the first node, the problem of waste of resources in the task scheduling cluster is solved, and resource integration and efficient utilization of computing resources are achieved.

CN120295745APending Publication Date: 2025-07-11SAIC MOTOR
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410039539.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the resource utilization rate of the task scheduling cluster is low, resulting in serious resource waste. Especially in deep learning model training tasks, the task scheduling cluster corresponding to the container is prone to the problem of resource fragmentation.

Method used

By obtaining the resource requirements of the first task, the second task executed on the first node is determined, and migrating it to the resource-rich second node to perform, while performing the first task on the first node, using the fragmented resources of the second node to achieve resource integration.

Benefits of technology

It improves the utilization rate of computing resources, reduces resource waste, and improves the processing efficiency of task queues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295745A_ABST
    Figure CN120295745A_ABST
Patent Text Reader

Abstract

The invention discloses a task scheduling method and device, equipment and a storage medium. The method is applied to a task scheduling cluster, and the task scheduling cluster comprises a first node and a second node. The method comprises the steps of obtaining a first task; the first resource quantity required by the first task is greater than the respective residual resource quantity of the first node and the second node; determining a second task executed by the first node according to the first resource quantity and the respective residual resource quantities of the first node and the second node; the second resource quantity required by the second task is less than the residual resource quantity of the second node, and the sum of the second resource quantity and the residual resource quantity of the first node is greater than the first resource quantity; and migrating the second task to the second node for execution, and executing the first task on the first node. In this way, the effect of fragment resource integration can be achieved, and therefore the utilization rate of computing resources is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a task scheduling method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of technology, the application of artificial intelligence technology is becoming more and more extensive. For example, autonomous driving technology, as one of the applications of artificial intelligence technology, plays an important role in the automotive industry. However, artificial intelligence technology usually requires a large amount of resources to support. Therefore, the problem of resource utilization has become one of the factors restricting the development of artificial intelligence technology.

[0003] In practical applications, the training tasks of deep learning models in the field of artificial intelligence tend to be platformized. For example, the training tasks of models are executed in containers using Nvidia-Docker (a container technology) to improve resource utilization using container technology. However, the scheduling methods of the task scheduling clusters corresponding to these containers are prone to cause serious waste of resources. Summary of the Invention

[0004] Embodiments of this application provide a task scheduling method, apparatus, device, and storage medium to execute tasks using fragmented resources and improve the utilization rate of computing resources.

[0005] In a first aspect, embodiments of this application provide a task scheduling method, which is applied to a task scheduling cluster. The task scheduling cluster includes a first node and a second node; the method includes:

[0006] Obtain a first task; the amount of the first resources required by the first task is greater than the remaining resources of each of the first node and the second node;

[0007] According to the amount of the first resources, the remaining resources of each of the first node and the second node, determine a second task executed by the first node; the amount of the second resources required by the second task is less than the remaining resources of the second node, and the sum of the amount of the second resources and the remaining resources of the first node is greater than the amount of the first resources;

[0008] Migrate the second task to the second node for execution, and execute the first task on the first node.

[0009] Optionally, before determining the second task executed by the first node according to the amount of the first resources, the remaining resources of each of the first node and the second node, and the task execution information of each of the first node and the second node, the method further includes:

[0010] Obtain the execution duration of each of the multiple tasks executed by the first node on the first node, and the total amount of resources of the first node and the second node respectively;

[0011] The determining the second task executed by the first node according to the first amount of resources and the remaining amounts of resources of the first node and the second node respectively includes:

[0012] According to the remaining amounts of resources of the first node and the second node respectively, and the total amounts of resources of the first node and the second node respectively, determine the fragmentation rate of the amount of resources jointly corresponding to the first node and the second node;

[0013] When the fragmentation rate of the amount of resources is greater than the preset fragmentation rate, determine the second task from the multiple tasks according to the execution duration of each of the multiple tasks executed by the first node on the first node; the execution duration of the second task on the first node is less than the preset duration.

[0014] Optionally, the remaining amounts of resources of the first node and the second node respectively include the remaining computing resource amount and the remaining storage resource amount; the total amounts of resources of the first node and the second node respectively include the total computing resource amount and the total storage resource amount;

[0015] The determining the fragmentation rate of the amount of resources jointly corresponding to the first node and the second node according to the remaining amounts of resources of the first node and the second node respectively, and the total amounts of resources of the first node and the second node respectively includes:

[0016] Based on the remaining computing resource amount and the total computing resource amount, determine the fragmentation rate of the computing resources jointly corresponding to the first node and the second node; and, based on the remaining storage resource amount and the total storage resource amount, determine the fragmentation rate of the storage resources jointly corresponding to the first node and the second node;

[0017] Based on the fragmentation rate of the computing resources and the fragmentation rate of the storage resources, determine the fragmentation rate of the amount of resources.

[0018] Optionally, the obtaining the first task includes:

[0019] Obtain the priority information of the multiple to-be-executed tasks included in the task queue;

[0020] Based on the priority information of the multiple to-be-executed tasks, obtain the first task from the task queue; the first task is the task with the highest priority among the multiple to-be-executed tasks.

[0021] Optionally, the method further includes:

[0022] Determine a third task executed by the first node according to the first resource amount and the remaining resource amounts of the first node and the second node respectively; the third resource amount required for the third task is greater than the remaining resource amount of the second node, or the sum of the third resource amount and the remaining resource amount of the first node is less than the first resource amount;

[0023] Determine a fourth task from the multiple tasks to be executed, and a node for executing the fourth task; the fourth resource amount required for the fourth task is less than the remaining resource amount of each of the first node or the second node; the node for executing the fourth task is the first node or the second node with the remaining resource amount greater than the fourth resource amount;

[0024] Adjust the fourth task to be the task with the highest priority among the multiple tasks to be executed, and execute the fourth task on the node for executing the fourth task.

[0025] Optionally, the migrating the second task to be executed on the second node includes:

[0026] Obtain the execution result generated during the execution of the second task, and stop executing the second task on the first node;

[0027] Continue to execute the second task on the second node based on the execution result; the execution result is sent from the first node to the second node.

[0028] Optionally, the method further includes:

[0029] When the execution status of the first task changes, update the remaining resource amounts of the first node and the second node respectively.

[0030] In a second aspect, an embodiment of the present application provides a task scheduling device, which is applied to a task scheduling cluster, and the task scheduling cluster includes a first node and a second node; the device includes:

[0031] A task acquisition module, configured to acquire a first task; the first resource amount required for the first task is greater than the remaining resource amounts of the first node and the second node respectively;

[0032] A task determination module, configured to determine a second task executed by the first node according to the first resource amount, the remaining resource amounts of the first node and the second node respectively, and the task execution information of the first node and the second node respectively; the second resource amount required for the second task is less than the remaining resource amount of the second node, and the sum of the second resource amount and the remaining resource amount of the first node is greater than the first resource amount;

[0033] A task execution module, configured to migrate the second task to the second node for execution and execute the first task on the first node.

[0034] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor, a memory, and a system bus;

[0035] The processor and the memory are connected through the system bus;

[0036] The memory is used to store a program, and the program includes instructions, which when executed by the processor, cause the processor to execute any implementation of the above task scheduling method.

[0037] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which instructions are stored, and when the instructions run on an electronic device, the electronic device is caused to execute any implementation of the above task scheduling method.

[0038] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:

[0039] In the embodiment of the present application, the first task can be obtained first. The first resource amount required by the first task is greater than the remaining resource amounts of the first node and the second node respectively. Then, according to the first resource amount, the remaining resource amounts of the first node and the second node respectively, the second task executed by the first node can be determined; the second resource amount required by the second task is less than the remaining resource amount of the second node, and the sum of the second resource amount and the remaining resource amount of the first node is greater than the first resource amount. Finally, the second task can be migrated to the second node for execution, and the first task can be executed on the first node.

[0040] In this way, although the first resource amount required by the first task is relatively large, and the remaining resource amounts of the first node and the second node respectively cannot support the execution of the first task, by migrating the second task executed by the first node to the second node, the fragmented resources on the second node can be effectively utilized, and the second task can be continued to be executed by the remaining resources of the second node. Moreover, after migrating the second task originally executed on the first node to the second node, the resources of the first node can be released, and the remaining resource amount of the first node after the resource release can be used to execute the first task. In this way, the effect of integrating fragmented resources can be achieved, thereby improving the utilization rate of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a flowchart of a task scheduling method provided by an embodiment of the present application;

[0042] Figure 2Flow chart of another task scheduling method provided by an embodiment of this application;

[0043] Figure 3 Structural schematic diagram of a task scheduling device provided by an embodiment of this application. Detailed implementation manners

[0044] As mentioned above, in practical applications, the training tasks of deep learning models in the field of artificial intelligence tend to be platform-based. For example, Nvidia-Docker (a container technology) is used to execute the model training tasks in containers to improve resource utilization using container technology. However, the scheduling methods of the task scheduling clusters corresponding to these containers are prone to cause serious waste of resources. This is because when the current task scheduling cluster performs task scheduling, when the remaining resource amount of a node is insufficient, the to-be-executed task will be directly assigned to other nodes, resulting in the remaining resource amount not being utilized, and thus there is a problem of resource fragmentation, leading to serious waste of resources.

[0045] To solve the above problems, an embodiment of this application provides a task scheduling method, which may include: first, obtain a first task, where the first resource amount required by the first task is greater than the remaining resource amounts of the first node and the second node respectively. Then, according to the first resource amount, the remaining resource amounts of the first node and the second node respectively, the second task executed by the first node can be determined; the second resource amount required by the second task is less than the remaining resource amount of the second node, and the sum of the second resource amount and the remaining resource amount of the first node is greater than the first resource amount. Finally, the second task can be migrated to the second node for execution, and the first task can be executed on the first node.

[0046] In this way, although the first resource amount required by the first task is relatively large and the remaining resource amounts of the first node and the second node respectively cannot support the execution of the first task, by migrating the second task executed by the first node to the second node, the fragmented resources on the second node can be effectively utilized, and the second task can continue to be executed using the remaining resources of the second node. Moreover, after migrating the second task originally executed on the first node to the second node, the resources of the first node can be released, and the remaining resource amount of the first node after the resource release can be used to execute the first task. In this way, the effect of integrating fragmented resources can be achieved, thereby improving the utilization rate of computing resources.

[0047] It should be noted that the execution subject of the task scheduling method in the embodiments of the present application is not limited. For example, the task scheduling method in the embodiments of the present application can be applied to data processing devices such as terminal devices or servers. Among them, the terminal device can be an electronic device such as a smart phone, a computer, a personal digital assistant (PDA), or a tablet computer. The server can be an independent server, a cluster server, or a cloud server.

[0048] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0049] Figure 1 It is a flowchart of a task scheduling method provided by an embodiment of the present application. In combination with Figure 1 As shown, the task scheduling method provided by the embodiment of the present application can be applied to a task scheduling cluster, and the task scheduling cluster includes a first node and a second node. Based on this, in the embodiment of the present application, the task scheduling method may include:

[0050] S101: Obtain a first task.

[0051] Among them, the first resource amount required for the first task is greater than the remaining resource amounts of the first node and the second node respectively. In this way, neither the first node nor the second node can support the execution of the first task, and fragmentation integration is required.

[0052] The first resource amount required for the above-mentioned first task can be configured in advance when the task is created. When configuring the first resource amount, a preset resource amount can be selected as the first resource amount, or the first resource amount can be customized. In addition, various information such as the identifier of the task, the mode of the task, the priority of the task, the data mounted by the task, the training code version of the task, the working image of the task, and the running time of the task can also be configured when the task is created. Among them, the mode of the task includes a single-machine training task and a distributed training task.

[0053] Based on this, in the embodiment of the present application, the process of obtaining the first task, that is, S101, may include: obtaining the priority information of multiple tasks to be executed included in the task queue; based on the priority information of the multiple tasks to be executed, obtaining the first task from the task queue; the first task is the task with the highest priority among the multiple tasks to be executed. Here, the higher the priority of the task to be executed, the more forward its position in the task queue, so it can be dequeued earlier and executed by the corresponding node; the lower the priority of the task to be executed, the more backward its position in the task queue, so it can be dequeued later and executed by the corresponding node. In this way, obtaining tasks according to the priority information of the tasks to be executed can achieve more efficient task scheduling and processing.

[0054] S102: Determine the second task executed by the first node according to the first resource amount and the remaining resource amounts of the first node and the second node respectively.

[0055] Among them, the second resource amount required for the second task is less than the remaining resource amount of the second node, and the sum of the second resource amount and the remaining resource amount of the first node is greater than the first resource amount. In this way, the second task executed on the first node can be migrated to the second node for execution, and the resources of the first node after the migration of the second task can also support the execution of the first task.

[0056] In addition, in the embodiment of the present application, it is also possible to first obtain the execution duration of each of the multiple tasks executed by the first node on the first node, and the total resource amounts of the first node and the second node respectively. Correspondingly, the process of determining the second task above, that is, step S102, may include: determining the fragmentation rate of the resource amount jointly corresponding to the first node and the second node according to the remaining resource amounts of the first node and the second node respectively and the total resource amounts of the first node and the second node; when the fragmentation rate of the resource amount is greater than the preset fragmentation rate, determining the second task from the multiple tasks according to the execution duration of each of the multiple tasks executed by the first node on the first node; the execution duration of the second task on the first node is less than the preset duration.

[0057] In this way, the second task needs to meet the following conditions: (1) The second resource amount required for the second task is less than the remaining resource amount of the second node; (2) The sum of the second resource amount of the second task and the remaining resource amount of the first node is greater than the first resource amount; (3) The fragmentation rate of the resource amount commonly corresponding to the first node and the second node exceeds the preset fragmentation rate; (4) The execution duration of the second task on the first node is less than the preset duration. For condition (3), when the fragmentation rate of the resource amount commonly corresponding to the first node and the second node exceeds the preset fragmentation rate, it can be indicated that there is a serious problem of resource fragmentation in the task scheduling cluster, which can be solved by means of task migration. For condition (4), when the execution duration of the second task on the first node is less than the preset duration, it can be indicated that the second task is not blocked and can still be executed. Based on this, by determining the second task, the second task can be migrated subsequently, so that the resources occupied by the second task on the first node are released, thus facilitating the subsequent execution of the first task.

[0058] Furthermore, the remaining resource amounts of the first node and the second node respectively include the remaining computing resource amount and the remaining storage resource amount; the total resource amounts of the first node and the second node respectively include the total computing resource amount and the total storage resource amount. Based on this, the process of determining the fragmentation rate of the resource amount commonly corresponding to the first node and the second node may include: based on the remaining computing resource amount and the total computing resource amount, determining the fragmentation rate of the computing resource commonly corresponding to the first node and the second node; and, based on the remaining storage resource amount and the total storage resource amount, determining the fragmentation rate of the storage resource commonly corresponding to the first node and the second node; based on the fragmentation rate of the computing resource and the fragmentation rate of the storage resource, determining the fragmentation rate of the resource amount.

[0059] In practical applications, the computing resources may include Graphics Processing Unit (GPU) resources and Central Processing Unit (CPU) resources; the storage resources may include memory resources and persistent storage resources. Based on this, for the sake of easy understanding, the following describes the process of obtaining the fragmentation rate of the resource amount in conjunction with formulas (1)-(5):

[0060]

[0061]

[0062]

[0063]

[0064] R avg = w GPU * S GPU + wCPU *S CPU +w Mem *S Men +w Storage *S Storage (5)

[0065] Among them, S GPU is the fragmentation rate of GPU resources, Node_CPU left_i is the remaining GPU resource amount of the i-th node in the task scheduling cluster, Node_CPU is the total GPU resource amount of any node in the task scheduling cluster, n represents the number of nodes in the task scheduling cluster, S CPU is the fragmentation rate of CPU resources, Node_CPU left_i is the remaining CPU resource amount of the i-th node, Node_CPU is the total CPU resource amount of any node, S Mem is the fragmentation rate of memory resources, Node_Mem left_i is the remaining memory resource amount of the i-th node, Node_Mem is the total memory resource amount of any node, S Storage is the fragmentation rate of persistent storage resources, Node_Storage left_i is the total remaining persistent storage resource amount of the i-th node, Node_Storage is the total persistent storage resource amount of any node, R avg is the fragmentation rate of the resource amount of the task scheduling cluster, w GPU is the weight of GPU resources, w CPU is the weight of CPU resources, w Mem is the weight of memory resources, w Storage is the weight of persistent storage resources.

[0066] It should be noted that in the embodiments of the present application, the task scheduling cluster may include multiple nodes, and n in the above formula is used to represent the number of multiple nodes. For the sake of understanding, the task scheduling process is only exemplarily described by taking the first node and the second node as examples.

[0067] S103: Migrate the second task to the second node for execution, and execute the first task on the first node.

[0068] In the embodiments of the present application, for the process of migrating the second task to the second node for execution, that is, S103, multiple possible implementation manners can be provided, which are described in detail below.

[0069] As a possible implementation manner, S103 may include: Stop executing the second task on the first node and re-execute the second task on the second node. In this way, the second task can be re-executed on the second node to quickly achieve the migration of the second task.

[0070] As another possible implementation, S103 may include: obtaining the execution result generated during the execution of the second task, and stopping the execution of the second task on the first node; continuing to execute the second task on the second node based on the execution result; the execution result is sent from the first node to the second node. In this way, the first node sends the execution result to the second node, and then the second task can be continuously executed on the second node, realizing the seamless switching and execution of the second task between nodes. And. Compared with re-executing the second task, this implementation can reduce the execution duration of the second task and improve the execution efficiency.

[0071] In addition, in the embodiments of the present application, the above task scheduling method may further include: when the execution state of the first task changes, updating the remaining resource amounts of the first node and the second node respectively. Here, the execution state of the first task may be to-be-executed, executing, or executed. When the first task is in the to-be-executed state, it does not occupy the resource amounts of the first node and the second node, and at this time, the remaining resource amounts of the first node and the second node do not change; when the first task is in the executing state, it occupies the resource amount of the node executing the first task, and at this time, the remaining resource amount of the node executing the first task, that is, the first node or the second node, changes; and when the first task is in the executed state, it releases the occupied resource amount during execution, causing the remaining resource amount of the node executing the first task, that is, the first node or the second node, to change. Therefore, updating the remaining resource amounts of the first node and the second node respectively through the change of the execution state of the first task helps to perform resource statistics in a timely and accurate manner, facilitates task migration based on the remaining resource amounts of each node, thereby realizing the integration of fragmented resources and improving the utilization rate of computing resources.

[0072] Based on the relevant content of the above steps S101 - S103, it can be known that in the embodiment of the present application, the first task can be obtained first. The first resource amount required for the first task is greater than the remaining resource amounts of the first node and the second node respectively. Then, according to the first resource amount, the remaining resource amounts of the first node and the second node respectively, the second task executed by the first node can be determined. The second resource amount required for the second task is less than the remaining resource amount of the second node, and the sum of the second resource amount and the remaining resource amount of the first node is greater than the first resource amount. Finally, the second task can be migrated to the second node for execution, and the first task can be executed on the first node. In this way, although the first resource amount required for the first task is relatively large, and the remaining resource amounts of the first node and the second node respectively cannot support the execution of the first task, by migrating the second task executed by the first node to the second node, the fragmented resources on the second node can be effectively utilized, and the second task can be continued to be executed with the remaining resources of the second node. Moreover, after migrating the second task originally executed on the first node to the second node, the resources of the first node can be released, and the remaining resource amount of the first node after resource release can be used to execute the first task. In this way, the effect of integrating fragmented resources can be achieved, thereby improving the utilization rate of computing resources.

[0073] Further, as mentioned above, the task queue of the task scheduling cluster can arrange each task to be executed according to the priority information of the task to be executed to determine the execution order of each task to be executed. Based on this, when executing the first task with the highest priority, if the tasks in the tasks to be executed cannot perform fragmented resource integration, that is, cannot be migrated, the execution of the first task can be postponed, and the tasks with less required resource amount can be executed first to avoid blocking of multiple tasks to be executed in the task queue, thereby improving the processing efficiency of the task queue. To facilitate the understanding of the above process, the following can be described in combination with embodiments and drawings.

[0074] Figure 2 It is a flowchart of another task scheduling method provided by the embodiment of the present application. Combining Figure 2 As shown, the task scheduling method provided by the embodiment of the present application can be applied to a task scheduling cluster, and the task scheduling cluster includes a first node and a second node. Based on this, in the embodiment of the present application, the task scheduling method can include:

[0075] S201: Obtain the first task.

[0076] In the embodiment of the present application, the implementation process of S201 can refer to the relevant description of S101 in the above embodiment, and will not be elaborated here.

[0077] S202: Determine the third task executed by the first node according to the first resource amount, the remaining resource amounts of the first node and the second node respectively.

[0078] Among them, the amount of the third resource required for the third task is greater than the remaining resources of the second node, or the sum of the amount of the third resource and the remaining resources of the first node is less than the amount of the first resource. Thus, the remaining resources of the second node cannot meet the resource requirement of the third task. Therefore, the third task executed on the first node cannot be migrated to the second node for execution. Or, the resources of the first node after migrating the third task cannot meet the resource requirement of the first task. That is to say, the third task executed on the first node cannot be migrated to the second node, and the resource requirements of the first task and the third task cannot be met.

[0079] S203: Determine a fourth task from multiple tasks to be executed, and a node for executing the fourth task.

[0080] Among them, the amount of the fourth resource required for the fourth task is less than the remaining resources of the first node or the second node respectively. Here, since the third task cannot be migrated, the first task cannot be executed. Therefore, the fourth task with a smaller resource requirement can be executed first to avoid blocking of multiple tasks to be executed. Of course, the node for executing the above-mentioned fourth task can be the first node or the second node with the remaining resources greater than the amount of the fourth resource. That is, if the remaining resources of the first node are greater than the amount of the fourth resource, the fourth task can be directly executed on the first node subsequently; if the remaining resources of the second node are greater than the amount of the fourth resource, the fourth task can be directly executed on the second node subsequently; if the remaining resources of the first node and the second node are both greater than the amount of the fourth resource, the fourth task can be selected from the first node and the second node for execution subsequently.

[0081] S204: Adjust the fourth task to the task with the highest priority among multiple tasks to be executed, and execute the fourth task on the node for executing the fourth task.

[0082] Since the fourth task has a smaller resource requirement, it can be adjusted to the task with the highest priority for priority execution, so as to avoid blocking of multiple tasks to be executed and further improve the processing efficiency of the task queue.

[0083] Furthermore, based on the task scheduling method provided in the above embodiments, an embodiment of the present application can also provide a task scheduling device. The task scheduling device will be described below in combination with the embodiments and the drawings respectively.

[0084] Figure 3 It is a schematic structural diagram of a task scheduling device provided by an embodiment of the present application. In combination with Figure 3 As shown, the task scheduling device 300 provided by an embodiment of the present application is applied to a task scheduling cluster, and the task scheduling cluster includes a first node and a second node; the task scheduling device 300 includes:

[0085] A task acquisition module 301, configured to acquire a first task; a first resource amount required by the first task is greater than remaining resource amounts of the first node and the second node respectively;

[0086] A task determination module 302, configured to determine a second task executed by the first node according to the first resource amount, remaining resource amounts of the first node and the second node respectively, and task execution information of the first node and the second node respectively; a second resource amount required by the second task is less than the remaining resource amount of the second node, and a sum of the second resource amount and the remaining resource amount of the first node is greater than the first resource amount;

[0087] A task execution module 303, configured to migrate the second task to the second node for execution, and execute the first task on the first node.

[0088] As an implementation manner, the task scheduling device 300 further includes:

[0089] An information acquisition module, configured to acquire execution durations of multiple tasks executed by the first node on the first node respectively, and total resource amounts of the first node and the second node respectively;

[0090] The task determination module 302 includes:

[0091] A fragmentation rate determination module, configured to determine a fragmentation rate of a resource amount commonly corresponding to the first node and the second node according to remaining resource amounts of the first node and the second node respectively, and total resource amounts of the first node and the second node respectively;

[0092] A task determination sub-module, configured to, when the fragmentation rate of the resource amount is greater than a preset fragmentation rate, determine the second task from the multiple tasks according to execution durations of the multiple tasks executed by the first node on the first node respectively; an execution duration of the second task on the first node is less than a preset duration.

[0093] As an implementation manner, remaining resource amounts of the first node and the second node respectively include remaining computing resource amounts and remaining storage resource amounts; total resource amounts of the first node and the second node respectively include total computing resource amounts and total storage resource amounts;

[0094] The fragmentation rate determination module includes:

[0095] The first fragmentation rate determination sub-module is used to determine the fragmentation rate of the computing resources jointly corresponding to the first node and the second node based on the remaining computing resource amount and the total computing resource amount; and determine the fragmentation rate of the storage resources jointly corresponding to the first node and the second node based on the remaining storage resource amount and the total storage resource amount;

[0096] The second fragmentation rate determination sub-module is used to determine the fragmentation rate of the resource amount based on the fragmentation rate of the computing resources and the fragmentation rate of the storage resources.

[0097] As an implementation manner, the task acquisition module 301 includes:

[0098] The priority acquisition module is used to acquire the priority information of multiple to-be-executed tasks included in the task queue;

[0099] The task acquisition sub-module is used to acquire the first task from the task queue based on the priority information of the multiple to-be-executed tasks; the first task is the task with the highest priority among the multiple to-be-executed tasks.

[0100] As an implementation manner, the task determination module 302 is further used for:

[0101] Determine the third task executed by the first node according to the first resource amount and the remaining resource amounts of the first node and the second node respectively; the third resource amount required by the third task is greater than the remaining resource amount of the second node, or the sum of the third resource amount and the remaining resource amount of the first node is less than the first resource amount;

[0102] Determine the fourth task from the multiple to-be-executed tasks and the node for executing the fourth task; the fourth resource amount required by the fourth task is less than the remaining resource amounts of the first node or the second node respectively; the node for executing the fourth task is the first node or the second node with the remaining resource amount greater than the fourth resource amount;

[0103] The task execution module 303 is further used for:

[0104] Adjust the fourth task to the task with the highest priority among the multiple to-be-executed tasks and execute the fourth task on the node for executing the fourth task.

[0105] As an implementation manner, the task execution module 303 includes:

[0106] The result acquisition module is used to acquire the execution result generated during the execution of the second task and stop executing the second task on the first node;

[0107] A task execution sub-module, configured to continue to execute the second task on the second node based on the execution result; the execution result is sent from the first node to the second node.

[0108] As an implementation manner, the task scheduling device 300 further includes:

[0109] An information update module, configured to update the remaining resource amounts of the first node and the second node respectively when the execution status of the first task changes.

[0110] Furthermore, an embodiment of the present application further provides an electronic device, including: a processor, a memory, and a system bus;

[0111] The processor and the memory are connected through the system bus;

[0112] The memory is used to store a program, the program includes instructions, and when the instructions are executed by the processor, the processor is caused to execute any implementation manner of the above task scheduling method.

[0113] Furthermore, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored, and when the instructions run on an electronic device, the terminal device is caused to execute any implementation manner of the above task scheduling method.

[0114] From the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present application. It should be noted that the embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other.

[0115] For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0116] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0117] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A task scheduling method, characterized in that, Applied to a task scheduling cluster, the task scheduling cluster includes a first node and a second node; the method includes: Obtain a first task; the first resource amount required by the first task is greater than the remaining resource amounts of the first node and the second node respectively; Determine a second task executed by the first node according to the first resource amount, the remaining resource amounts of the first node and the second node respectively; the second resource amount required by the second task is less than the remaining resource amount of the second node, and the sum of the second resource amount and the remaining resource amount of the first node is greater than the first resource amount; Migrate the second task to the second node for execution, and execute the first task on the first node.

2. The task scheduling method according to claim 1, wherein Before determining the second task executed by the first node according to the first resource amount, the remaining resource amounts of the first node and the second node respectively, and the task execution information of the first node and the second node respectively, the method further includes: Obtain the execution durations of multiple tasks executed by the first node on the first node respectively, and the total resource amounts of the first node and the second node respectively; The determining the second task executed by the first node according to the first resource amount, the remaining resource amounts of the first node and the second node respectively, includes: Determine the fragmentation rate of the resource amount commonly corresponding to the first node and the second node according to the remaining resource amounts of the first node and the second node respectively, and the total resource amounts of the first node and the second node respectively; When the fragmentation rate of the resource amount is greater than the preset fragmentation rate, determine the second task from the multiple tasks according to the execution durations of the multiple tasks executed by the first node on the first node respectively; the execution duration of the second task on the first node is less than the preset duration.

3. The task scheduling method according to claim 2, wherein The remaining resource amounts of the first node and the second node respectively include the remaining computing resource amount and the remaining storage resource amount; the total resource amounts of the first node and the second node respectively include the total computing resource amount and the total storage resource amount; The determining the fragmentation rate of the resource amount commonly corresponding to the first node and the second node according to the remaining resource amounts of the first node and the second node respectively, and the total resource amounts of the first node and the second node respectively, includes: Based on the remaining computing resource amount and the total computing resource amount, determine the fragmentation rate of the computing resource commonly corresponding to the first node and the second node; and, based on the remaining storage resource amount and the total storage resource amount, determine the fragmentation rate of the storage resource commonly corresponding to the first node and the second node; Based on the fragmentation rate of the computing resource and the fragmentation rate of the storage resource, determine the fragmentation rate of the resource amount.

4. The task scheduling method according to claim 1, wherein The obtaining the first task includes: Obtain the priority information of multiple to-be-executed tasks included in the task queue; Based on the priority information of the multiple to-be-executed tasks, obtain the first task from the task queue; the first task is the task with the highest priority among the multiple to-be-executed tasks.

5. The task scheduling method according to claim 4, wherein, The method further includes: Determine a third task executed by the first node according to the first resource amount and the remaining resource amounts of the first node and the second node respectively; the third resource amount required for the third task is greater than the remaining resource amount of the second node, or the sum of the third resource amount and the remaining resource amount of the first node is less than the first resource amount; Determine a fourth task from the multiple tasks to be executed, and a node for executing the fourth task; the fourth resource amount required for the fourth task is less than the remaining resource amounts of the first node or the second node respectively; the node for executing the fourth task is the first node or the second node with the remaining resource amount greater than the fourth resource amount; Adjust the fourth task to be the task with the highest priority among the multiple tasks to be executed, and execute the fourth task on the node for executing the fourth task.

6. The task scheduling method according to claim 1, characterized in that The migrating the second task to be executed on the second node includes: Obtain the execution result generated during the execution of the second task, and stop executing the second task on the first node; Continue to execute the second task on the second node based on the execution result; the execution result is sent from the first node to the second node.

7. The task scheduling method according to any one of claims 1 to 6, characterized in that, The method further includes: When the execution status of the first task changes, update the remaining resource amounts of the first node and the second node respectively.

8. A task scheduling device, characterized in that, Applied to a task scheduling cluster, the task scheduling cluster includes a first node and a second node; the apparatus includes: A task acquisition module, configured to acquire a first task; the first resource amount required for the first task is greater than the remaining resource amounts of the first node and the second node respectively; A task determination module, configured to determine a second task executed by the first node according to the first resource amount, the remaining resource amounts of the first node and the second node respectively, and the task execution information of the first node and the second node respectively; the second resource amount required for the second task is less than the remaining resource amount of the second node, and the sum of the second resource amount and the remaining resource amount of the first node is greater than the first resource amount; A task execution module, configured to migrate the second task to be executed on the second node, and execute the first task on the first node.

9. An electronic device, characterized in that, The device includes: a processor, a memory, and a system bus; The processor and the memory are connected through the system bus; The memory is used to store a program, the program includes instructions, and when the instructions are executed by the processor, the processor is caused to execute the task scheduling method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions run on an electronic device, the electronic device is caused to execute the task scheduling method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Resource defragmentation method, controller, control node and computer cluster

    CN120950224A

  • Resource fragmentation consolidation method, controller, control node, and computer cluster

    CN120950224B