Task scheduling method and related equipment
By actively receiving task information and deciding when sending request information based on its own load situation, requesting task execution to the second computing node, and sending instructions information regularly during the execution process, the problem of low task scheduling efficiency in the prior art and inability to perceive task execution between computing nodes is solved, and efficient and uniform task allocation and fault migration are achieved.
Patent Information
- Application Number
- CN202510103779.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-01-21
AI Technical Summary
When the existing task scheduling methods have a large number of tasks or a large number of computing nodes, the scheduling efficiency is low, and the task execution cannot be effectively sensed between computing nodes, resulting in uneven task allocation and affecting execution efficiency.
The first computing node actively receives the task information and decides the time when the request information is sent based on its own load situation, and requests the task to be executed from the second computing node. Meanwhile, when performing a task, the first computing node regularly sends instructions to the second computing node so that the second computing node can monitor the task status and avoid duplicate assignment of tasks.
It improves the efficiency of task scheduling, avoids the problem of uneven task allocation, and realizes task failure migration when a computing node fails, ensuring the continuity and efficiency of task execution.
Smart Images

Figure CN120144243A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud computing, and more particularly, to a task scheduling method, a computing device, a computing device cluster, a computer program product, and a computer-readable storage medium. Background Art
[0002] Task scheduling methods are used to split a larger task into multiple smaller tasks, and each computing node among multiple computing nodes executes some of the multiple smaller tasks, thereby improving the execution efficiency of the task. The current task scheduling methods mainly include the following two. The first task scheduling method is that the task scheduling node obtains the resource utilization of each computing node among multiple computing nodes, so as to allocate appropriate tasks to each computing node. The second task scheduling method is that each computing node actively obtains and executes tasks from the task queue. However, in the first task scheduling method, the execution efficiency of the task highly depends on the task scheduling node. When the number of tasks or computing nodes is large, the task scheduling node may not be able to schedule tasks in time, resulting in a low execution efficiency of the task. In the second task scheduling method, multiple computing nodes cannot perceive the execution status of tasks from each other, which easily leads to uneven task allocation, resulting in a low execution efficiency of the task. Moreover, when a computing node fails, the tasks in that computing node cannot be switched to other computing nodes for execution, thus affecting the execution efficiency of the task.
[0003] Therefore, how to schedule tasks and improve the execution efficiency of tasks has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a task scheduling method, a computing device, a computing device cluster, a computer program product, and a computer-readable storage medium, which can schedule tasks, thereby improving the execution efficiency of tasks.
[0005] In a first aspect, a method for scheduling tasks is provided. This method is executed by a first computing node, which belongs to the infrastructure managed by a cloud management platform. The infrastructure is used to provide cloud services and includes multiple computing nodes. The method includes: receiving information about a first task from a second computing node, where the information about the first task includes the information required to execute the first task. The moment of receiving the information about the first task is the first moment, and the first task belongs to a first set, and the first set includes at least one uncompleted task; at a second moment, sending a first request message to the second computing node, where the first request message is used to request the execution of a task in the first set. There is at least a first time length between the second moment and the first moment, and the first time length is determined according to a first load value of the first computing node. The first load value is determined according to the number of tasks currently executed by the first computing node and / or the resource utilization rate of the first computing node; receiving a first response message from the second computing node, where the first response message is used to instruct the first computing node to execute a second task. The second task belongs to the tasks in the first set that have not been executed, or the first response message is used to indicate that each task in the first set is being executed; when the first response message is used to instruct the first computing node to execute the second task, execute the second task according to the first response message.
[0006] In the embodiments of the present application, the first computing node can obtain the information of each task to be executed, and thus request the execution of the task by sending a request message to the second computing node. Therefore, it is not necessary for the second computing node to actively allocate tasks to the first computing node for execution, thereby avoiding the problem that tasks cannot be timely allocated to the first computing node for execution due to over-reliance on the scheduling of the second computing node. Moreover, the first computing node needs to determine the first time length according to its own load situation, so that after the first time length after receiving the information of the task, it can send a request message to the second computing node, so that the computing node with a lower load can give priority to sending a request message to execute the task, thereby improving the execution efficiency of the task.
[0007] In combination with the first aspect, in some implementation manners, the first load value is positively correlated with the number of tasks currently executed by the first computing node and / or the resource utilization rate of the first computing node, and the first time length is positively correlated with the first load value.
[0008] In the embodiments of the present application, when the number of tasks currently executed by the first computing node is relatively large and / or the resource utilization rate is relatively high, the time for sending the request message is relatively late; when the number of tasks currently executed by the first computing node is relatively small and / or the resource utilization rate is relatively low, the time for sending the request message is relatively early. Therefore, it is possible to make the computing node with a relatively small number of currently executed tasks and / or a relatively low resource utilization rate give priority to sending a request message to give priority to applying for task execution.
[0009] In combination with the first aspect, in some implementation manners, when the first response message is used to indicate that each task in the first set is being executed, at a third moment, a first request message is sent to a second computing node. There is at least an interval of N second time lengths between the third moment and the second moment. The second time length is a preset time length, and N is a positive integer.
[0010] In the embodiments of the present application, when the first response message is used to indicate that each task in the first set is being executed, the first computing node may repeatedly send a request message at regular intervals to request task execution, so as to take over the execution of the task by this computing node when the computing node originally used to execute the task can no longer execute the task, thereby realizing fault migration.
[0011] In combination with the first aspect, in some implementation manners, when the first response message is used to indicate that the first computing node executes a second task, at a fourth moment, a first indication message is sent to the second computing node. The first indication message is used to indicate that the first computing node is executing the second task. There is an interval of M third time lengths between the fourth moment and the fifth moment. The fifth moment is the moment when the first response message is received. The third time length is a preset time length, and M is a positive integer.
[0012] In the embodiments of the present application, during the execution of the second task by the first computing node, the first computing node may send the first indication message to the second computing node at regular intervals, so that the second computing node determines that the second task is being executed, thereby avoiding the problem that the second computing node repeatedly instructs other computing nodes to execute the second task, or avoiding the situation that the second computing node cannot perceive when the first computing node is unable to execute the second task.
[0013] In combination with the first aspect, in some implementation manners, when the first response message is used to indicate that the first computing node executes the second task, after the execution of the second task is completed, a second indication message is sent to the second computing node. The second indication message is used to indicate that the second task has been executed.
[0014] In the embodiments of the present application, after the execution of the second task is completed, the first computing node sends the second indication message to the second computing node, so that the second computing node determines that the second task has been executed, thereby facilitating the second computing node to delete the second task from the first set.
[0015] In combination with the first aspect, in some implementation manners, a second task is determined from the first set according to the first response message; and the second task is executed according to the information of the second task.
[0016] In combination with the first aspect, in some implementations, the second task is the same as or different from the first task. When the second task is different from the first task, the method further includes: receiving information about the second task from a second computing node, where the information about the second task includes the information required to execute the second task.
[0017] In the embodiments of the present application, the first computing node can not only receive the information of the tasks executed by the first computing node, but also receive the information of the tasks executed by other computing nodes except the first computing node, that is, the first computing node can obtain the information of each task in the first set. After sending the first request information, the first computing node can directly determine the information of the second task from the information of each task in the stored first set according to the indication of the second computing node, and thus execute the second task according to the information of the second task.
[0018] In a second aspect, a method for scheduling tasks is provided. The method is executed by a second computing node, and the second computing node belongs to the infrastructure managed by the cloud management platform. The infrastructure is used to provide cloud services and includes multiple computing nodes. The method includes: sending the information of each task in the first set to each computing node in the computing node set, where the first set includes at least one uncompleted task, and the information of each task includes the information required to execute each task. The computing node set includes multiple computing nodes, and each computing node in the computing node set is used to execute tasks; receiving a first request information from a first computing node, where the first request information is used to request to execute a task in the first set, and the first computing node belongs to the computing node set; in response to the first request information, sending a first response information to the first computing node; where, when there is at least one unexecuted task in the first set, the first response information is used to instruct the first computing node to execute a second task, and the second task belongs to at least one unexecuted task; or, when each task in the first set is being executed, the first response information is used to indicate that each task in the first set is being executed.
[0019] In the embodiments of the present application, each computing node in the computing node set can obtain the information of each task in the first set, so as to request to execute a task by sending request information to the second computing node. Therefore, it is not necessary for the second computing node to actively assign tasks to the computing nodes in the computing node set for execution, thereby avoiding the problem that tasks cannot be timely assigned to the computing nodes for execution due to over-reliance on the scheduling of the second computing node.
[0020] In combination with the second aspect, in some implementations, each task in the first set corresponds to a first identification information, and the first identification information of each task is used to indicate whether each task is being executed.
[0021] In the embodiments of the present application, the second computing node determines whether one or more unexecuted tasks are included in the first set through the first identification information of each task, so as to facilitate indicating the one or more unexecuted tasks to the computing node for execution.
[0022] In combination with the second aspect, in some implementation manners, when the first response information is used to instruct the first computing node to execute the second task, set the first identification information of the second task to be used to indicate that the second task is being executed within the first time period. The starting moment of the first time period includes the moment of sending the first response information, or a moment before or after the moment of sending the first response information. The length of the first time period is the second time length, and the second time length is a preset time length.
[0023] In the embodiments of the present application, after the second computing node instructs the first computing node to execute the second task, set the first identification information of the second task to be used to indicate that the second task is being executed, and the first identification information of the second task cannot be modified within the first time period, so as to prevent the second computing node from repeatedly instructing the second task to other computing nodes for execution.
[0024] In combination with the second aspect, in some implementation manners, before the end of the first time period, receive the first indication information from the first computing node. The first indication information is used to indicate that the first computing node is executing the second task; according to the first indication information, update the first time period. The starting moment of the updated first time period includes the moment of receiving the first indication information, or a moment before or after the moment of receiving the first indication information. The length of the updated first time period is the second time length; set the first identification information of the second task to be used to indicate that the second task is being executed within the updated first time period.
[0025] In the embodiments of the present application, during the process of the first computing node executing the second task, the first computing node may send the first indication information to the second computing node before the end of the first time period or before the end of the updated first time period, so that the second computing node determines that the second task is being executed, thereby avoiding the problem that the second computing node repeatedly instructs the second task to other computing nodes for execution, or avoiding the problem that the second computing node cannot perceive when the first computing node cannot execute the second task.
[0026] In combination with the second aspect, in some implementation manners, when the first indication information from the first computing node is not received before the end of the first time period, at the end of the first time period, set the first identification information of the second task to be used to indicate that the second task is not executed, and the first indication information is used to indicate that the first computing node is executing the second task.
[0027] In the embodiment of the present application, when the second computing node does not receive the indication information from the first computing node before the end of the first time period or before the end of the updated first time period, it is determined that the first computing node is unable to execute the second task. Therefore, the first identification information of the second task is set to indicate that the second task has not been executed, so as to facilitate instructing the second task to be executed by computing nodes other than the first computing node, thereby realizing failover.
[0028] In combination with the second aspect, in some implementation manners, the second indication information from the first computing node is received, and the second indication information is used to indicate that the second task has been executed and completed; the second task in the first set is deleted; the third indication information is sent to other computing nodes in the computing node set except the first computing node, and the third indication information is used to indicate that the second task has been executed and completed, or the third indication information is used to indicate deleting the second task in the first set.
[0029] In the embodiment of the present application, after the first computing node executes the second task, the second computing node receives the second indication information to determine that the second task has been executed and completed, and thus deletes the second task from the first set. After the second computing node deletes the second task, it can also notify other computing nodes except the first computing node to delete the second task and / or the information of the second task.
[0030] In a third aspect, a computing device is provided. The device includes a module for implementing the first aspect or any possible implementation manner of the first aspect.
[0031] In a fourth aspect, a computing device is provided. The device includes a module for implementing the second aspect or any possible implementation manner of the second aspect.
[0032] In a fifth aspect, a task scheduling system is provided. The system includes the computing device described in the third aspect and the device described in the fourth aspect.
[0033] In a sixth aspect, a computing device cluster is provided, including at least one computing device, and each computing device includes a processor and a memory; the processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the method described in any one of the first aspect or the second aspect or any possible implementation manner in any one of the aspects.
[0034] In a seventh aspect, a computer program product including instructions is provided. When the instructions are run by a computing device cluster, the computing device cluster is caused to execute the method described in any one of the first aspect or the second aspect or any possible implementation manner in any one of the aspects.
[0035] In an eighth aspect, there is provided a computer-readable storage medium including computer program instructions which, when executed by a cluster of computing devices, cause the cluster of computing devices to execute the method described in any one of the above first aspect or second aspect or any implementation manner in any one of the aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 FIG. is a schematic structural diagram of a cloud scenario according to an embodiment of the present application.
[0037] Figure 2 FIG. is a schematic structural diagram of a task scheduling system according to an embodiment of the present application.
[0038] Figure 3 FIG. is a schematic flowchart of a task scheduling method according to an embodiment of the present application.
[0039] Figure 4 FIG. is a schematic flowchart of a task scheduling method according to another embodiment of the present application.
[0040] Figure 5 FIG. is a schematic diagram of a fourth moment and a fifth moment according to an embodiment of the present application.
[0041] Figure 6 FIG. is a schematic structural block diagram of a computing device according to an embodiment of the present application.
[0042] Figure 7 FIG. is a schematic structural diagram of a computing device according to an embodiment of the present application.
[0043] Figure 8 FIG. is a schematic structural diagram of a cluster of computing devices according to an embodiment of the present application.
[0044] Figure 9 FIG. is a schematic diagram of a connection between computing devices 700A and 700B via a network according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0046] Embodiments of the present application will present various aspects, embodiments or features around a system including multiple devices, components, modules, etc. It should be understood and clear that each system may include additional devices, components, modules, etc., and / or may not include all the devices, components, modules, etc. discussed in conjunction with the accompanying drawings. In addition, combinations of these solutions may also be used.
[0047] In addition, in the embodiments of the present application, words such as "exemplary" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "exemplary" is intended to present concepts in a specific manner.
[0048] The business scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art can know that with the evolution of technology and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0049] References to "one embodiment" or "some embodiments" etc. described in this specification mean that specific features, structures or characteristics described in connection with that embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0050] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: including the case where A exists alone, the case where A and B exist simultaneously, and the case where B exists alone, where A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (item)" or its similar expression below refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0051] The methods in the embodiments of the present application can be applied to various computing nodes. The computing nodes belong to the infrastructure managed by the cloud management platform, and the infrastructure is used to provide cloud services. That is, the cloud management platform is used to manage the infrastructure that provides cloud services. The infrastructure includes at least one cloud data center. Each cloud data center includes at least one computing node. Each computing node is, for example, a container, a virtual machine, a server, a computing device, etc. Each computing node respectively includes cloud service resources, so as to provide corresponding cloud services for tenants.
[0052] The cloud management platform can be located in the cloud data center and can provide access interfaces (such as interfaces or application program interfaces (APIs)). Tenants can remotely access the access interfaces to register cloud accounts and passwords on the cloud management platform and log in to the cloud management platform. After the cloud management platform successfully authenticates the cloud accounts and passwords, tenants can further pay on the cloud management platform to select and purchase computing nodes (such as containers, virtual machines, servers, computing devices, etc.) of specific specifications (such as processors, memory, disks). After the successful payment and purchase, the cloud management platform provides the remote login account password of the purchased computing node, and tenants can remotely log in to the computing node and install and run the tenants' applications on the computing node. Therefore, tenants can create, manage, log in to, and operate computing nodes in the cloud data center through the cloud management platform. Among them, the computing node can also be called an elastic compute service (ECS), an elastic instance (with different names in different cloud service providers).
[0053] It should be understood that the tenants of cloud services can be individuals, enterprises, schools, hospitals, administrative organs, etc.
[0054] The functions of the cloud management platform include but are not limited to user consoles, computing management services, network management services, storage management services, authentication services, and image management services. The user console provides interfaces or APIs to interact with tenants. The computing management service is used to manage servers running virtual machines and containers and bare metal servers. The network management service is used to manage network services (such as gateways, firewalls, etc.). The storage management service is used to manage storage services (such as data bucket services). The authentication service is used to manage tenants' account passwords. The image management service is used to manage virtual machine images.
[0055] The embodiments of the present application do not limit the meaning of a moment. For example, a moment includes milliseconds, seconds, minutes, hours, days, etc. A moment in time period A includes, for example: one millisecond in time period A, one second in time period A, one minute in time period A, one hour in time period A, one day in time period A, etc. The starting moment in time period A includes, for example: the first millisecond in time period A, the first second in time period A, the first minute in time period A, the first hour in time period A, the first day in time period A, etc. The ending moment in time period A includes, for example: the last millisecond in time period A, the last second in time period A, the last minute in time period A, the last hour in time period A, the last day in time period A, etc. The moment of receiving information includes: any moment within the time period for receiving information (such as the first moment or the last moment). The moment of sending information includes: any moment within the time period for sending information (such as the first moment or the last moment). The embodiments of the present application do not limit the meaning of a time period. For example, a time period includes at least one moment starting from the starting moment of the time period. The embodiments of the present application do not limit the meaning of a time length. For example, a time length includes at least one moment.
[0056] Figure 1 It is a schematic structural diagram of cloud scenario 100 provided by the embodiments of the present application. Figure 1 The cloud scenario 100 in [description] includes a cloud management platform 110. The cloud management platform 110 is used to manage the infrastructure that provides cloud services. The infrastructure includes at least one data center (such as data center 120). The data center 120 includes at least one computing node cluster, and each computing node cluster includes multiple computing nodes. The computing nodes are, for example, containers, virtual machines, servers, computing devices, etc. Tenants can apply to use the resources in the data center 120 through the cloud management platform 110. The tenant is a tenant of a public cloud that has registered a public cloud account and purchased public cloud resources. Each computing node managed by the cloud management platform 110 is used to execute the method provided in the embodiments of the present application, such as Figure 3 or Figure 4 the method in [description].
[0057] For example, the data center 120 includes a computing node cluster 130 and / or a computing node cluster 140. The computing node cluster 130 includes computing nodes 131 and 132, and the computing node cluster 140 includes computing nodes 141 and 142. Among them, the network bandwidth and / or data transmission speed between computing nodes belonging to the same computing node cluster is greater than the network bandwidth and / or data transmission speed between computing nodes belonging to different computing node clusters. For example, the network bandwidth and / or data transmission speed between computing nodes 131 and 132 is greater than the network bandwidth and / or data transmission speed between computing nodes 131 and 141.
[0058] In some embodiments, multiple computing nodes in a computing node cluster are directly connected or connected through a network, such as a wide area network or a local area network.
[0059] In some embodiments, the data center 120 may further include at least one storage node cluster, and each storage node cluster includes at least one storage node. The embodiments of the present application do not limit the type of the storage node. For example, the storage node is a centralized storage node or a distributed storage node. The storage node is used to store the data of the tenant, or the storage node is used to store the data and / or the generated data required when executing the method in the embodiments of the present application. Exemplarily, multiple storage nodes in the storage node cluster are directly connected or connected through a network, such as a wide area network or a local area network.
[0060] In some embodiments, multiple computing nodes in the data center 120 are used to execute tasks, such as tasks generated during the operation of the tenant's business. When the multiple computing nodes execute tasks, the tasks need to be scheduled to improve the execution efficiency of the tasks. For example, the task scheduling system is as Figure 2 shown.
[0061] Figure 2 is a schematic structural block diagram of the task scheduling system provided by the embodiments of the present application. Figure 2 The task scheduling system 200 in includes computing nodes 210, 220, and 230. Among them, at least two of the computing nodes 210, 220, and 230 belong to the same computing node cluster or different computing node clusters.
[0062] For example, when computing nodes 210, 220, and 230 are containers, at least two of the computing nodes 210, 220, and 230 are deployed in the same virtual machine or physical server, or at least two of the computing nodes 210, 220, and 230 are deployed in different virtual machines or physical servers. Alternatively, when computing nodes 210, 220, and 230 are virtual machines, at least two of the computing nodes 210, 220, and 230 are deployed in the same physical server, or at least two of the computing nodes 210, 220, and 230 are deployed in different physical servers.
[0063] Computing node 210 is used to receive information of at least one task to be executed, and the at least one task to be executed belongs to the first set. Information of each task to be executed is the information required to execute the task. For example, information of each task to be executed includes at least one of the following: the content of the task, the computing method required to execute the task, the data to be accessed for executing the task, etc. After obtaining information of each task to be executed, computing node 210 sends the information of the task to be executed to each computing node (such as computing nodes 220 and 230) in the set of computing nodes connected to the computing node 210. The set of computing nodes includes multiple computing nodes, and each computing node is used to execute tasks. After each computing node in the set of computing nodes receives information of at least one task to be executed from computing node 210, it sends a first request message to computing node 210, and the first request message is used to request to execute a task in the first set.
[0064] In some embodiments, at a first moment, computing node 220 receives information of a first task from computing node 210. The first task belongs to the first set, and the information of the first task includes the information required to execute the first task. At a second moment, computing node 220 sends a first request message to computing node 210. There is at least a first time length between the second moment and the first moment, and the first time length is determined according to the first load value of computing node 220. The first load value is determined according to the number of tasks currently executed by computing node 220 and / or the resource utilization rate of computing node 220. The way to determine the moment when computing node 230 sends a first request message to computing node 210 is similar to that of the second moment, and will not be elaborated here. In other words, the computing nodes for executing tasks determine the moment to send request messages according to their own load conditions, so that the computing nodes with lower load can execute tasks first, thereby improving the execution efficiency of tasks.
[0065] Exemplarily, the first load value is positively correlated with the number of tasks currently executed by the computing node 220 and / or the resource utilization rate of the computing node 220, and the first time length is positively correlated with the first load value.
[0066] After receiving the first request information from the computing node 220, the computing node 210 sends a first response information to the computing node 220 in response to the first request information. The first response information is used to instruct the computing node 220 to execute a second task, or the first response information is used to indicate that each task in the first set is being executed. In the case where the first response information is used to instruct the computing node 220 to execute the second task, the computing node 220 executes the second task according to the first response information. Similarly, after receiving the request information from the computing node 230, the computing node 210 sends a response information to the computing node 230 in response to the request information, and the response information is used to instruct the computing node 230 to execute a task, or is used to indicate that each task in the first set is being executed.
[0067] Exemplarily, in the case where the first response information is used to instruct the computing node 220 to execute the second task, the first response information includes the identification information of the second task.
[0068] In some embodiments, each task in the first set has a first identification information, and the first identification information of each task is used to indicate whether the task is being executed.
[0069] Exemplarily, when the first identification information includes a first string, the first identification information is used to indicate that the task is being executed. When the first identification information includes a second string, the first identification information is used to indicate that the task is not being executed. The first string and the second string are different. The first string and the second string include at least one character, and the at least one character includes: numbers, letters, words, symbols, etc.
[0070] In some embodiments, after receiving the first request information, the computing node 210 determines whether the first set includes tasks that have not been executed. Specifically, the computing node determines whether the first identification information of each task in the first set is used to indicate that the task is being executed. When the first set includes one or more tasks that have not been executed, the computing node 210 sends the first response information to the computing node 220, and the first response information is used to instruct the computing node 220 to execute the second task. The second task belongs to the one or more tasks that have not been executed. In other words, when the computing node 210 determines that the second task has not been executed, it sends the first response information to the computing node 220 to instruct the computing node 220 to execute the second task. When each task in the first set is being executed, the computing node 210 sends the first response information to the computing node 220, and the first response information is used to indicate that each task in the first set is being executed. In other words, when the computing node 210 determines that there are no tasks that have not been executed currently, it sends the first response information to the computing node 220 to instruct the computing node 220 that there is no need to execute new tasks. Alternatively, when the computing node 210 determines that there are no tasks that have not been executed currently, it does not need to send the first response information to the computing node 220.
[0071] Exemplarily, "the first response information is used to indicate that each task in the first set is being executed" and "the first response information is used to indicate that there are no tasks that need to be executed currently" can be replaced with each other.
[0072] In some embodiments, when the second task has not been executed, before or after sending the first response information, the computing node 210 sets the first identification information of the second task to be used to indicate that the second task is being executed within the first time period. The starting moment of the first time period includes the moment of sending the first response information, or the starting moment of the first time period is the moment before the moment of sending the first response information, or the starting moment of the first time period is the moment after the moment of sending the first response information. The time length of the first time period is the second time length, and the second time length is a preset time length. The present application embodiment does not limit the specific value of the second time length. For example, it is 30 seconds, 1 minute, 5 minutes, etc.
[0073] In some embodiments, when the first response information is used to instruct the computing node 220 to execute the second task, after receiving the first response information, the computing node 220 determines the second task from the first set and executes the second task according to the information of the second task.
[0074] In some embodiments, the first task and the second task are the same or different. When the first task and the second task are different, the computing node 220 may also receive the information of the second task from the computing node 210, and the information of the second task includes the information required to execute the second task.
[0075] In some embodiments, when the computing node 220 is executing the second task, at the fourth moment, the computing node 220 sends first indication information to the computing node 210. The first indication information is used to indicate that the computing node 220 is executing the second task. There is an interval of M third time lengths between the fourth moment and the fifth moment, where the fifth moment is the moment of receiving the first response information, the third time length is a preset time length, and M is a positive integer. The specific value of the third time length in the embodiments of the present application is not limited. For example, it can be 30 seconds, 1 minute, 5 minutes, etc. In other words, after receiving the first response information, the computing node 220 sends the first indication information to the computing node 210 at regular intervals to indicate that the computing node 220 is executing the second task.
[0076] Exemplarily, at the fourth moment, the computing node 220 is executing the second task.
[0077] Exemplarily, the third time length is less than the second time length. For example, the third time length is half of the second time length.
[0078] In some embodiments, before the end of the first time period, the computing node 210 receives the first indication information from the computing node 220. In response to the first indication information, the computing node 210 updates the first time period. The start moment of the updated first time period includes the moment of receiving the first indication information, or the start moment of the updated first time period includes a moment before the moment of receiving the first indication information, or the start moment of the updated first time period includes a moment after the moment of receiving the first indication information. The time length of the updated first time period is the second time length. Before the end of the first time period, if the computing node 210 does not receive the first indication information from the computing node 220, at the end of the first time period, the computing node 210 sets the first identification information of the second task to indicate that the second task has not been executed. In other words, the computing node 210 can set a time lock for the second task, so that the first identification information of the second task indicates that the second task is being executed within the first time period. And, the computing node 210 can receive a relock indication (i.e., the first indication information) from the computing node 220 before the end of the first time period, so as to extend the time lock of the second task. If the computing node 210 does not receive a relock request before the end of the first time period, it means that the computing node 220 cannot continue to execute the second task. After the end of the first time period, the computing node 210 cancels the time lock of the second task so that other computing nodes (such as the computing node 230) can request to execute the task.
[0079] In some embodiments, after updating the first time period, the computing node 210 sets the first identification information of the second task to indicate that the second task is being executed within the updated first time period. Before the end of the updated first time period, the computing node 210 receives the first indication information from the computing node 220. In response to the first indication information, the computing node 210 updates the first time period again. Before the end of the updated first time period, if the computing node 210 does not receive the first indication information from the computing node 220, at the end of the updated first time period, the computing node 210 sets the first identification information of the second task to indicate that the second task is not being executed.
[0080] In some embodiments, after the computing node 220 finishes executing the second task, it sends the second indication information to the computing node 210, and the second indication information is used to indicate that the second task has been executed. After receiving the second indication information, the computing node 210 deletes the second task from the first set. The computing node 210 sends the third indication information to other computing nodes (such as the computing node 230) except the computing node 220. The third indication information is used to indicate that the second task has been executed, or the third indication information is used to indicate deleting the second task in the first set.
[0081] Exemplarily, the first set is stored in each of the computing node 210, the computing node 220, and the computing node 230. Information of each task in the first set is also stored in the computing node 210, the computing node 220, and the computing node 230.
[0082] In some embodiments, when the first response information is used to indicate that each task in the first set is being executed, at the third moment, the computing node 220 sends the first request information to the computing node 210. The third moment is at least N second time lengths apart from the second moment, and N is a positive integer. In other words, when there is no task to be executed in the first set currently, the computing node 220 can repeat sending the first request information at regular intervals to request task execution.
[0083] In Figure 2In the system 200, each computing node (e.g., computing node 220, computing node 230) used to execute a task can obtain information about each task to be executed, and thus request the execution of the task by sending request information to the computing node 210. Therefore, there is no need for the computing node 210 to actively allocate tasks to the computing nodes used to execute tasks, thereby avoiding the problem that tasks cannot be timely allocated to the computing nodes used to execute tasks due to over-reliance on the scheduling of the computing node 210. Moreover, the computing node used to execute a task needs to send request information to the computing node 210 after a first time period after receiving the information of the task. Therefore, when the number of tasks currently executed by the computing node used to execute tasks is relatively large and / or the resource utilization rate is relatively high, the time to send the request information is relatively late, so that the computing node with a relatively small number of currently executed tasks and / or a relatively low resource utilization rate can give priority to sending request information to execute the task, thereby improving the execution efficiency of the task.
[0084] Figure 3 is a schematic flowchart of a task scheduling method provided by an embodiment of the present application. Figure 3 The method in can be executed by a first computing node and a second computing node, and the first computing node and the second computing node belong to the infrastructure managed by a cloud management platform, and the infrastructure is used to provide cloud services. The first computing node and the second computing node are, for example, Figure 1 two different computing nodes in. Or, the first computing node is, for example, Figure 2 the computing node 220 or the computing node 230 in, and the second computing node is, for example, Figure 2 the computing node 210 in. Figure 3 The method in includes the following steps.
[0085] 310, send information about the first task to the first computing node.
[0086] After obtaining the information of the first task, the second computing node sends the information of the first task to each computing node in the computing node set. Correspondingly, each computing node in the computing node set receives the information of the first task. The computing node set includes multiple computing nodes, and each computing node in the multiple computing nodes is used to execute tasks. The first task is a task that has not been executed to completion, that is, the first task is a task to be executed. The information of the first task includes the information required to execute the first task.
[0087] Exemplarily, the information of the first task includes at least one of the following: the content of the first task, the calculation method required to execute the first task, the data to be accessed required to execute the first task, etc.
[0088] Exemplarily, the second computing node sends information about each task in the first set to each computing node in the set of computing nodes, where the first set includes at least one task that has not been executed to completion. The first task belongs to the first set.
[0089] Exemplarily, the first set is stored in the second computing node. After the second computing node obtains a task that has not been executed, it adds the task to the first set.
[0090] 320. At the second moment, send the first request message to the second computing node.
[0091] After receiving the information about the first task, the first computing node determines the second moment and sends the first request message to the second computing node at the second moment. Correspondingly, the second computing node receives the first request message. The first request message is used to request the execution of a task in the first set. The time interval between the second moment and the first moment is at least the first time length, and the first moment is the moment when the first computing node receives the information about the first task. The first time length is determined according to the first load value of the first computing node, and the first load value is determined according to the number of tasks currently executed by the first computing node and / or the resource utilization rate of the first computing node.
[0092] Exemplarily, the first request message is used to request the execution of a specified task, and the first request message includes the identification information of the specified task. Alternatively, the first request message is not used to request the execution of a specified task, and the first request message does not include the identification information of the task.
[0093] In some embodiments, the first computing node determines the first load value according to the number of tasks currently executed and / or the resource utilization rate. The first computing node determines the first time length according to the first load value. Among them, the first load value is positively correlated with the number of tasks currently executed by the first computing node and / or the resource utilization rate, and the first time length is positively correlated with the first load value. In other words, the first time length corresponding to a computing node with a relatively large number of currently executed tasks and / or a relatively high resource utilization rate is relatively long, so the moment of sending the first request message is relatively late. The first time length corresponding to a computing node with a relatively small number of currently executed tasks and / or a relatively low resource utilization rate is relatively short, so the moment of sending the first request message is relatively early.
[0094] Exemplarily, the calculation formula for the first time length of the first computing node is: L = a1x1 + a2x2. Where L represents the first time length, x1 represents the number of tasks currently being executed by the first computing node, and x2 represents the current resource utilization rate of the first computing node. a1 and a2 are preset values, and a1 and a2 may be the same or different. The embodiments of the present application do not limit the specific values of a1 and a2.
[0095] Exemplarily, the resource utilization rate of the first computing node is determined according to at least one of the following: the resource utilization rate of computing resources (such as the resource utilization rate of a processing unit), the resource utilization rate of storage resources (such as the resource utilization rate of memory), the resource utilization rate of data transmission resources (such as the resource utilization rate of input-output (IO) resources), etc.
[0096] For example, the calculation formula for the resource utilization rate of the first computing node is:
[0097]
[0098] Wherein, represents the total amount of computing resources of the first computing node, represents the amount of computing resources that have been used currently in the first computing node. represents the total amount of storage resources of the first computing node, represents the amount of storage resources that have been used currently in the first computing node. represents the total amount of data transmission resources of the first computing node, represents the amount of data transmission resources that have been used currently in the first computing node. b1, b2, and b3 are preset values, and at least two of b1, b2, and b3 are the same or different. The embodiments of the present application do not limit the specific values of b1, b2, and b3.
[0099] Optionally, when the first load value is greater than or equal to the first preset threshold, the first computing node does not send the first request message to the second computing node, that is, the first computing node does not execute step 320. When the first load value is less than the first preset threshold, the first computing node sends the first request message to the second computing node, that is, the first computing node executes step 320. In other words, when the load of the first computing node itself is relatively high, the first computing node does not request to execute tasks. The embodiments of the present application do not limit the specific value of the first preset threshold.
[0100] 330, send a first response message to the first computing node.
[0101] After receiving the first request message, the second computing node sends a first response message to the first computing node in response to the first request message. Correspondingly, the first computing node receives the first response message. This first response message is used to instruct the first computing node to execute a second task, or this first response message is used to indicate that each task in the first set is being executed. This second task belongs to the tasks in the first set that have not been executed. This second task is the same as or different from the first task.
[0102] Exemplarily, in the case where the first response message is used to instruct the first computing node to execute a second task, the first response message includes identification information of the second task.
[0103] Exemplarily, "the first response message is used to indicate that each task in the first set is being executed" and "the first response message is used to indicate that there are no tasks to be executed currently" can be replaced with each other.
[0104] In some embodiments, after the second computing node receives the first request message, it determines whether there is one or more unexecuted tasks in the first set. When there is one or more unexecuted tasks in the first set, the first response message sent by the second computing node is used to instruct the first computing node to execute a second task. The second task belongs to the one or more unexecuted tasks. When there are no unexecuted tasks in the first set, the first response message sent by the second computing node is used to indicate that each task in the first set is being executed.
[0105] Exemplarily, each task in the first set corresponds to a first identification information. The first identification information of each task is used to indicate whether the task is being executed. The second computing node determines whether each task is being executed through the first identification information of each task in the first set.
[0106] Exemplarily, when the first identification information includes a first string, the first identification information is used to indicate that the task is being executed. When the first identification information includes a second string, the first identification information is used to indicate that the task is not executed. The first string and the second string are different. The first string and the second string include at least one character, and the at least one character includes: numbers, letters, words, symbols, etc.
[0107] In some embodiments, when the second computing node determines that there are no unexecuted tasks in the first set, there is no need to send the first response message to the first computing node. That is, step 330 may not be executed.
[0108] In some embodiments, in the case where the first response message is used to instruct the first computing node to execute a second task, the first computing node may further execute step 340. That is, step 340 is an optional step and may not be executed.
[0109] 340. Execute the second task according to the first response message.
[0110] After receiving the first response message, the first computing node executes the second task according to the first response message. Specifically, the first computing node determines the second task from the first set according to the first response message. The first computing node executes the second task according to the information of the second task.
[0111] In some embodiments, before step 340, the first computing node receives information about a second task from the second computing node. The information about the second task includes the information required to execute the second task. The information about the second task is similar to the information about the first task and will not be elaborated here.
[0112] In some embodiments, a first set is stored in the first computing node. After the first computing node receives information about a task to be executed from the second computing node, it adds the task to be executed to the first set. Exemplarily, information about each task in the first set is also stored in the first computing node.
[0113] In some embodiments, when the first response message is used to indicate that each task in the first set is being executed, the first computing node sends a first request message at a third time. The third time is at least N second time lengths apart from the first time. The second time length is a preset time length, and N is a positive integer. In other words, when the first response message is used to indicate that each task in the first set is being executed, the first computing node can repeatedly send request messages at regular intervals to request task execution, so that when the computing node originally used to execute the task can no longer execute the task, it takes over the computing node to execute the task, thereby achieving fault migration.
[0114] Exemplarily, at the third time, if the first load value of the first computing node is greater than or equal to the first preset threshold, the first computing node does not send the first request message. At the third time, if the first load value of the first computing node is less than the first preset threshold, the first computing node sends the first request message.
[0115] In Figure 3 's method, the first computing node can obtain information about each task to be executed, and thus request to execute the task by sending a request message to the second computing node. Therefore, it is not necessary for the second computing node to actively allocate tasks to the first computing node for execution, thereby avoiding the problem that tasks cannot be allocated to the first computing node in a timely manner due to over-reliance on the scheduling of the second computing node. Moreover, the first computing node needs to wait for the first time length after receiving the task information before sending a request message to the second computing node. Therefore, when the number of tasks currently executed by the first computing node is large and / or the resource utilization rate is high, the time to send the request message is relatively late, so that a computing node with a relatively small number of currently executed tasks and / or a relatively low resource utilization rate can send a request message first to execute the task, thereby improving the task execution efficiency.
[0116] When one or more unexecuted tasks are included in the first set, the first computing node and the second computing node can also execute Figure 4 's method. That is, after step 320,Figure 4 The method in
[0117] Figure 4 is a schematic flowchart of a task scheduling method provided by an embodiment of the present application. Figure 4 The method in can be executed by a first computing node and a second computing node, which belong to the infrastructure managed by a cloud management platform, and the infrastructure is used to provide cloud services. The first computing node and the second computing node are, for example, Figure 1 two different computing nodes in Figure 2 The computing node 220 or the computing node 230 in, and the second computing node is, for example, Figure 2 the computing node 210 in Figure 4 The method in includes the following steps.
[0118] 410, set the first identification information of the second task to indicate that the second task is being executed within the first time period.
[0119] Before or after executing step 330, the second computing node sets the first identification information of the second task to indicate that the second task is being executed within the first time period. The start time of the first time period includes the time of executing step 330, or a time before or after the time of executing step 330. The time length of the first time period is the second time length. The second time length is a preset time length. The specific value of the second time length in the embodiment of the present application is not limited, and is, for example, 30 seconds, 1 minute, 5 minutes, etc.
[0120] In some embodiments, within the first time period, the first identification information of the second task cannot be modified.
[0121] Optionally, at the fourth moment, if the first computing node is executing the second task, the first computing node executes step 420. At the fourth moment, if the first computing node cannot execute the second task, the first computing node does not execute step 420. That is, step 420 is an optional step and can be not executed. The fourth moment is as described below.
[0122] 420, at the fourth moment, send the first indication information to the second computing node.
[0123] At the fourth moment, if the first computing node is executing the second task, the first computing node sends first indication information to the second computing node. Correspondingly, the second computing node receives the first indication information from the first computing node. The first indication information is used to indicate that the first computing node is executing the second task. There is an interval of M third time lengths between the fourth moment and the fifth moment. The fifth moment is the moment when the first computing node receives the first response information. The third time length is a preset time length, and M is a positive integer. The specific value of the third time length in the embodiments of the present application is not limited, for example, it can be 30 seconds, 1 minute, 5 minutes, etc. The first response information is used to indicate that the first computing node executes the second task. In other words, during the process of the first computing node executing the second task, the first computing node sends the first indication information to the second computing node at regular intervals to indicate that it is executing the second task, so as to avoid the problem that when the first computing node cannot execute the second task, the second computing node cannot perceive it, resulting in the second task not being able to be executed by other computing nodes.
[0124] In some embodiments, the third time length is less than the second time length. For example, the third time length is half of the second time length.
[0125] For example, the fifth moment and the fourth moment are as Figure 5 shown. Figure 5 It is a schematic diagram of the fourth moment and the fifth moment provided by the embodiments of the present application. As Figure 5 shown, Figure 5 it includes four moments, namely moment t 1 , moment t 2 , moment t 3 and moment t 4 . Among them, there is an interval of one third time length between t 1 and t 2 . There is an interval of two third time lengths between t 3 and t 1 . There is an interval of one third time length between t 3 and t 2 . There is an interval of three third time lengths between t 1 and t 4 . There is an interval of one third time length between t 4 and t 3 . When t 1 is the fifth moment, at least one of t 2 , t 3 , t 4 is the fourth moment.
[0126] 430, update the first time period according to the first indication information.
[0127] After receiving the first indication information, the second computing node updates the first time period according to the first indication information. The start time of the updated first time period includes the time when the first indication information is received, or includes a time before or after the time when the first indication information is received. The time length of the updated first time period is the second time length.
[0128] In some embodiments, after the second computing node updates the first time period, the first identification information of the second task is set to indicate that the second task is being executed within the updated first time period.
[0129] Exemplarily, within the updated first time period, the first identification information of the second task cannot be modified.
[0130] In some embodiments, the first computing node and the second computing node may repeatedly execute steps 420 and 430 until the second task is completed.
[0131] Optionally, before the end of the first time period or the updated first time period, if the second computing node does not receive the first indication information from the first computing node, the second computing node executes step 440. That is, step 440 is an optional step and may not be executed.
[0132] 440, set the first identification information of the second task to indicate that the second task is not executed.
[0133] Before the end of the first time period, if the second computing node does not receive the first indication information from the first computing node, at the end of the first time period, the second computing node sets the first identification information of the second task to indicate that the second task is not executed. Or, before the end of the updated first time period, if the second computing node does not receive the first indication information from the first computing node, at the end of the updated first time period, the second computing node sets the first identification information of the second task to indicate that the second task is not executed.
[0134] In some embodiments, after the second computing node sets the first identification information of the second task to indicate that the second task is not executed, the first set includes one or more unexecuted tasks. The second computing node receives the first request information from the third computing node and sends the first response information to the third computing node, and the first response information is used to instruct the third computing node to execute the second task. The third computing node is the same as or different from the first computing node.
[0135] Optionally, after the first computing node executes the second task, the first computing node executes step 450. That is, step 450 is an optional step and may not be executed.
[0136] 450, send the second indication information to the second computing node.
[0137] After the second task is completed, the first computing node sends second indication information to the second computing node. Correspondingly, the second computing node receives the second indication information from the first computing node. The second indication information is used to indicate that the second task has been completed.
[0138] In some embodiments, after the second task is completed, the first computing node deletes the second task in the first set in the first computing node. The first computing node may also delete the information of the second task.
[0139] 460, delete the second task in the first set.
[0140] After receiving the second indication information, the second computing node deletes the second task in the first set in the second computing node. The second computing node may also delete the information of the second task.
[0141] In some embodiments, the second computing node sends third indication information to the fourth computing node. The third indication information is used to indicate that the second task has been completed, or the third indication information is used to indicate that the fourth computing node deletes the second task in the first set in the fourth computing node. The fourth computing node is different from the first computing node. After receiving the third indication information, the fourth computing node deletes the second task in the first set in the fourth computing node. The first set is stored in the fourth computing node. The fourth computing node may also delete the information of the second task.
[0142] In Figure 4 's method, during the execution of the second task, the first computing node sends first indication information to the second computing node so that the second computing node determines that the second task is being executed. The second computing node may set the first identification information of the second task within a time period to indicate that the second task is being executed, thereby avoiding repeatedly instructing the second task to be executed by computing nodes other than the first computing node. The second computing node may also set the first identification information of the second task to indicate that the second task is not being executed if it does not receive the indication information from the first computing node before the end of the time period, thereby facilitating instructing the second task to be executed by computing nodes other than the first computing node, and thus realizing failover.
[0143] Figure 6 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. Figure 6 The computing device 600 in Figure 6 includes a transceiver module 610 and a processing module 620. Figure 3 or Figure 4 The method executed by the first computing node or the second computing node in
[0144] When the computing device 600 is used to execute the method executed by the first computing node in Figure 3 the transceiver module 610 is used to: receive information about the first task; at a second moment, send first request information to the second computing node; receive first response information. The transceiver module 610 is used to execute step 320 in Figure 3
[0145] In some embodiments, when the computing device 600 is used to execute the method executed by the first computing node in Figure 3 the processing module 620 is used to execute a second task according to the first response information. The processing module 620 is used to execute step 340 in Figure 3
[0146] When the computing device 600 is used to execute the method executed by the second computing node in Figure 3 the transceiver module 610 is used to: send information about the first task to the first computing node; receive first request information; send first response information to the first computing node. The transceiver module 610 is used to execute steps 310 and 330 in Figure 3
[0147] When the computing device 600 is used to execute the method executed by the first computing node in Figure 4 the transceiver module 610 is used to: at a fourth moment, send first indication information to the second computing node; send second indication information to the second computing node. The transceiver module 610 is used to execute steps 420 and 450 in Figure 4
[0148] When the computing device 600 is used to execute the method executed by the second computing node in Figure 4 the processing module 620 is used to set the first identification information of the second task to indicate that the second task is being executed within a first time period. The processing module 620 is used to execute step 410 in Figure 4
[0149] In some embodiments, when the computing device 600 is used to execute the method executed by the second computing node in Figure 4 the transceiver module 610 is used to: receive first indication information and second indication information. The processing module is further used to: update the first time period according to the first indication information; set the first identification information of the second task to indicate that the second task is not being executed; delete the second task in the first set. The processing module 620 is used to execute steps 430, 440, and 460 in Figure 4
[0150] Among them, both the transceiver module 610 and the processing module 620 can be implemented by software or by hardware. Exemplarily, next, taking the transceiver module 610 as an example, the implementation manner of the transceiver module 610 will be introduced. Similarly, the implementation manner of the processing module 620 can refer to the implementation manner of the transceiver module 610.
[0151] As an example of a software functional unit, the transceiver module 610 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the transceiver module 610 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Among them, generally, one region may include multiple AZs.
[0152] Similarly, the multiple hosts / virtual machines / containers for running this code may be distributed in the same VPC or in multiple VPCs. Among them, generally, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.
[0153] As an example of a hardware functional unit, the transceiver module 610 may include at least one computing device, such as a server, etc. Alternatively, the transceiver module 610 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0154] The multiple computing devices included in the transceiver module 610 may be distributed in the same region or in different regions. The multiple computing devices included in the transceiver module 610 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the transceiver module 610 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), and generic array logic (GALs).
[0155] Therefore, the modules of the examples described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0156] It should be noted that when the device provided in the above embodiment executes the above method, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. For example, the transceiver module 610 can be used to execute any step in the above method, and the processing module 620 can be used to execute any step in the above method. The steps to be implemented by the transceiver module 610 and the processing module 620 can be specified according to needs, and all the functions of the above device can be implemented by respectively implementing different steps in the above method through the transceiver module 610 and the processing module 620.
[0157] In addition, the device provided in the above embodiment and the method embodiment belong to the same concept. For the specific implementation process, please refer to the method embodiment in the above text, which will not be repeated here.
[0158] The method provided by the embodiments of the present application can be executed by a computing device, which can also be referred to as a computer system. It includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a processing unit, a memory, and a memory control unit, and the functions and structures of the hardware will be described in detail subsequently. The operating system is any one or more computer operating systems that implement business processing through processes. For example, Linux operating system, Unix operating system, Android operating system, iOS operating system, or Windows operating system, etc. The application layer includes application programs such as a browser, an address book, a word processing software, and an instant messaging software. And, optionally, the computer system is a handheld device such as a smartphone, or a terminal device such as a personal computer. The present application is not particularly limited as long as it can implement the method provided by the embodiments of the present application. The execution subject of the method provided by the embodiments of the present application can be a computing device, or a functional module in the computing device that can call and execute a program.
[0159] Figure 7 FIG. 4 is a schematic structural block diagram of a computing device 700 provided by the embodiments of the present application. The computing device 700 can be a server, a computer, or other devices with computing capabilities. Figure 7 The shown computing device 700 includes: at least one processor 710 and a memory 720.
[0160] It should be understood that the present application does not limit the number of processors and memories in the computing device 700.
[0161] The processor 710 executes the instructions in the memory 720, so that the computing device 700 implements the method provided by the present application. Or, the processor 710 executes the instructions in the memory 720, so that the computing device 700 implements each functional module provided by the present application, thereby implementing the method provided by the present application.
[0162] Optionally, the computing device 700 further includes a communication interface 730. The communication interface 730 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 700 and other devices or a communication network.
[0163] Optionally, the computing device 700 further includes a system bus 740, where the processor 710, the memory 720, and the communication interface 730 are respectively connected to the system bus 740. The processor 710 can access the memory 720 through the system bus 740. For example, the processor 710 can read and write data or execute code in the memory 720 through the system bus 740. The system bus 740 is a Peripheral Component Interconnect Express (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus 740 is divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0164] In a possible implementation, the main function of the processor 710 is to interpret the instructions (or rather, the code) of a computer program and process the data in computer software. Among them, the instructions of the computer program and the data in the computer software can be stored in the memory 720 or the cache of the processor 710.
[0165] Optionally, the processor 710 may be an integrated circuit chip with signal processing capabilities. By way of example and not limitation, the processor 710 is a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Among them, the general-purpose processor is a microprocessor, etc. For example, the processor 710 is a central processing unit (CPU).
[0166] The memory 720 can provide a running space for the processes in the computing device 700. For example, the memory 720 stores the computer program (specifically, the code of the program) used to generate a process. After the computer program is run by the processor to generate a process, the processor allocates a corresponding storage space for the process in the memory 720. Further, the above storage space further includes a text segment, an initialized data segment, an uninitialized data segment, a stack segment, a heap segment, and so on. The memory 720 stores the data generated during the running of the process in the storage space corresponding to the above process, such as intermediate data, or process data, and so on.
[0167] Optionally, the memory, also known as internal memory, is used to temporarily store the operation data in the processor 710 and the data exchanged with external memories such as hard disks. As long as the computer is running, the processor 710 will transfer the data to be processed to the memory for processing, and then transfer the result out after the processing is completed.
[0168] By way of example and not limitation, the memory 720 is a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile storage medium may be, for example, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory, etc. The volatile memory is a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus dynamic random access memory (DRDRAM). It should be noted that the memory 720 of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.
[0169] The structure of the computing device 700 listed above is only an exemplary illustration, and the present application is not limited thereto. The computing device 700 in the embodiments of the present application includes various hardware in a computer system in the prior art. For example, the computing device 700 further includes other memories other than the memory 720, such as disk memories, etc. Those skilled in the art should understand that the computing device 700 may further include other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the above computing device 700 may further include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the above computing device 700 may also only include the devices necessary for implementing the embodiments of the present application, and do not have to include Figure 7All the devices shown in
[0170] An embodiment of the present application further provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0171] As Figure 8 shown, the computing device cluster includes at least one computing device 700. Instructions for executing the above method may be stored in the same manner in the memories 720 of one or more of the computing devices 700 in the computing device cluster.
[0172] In some possible implementation manners, the memories 720 of one or more of the computing devices 700 in the computing device cluster may also store partial instructions for executing the above method respectively. In other words, a combination of one or more computing devices 700 may jointly execute the instructions of the above method.
[0173] It should be noted that the memories 720 of different computing devices 700 in the computing device cluster may store different instructions respectively for executing partial functions of the above device. That is, the instructions stored in the memories 720 of different computing devices 700 may implement the functions of one or more modules in the above device.
[0174] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 9 shows a possible implementation manner. As Figure 9 shown, two computing devices 700A and 700B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device.
[0175] It should be understood that Figure 9 the functions of the computing device 700A shown in may also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700B may also be completed by multiple computing devices 700.
[0176] In this embodiment, a computer program product including instructions is further provided. The computer program product may be software or a program product including instructions that can run on the computing device cluster or be stored in any available medium. When it runs on the computing device cluster, it enables the computing device cluster to execute the above-provided method, or enables the computing device cluster to implement the functions of the above-provided device.
[0177] In this embodiment, a computer-readable storage medium is also provided. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that, when executed by a computing device cluster, cause the computing device cluster to execute the method provided above.
[0178] In an embodiment of the present application, a task scheduling system is also provided. The task scheduling system includes the above-mentioned first computing node and second computing node.
[0179] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0180] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0181] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0182] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0183] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, can exist physically alone for each unit, or two or more units can be integrated into one unit.
[0184] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0185] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A task scheduling method, characterized in that: The method is performed by a first computing node, the first computing node belongs to an infrastructure managed by a cloud management platform, the infrastructure is used to provide cloud services, the infrastructure includes multiple computing nodes, and the method includes: At a first moment, receiving information of a first task from a second computing node, the information of the first task including information required to execute the first task, the first task belonging to a first set, and the first set including at least one uncompleted task; At a second moment, sending a first request message to the second computing node, where the first request message is used to request execution of a task in the first set, where the second moment is separated from the first moment by at least a first time length, where the first time length is determined according to a first load value of the first computing node, and where the first load value is determined according to the number of tasks currently executed by the first computing node and / or the resource utilization rate of the first computing node; receiving first response information from the second computing node, the first response information being used to instruct the first computing node to execute a second task, the second task belonging to unexecuted tasks in the first set, or the first response information being used to indicate that each task in the first set is being executed; When the first response information is used to instruct the first computing node to execute a second task, the second task is executed according to the first response information.
2. The method according to claim 1, characterized in that The first load value is positively correlated with the number of tasks currently executed by the first computing node and / or the resource utilization rate of the first computing node, and the first time length is positively correlated with the first load value.
3. The method according to claim 1 or 2, characterized in that: In a case where the first response information is used to indicate that each task in the first set is being executed, the method further includes: At a third moment, the first request information is sent to the second computing node, the third moment is separated from the second moment by at least N second time lengths, the second time length is a preset time length, and N is a positive integer.
4. The method according to any one of claims 1 to 3, characterized in that In a case where the first response information is used to instruct the first computing node to perform a second task, the method further includes: At the fourth moment, a first indication message is sent to the second computing node, the first indication message is used to indicate that the first computing node is executing the second task, the fourth moment is separated from the fifth moment by M third time lengths, the fifth moment is the moment of receiving the first response message, the third time length is a preset time length, and M is a positive integer.
5. The method according to any one of claims 1 to 4, characterized in that When the first response information is used to instruct the first computing node to perform a second task, the method further includes: After the second task is executed and completed, second indication information is sent to the second computing node, where the second indication information is used to indicate that the second task has been executed and completed.
6. The method according to any one of claims 1 to 5, characterized in that The performing the second task according to the first response information includes: Determine the second task from the first set according to the first response information; The second task is executed according to the information of the second task.
7. The method according to claim 6, characterized in that The second task is the same as or different from the first task. When the second task is different from the first task, the method further includes: The information of the second task is received from the second computing node, where the information of the second task includes information required to execute the second task.
8. A task scheduling method, characterized in that: The method is performed by a second computing node, the second computing node belongs to an infrastructure managed by a cloud management platform, the infrastructure is used to provide cloud services, the infrastructure includes multiple computing nodes, and the method includes: Sending information of each task in a first set to each computing node in a computing node set, wherein the first set includes at least one uncompleted task, and the information of each task includes information required to execute each task, the computing node set includes a plurality of computing nodes, and each computing node in the computing node set is used to execute a task; receiving first request information from a first computing node, the first request information being used to request execution of a task in the first set, the first computing node belonging to the computing node set; In response to the first request information, sending first response information to the first computing node; When the first set includes at least one unexecuted task, the first response information is used to instruct the first computing node to execute a second task, and the second task belongs to the at least one unexecuted task; or When each task in the first set is being executed, the first response information is used to indicate that each task in the first set is being executed.
9. The method according to claim 8, characterized in that Each task in the first set corresponds to a first identification information, and the first identification information of each task is used to indicate whether the task is being executed.
10. The method according to claim 9, characterized in that In a case where the first response information is used to instruct the first computing node to perform a second task, the method further includes: The first identification information of the second task is set within a first time period to indicate that the second task is being executed, the starting time of the first time period includes the time when the first response information is sent, or the time before or after the time when the first response information is sent, the length of the first time period is the second time length, and the second time length is a preset time length.
11. The method according to claim 10, characterized in that The method further comprises: Before the first time period ends, receiving first indication information from the first computing node, the first indication information being used to indicate that the first computing node is executing the second task; According to the first indication information, update the first time period, the starting time of the updated first time period includes the time when the first indication information is received, or the time before or after the time when the first indication information is received, and the length of the updated first time period is the second time length; The first identification information of the second task is set within the updated first time period to indicate that the second task is being executed.
12. The method according to claim 10, characterized in that The method further comprises: In the event that the first indication information is not received from the first computing node before the end of the first time period, at the end of the first time period, the first identification information of the second task is set to indicate that the second task has not been executed, and the first indication information is used to indicate that the first computing node is executing the second task.
13. The method according to any one of claims 8 to 12, characterized in that The method further comprises: receiving second indication information from the first computing node, where the second indication information is used to indicate that the second task has been completed; deleting the second task in the first set; Sending third indication information to other computing nodes in the computing node set except the first computing node, wherein the third indication information is used to indicate that the second task has been completed, or the third indication information is used to indicate deletion of the second task in the first set.
14. A computing node, characterized in that: Comprising means for performing the method of any one of claims 1 to 7 or any one of claims 8 to 13.
15. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7 or any one of claims 8 to 13.
16. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7 or any one of claims 8 to 13.
17. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, which, when executed by a computing device cluster, perform the method as claimed in any one of claims 1 to 7 or any one of claims 8 to 13.
Citation Information
Patent Citations
Resource management method and system based on edge computing and electronic equipment
CN110688213A
Task resource reservation method and device
CN111143063A
Task scheduling method and device, electronic equipment and computer readable storage medium
CN111831420A
Parallel task initialization on dynamic computing resources
CN116348853A
Resource adjustment method and device, computing device cluster and readable storage medium
CN117290083A