Task processing method, first device node and computer cluster
By dynamically allocating artificial intelligence tasks to appropriate device nodes in computer clusters, the problem of insufficient computing resources in traditional personal computers is solved, computing power and resource utilization are improved, and large-scale artificial intelligence applications are supported.
Patent Information
- Application Number
- CN202510534050.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-08
AI Technical Summary
Due to limited computing resources, traditional personal computers are difficult to efficiently handle complex artificial intelligence tasks, resulting in insufficient computing power and waste of resources.
By determining the status and task type of the device node in the computer cluster, dynamically allocating tasks to the appropriate device node for processing, and using the collaborative computing power of multiple device nodes to optimize resource utilization.
It improves the computing power of computer clusters, shortens task response time, improves resource utilization efficiency, and supports large-scale artificial intelligence applications.
Smart Images

Figure CN120281772A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of computer devices, and in particular, to a task processing method, a first device node, and a computer cluster. Background Art
[0002] With the rapid development of artificial intelligence technology, the widespread use of AI (Artificial Intelligence) applications in various fields has driven the demand for high-performance computing capabilities. As the complexity and data volume of AI applications increase, the limitations of traditional personal computers gradually become apparent due to their limited computing resources when processing AI tasks. Therefore, how to solve the above problems has become a research hotspot in this field. Summary of the Invention
[0003] The embodiments of the present application provide a task processing method, a first device node, and a computer cluster, which can improve the processing efficiency of tasks and optimize the resource utilization rate.
[0004] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0005] In a first aspect, the embodiments of the present application provide a task processing method, which is applied to a first device node in a computer cluster. The method includes: in response to a target task, determining the current node status of each device node in the computer cluster; analyzing the target task to determine the task type of the target task; based on the task type of the target task and the current node status of each device node, determining a target device node from each device node; and using the target device node to process the target task.
[0006] Based on this solution, by in response to a target task, determining the current node status of each device node in the computer cluster; analyzing the target task to determine the task type of the target task; based on the task type of the target task and the current node status of each device node, determining a target device node from each device node; and using the target device node to process the target task, in this way, the computing power of the computer cluster system composed of device nodes can be improved, the inference response time of tasks can be shortened, and the resource utilization efficiency can be effectively improved, providing strong support for large-scale AI applications.
[0007] In another possible implementation, analyzing the target task to determine the task type of the target task includes: analyzing the target task based on at least one evaluation metric to determine the evaluation information of the target task under the at least one evaluation metric; and determining the task type of the target task based on the evaluation information of the target task under the at least one evaluation metric.
[0008] Based on this solution, by analyzing the target task based on at least one evaluation metric to determine the evaluation information of the target task under the at least one evaluation metric, and determining the task type of the target task based on the evaluation information of the target task under the at least one evaluation metric, in this way, the task type of the target task can be determined quickly and accurately. Then, based on this task type, the target task is assigned to a suitable device node, thereby effectively reducing the task response time, while realizing the sharing and collaboration of computing resources, improving the computing power utilization efficiency, and reducing the waste of computing resources.
[0009] In yet another possible implementation, determining the target device node from the device nodes based on the task type of the target task and the current node states of the device nodes includes: determining the node state requirements corresponding to the target task based on the task type of the target task; comparing the current node states of the device nodes with the node state requirements to obtain a comparison result; and determining the target device node from the device nodes based on the comparison result, where the node state includes at least one of the following: the load condition of the device node, the network state of the device node, the computing resources of the device node, and the response time of the device node.
[0010] Based on this solution, by determining the node state requirements corresponding to the target task based on the task type of the target task, comparing the current node states of the device nodes with the node state requirements to obtain a comparison result, and determining the target device node from the device nodes based on the comparison result, where the node state includes at least one of the following: the load condition of the device node, the network state of the device node, the computing resources of the device node, and the response time of the device node, in this way, the topological structure of the workgroup for processing the target task can be determined accurately and quickly.
[0011] In yet another possible implementation manner, in the case of including multiple target device nodes, using the target device nodes to process the target task includes: determining the number of nodes of the multiple target device nodes and the amount of computation of the target task; dividing the target task into multiple target subtasks based on the number of nodes and the amount of computation; and allocating the multiple target subtasks to the multiple target device nodes for processing according to a preset task allocation strategy.
[0012] Based on this solution, by determining the number of nodes of the multiple target device nodes and the amount of computation of the target task; dividing the target task into multiple target subtasks based on the number of nodes and the amount of computation; and allocating the multiple target subtasks to the multiple target device nodes for processing according to a preset task allocation strategy; in this way, the target task can be divided into appropriate multiple target subtasks according to the amount of computation of the target task itself and the number of nodes of the target nodes, so as to use multiple device nodes in the computer cluster to process the target task, thereby improving the utilization rate of system resources.
[0013] In yet another possible implementation manner, allocating the multiple target subtasks to the multiple target device nodes for processing according to a preset task allocation strategy includes: determining the number of tasks of the multiple target subtasks; when the number of tasks is the same as the number of nodes, allocating the multiple target subtasks to each of the multiple target device nodes for processing according to a preset task allocation strategy.
[0014] Based on this solution, by allocating corresponding target subtasks to each target device node for processing when the number of tasks of the target subtasks is the same as the number of nodes of the target device nodes; in this way, when the amount of computation of the target task is large, the determined multiple target device nodes can be used to process each target subtask respectively, thereby improving the processing efficiency of the task and the utilization rate of resources.
[0015] In yet another possible implementation manner, allocating the multiple target subtasks to the multiple target device nodes for processing according to a preset task allocation strategy includes: determining the number of tasks of the multiple target subtasks; when the number of tasks is less than the number of nodes, allocating the multiple target subtasks to some of the multiple target device nodes for processing according to a preset task allocation strategy.
[0016] Based on this solution, when the number of tasks of the target subtask is less than the number of nodes of the target device nodes, corresponding target subtasks are assigned to some of the target device nodes for processing. In this way, when the computational load of the target task is small, some of the target device nodes can be used to process each target subtask respectively, thereby improving the task processing efficiency while saving computational resources.
[0017] In another possible implementation, the method further includes: in response to receiving a first exit message sent by a second device node among the multiple target device nodes, determining the target subtasks running in the second device node; wherein, the first exit message is used to indicate that the second device node needs to exit the computer cluster; performing migration processing on the target subtasks running in the second device node.
[0018] Based on this solution, by in response to receiving a first exit message sent by a second device node among the multiple target device nodes, determining the target subtasks running in the second device node; wherein, the first exit message is used to indicate that the second device node needs to exit the computer cluster; performing migration processing on the target subtasks running in the second device node; in this way, the system reliability can be enhanced. When a node fails or is overloaded, the above task migration and fault recovery mechanism can ensure that the tasks are not interrupted, thereby ensuring the high availability of the cluster is guaranteed.
[0019] In another possible implementation, when the number of nodes of the target device node is 1, the method further includes: in response to receiving a second exit message sent by the target device node, performing migration processing on the target task running in the target device node; wherein, the second exit message is used to indicate that the target device node needs to exit the computer cluster.
[0020] Based on this solution, by when the number of nodes of the target device node is 1, in response to receiving a second exit message sent by the target device node, performing migration processing on the target task running in the target device node; in this way, when the target node device fails or is overloaded, the above task migration and fault recovery mechanism can ensure that the tasks are not interrupted, thereby ensuring the high availability of the cluster is guaranteed.
[0021] In another possible implementation, the method further includes: determining the completion status of the target subtasks corresponding to each of the target device nodes among the multiple target device nodes; based on the completion status of each target subtask, determining the completion status of the target task; and notifying the user of the completion status of the target task.
[0022] Based on this solution, by determining the completion status of the target subtasks corresponding to each target node, the completion status of the entire target task is then determined, and finally the completion status of the entire target task is notified to the user; in this way, manual intervention can be reduced, and the operation and maintenance work is greatly simplified.
[0023] In another possible implementation manner, when the number of the target device nodes is 1, the method further includes: determining the completion status of the target task in the target device node; notifying the user of the completion status of the target task.
[0024] Based on this solution, by notifying the user of the completion status of the target task in the target device node; in this way, the user experience can be improved while reducing manual intervention.
[0025] In another possible implementation manner, the method further includes: in response to receiving a first join message sent by a third device node, performing identity authentication on the third device node to obtain an identity authentication result; wherein, the first join message is used to indicate that the third device node exists in the communication network corresponding to the computer cluster; in the case where the identity authentication result is authentication passed, obtaining the node information of the third device node; saving the node information of the third device node to add the third device node to the computer cluster.
[0026] Based on this solution, by in response to receiving a first join message sent by a third device node, performing identity authentication on the third device node, and when the identity authentication is passed, exchanging node information with the third device node and automatically adding the third device node to the computer cluster; in this way, new device nodes can be automatically added or unnecessary device nodes can be removed according to the computing requirements, and the system is supported to flexibly expand according to the actual load.
[0027] In another possible implementation manner, the method further includes: in response to the startup of the first device node, performing an access registration operation to join the communication network corresponding to the computer cluster; automatically sending a second join message through the communication network corresponding to the computer cluster to broadcast that the first device node is ready to join the computer cluster; wherein, the second join message is used to indicate that the first device node exists in the communication network corresponding to the computer cluster.
[0028] Based on this solution, by in response to the startup of the first device node, performing an access registration operation to join the communication network corresponding to the computer cluster; automatically sending a second join message through the communication network corresponding to the computer cluster to broadcast that the first device node is ready to join the computer cluster; in this way, new device nodes can automatically join the cluster, thus realizing the flexible expansion of nodes in the computer cluster without manual intervention.
[0029] Second aspect, an embodiment of the present application provides a task processing device, which includes: a current state determination module, configured to determine the current node states of each device node in the computer cluster in response to a target task; a task type determination module, configured to analyze the target task to determine the task type of the target task; a target node determination module, configured to determine a target device node from each of the device nodes based on the task type of the target task and the current node states of each device node; and a task processing module, configured to use the target device node to process the target task.
[0030] Third aspect, an embodiment of the present application provides a computer cluster, the computer cluster at least includes a first device node, and the first device node is configured to determine the current node states of each device node in the computer cluster in response to a target task; analyze the target task to determine the task type of the target task; determine a target device node from each of the device nodes based on the task type of the target task and the current node states of each device node; and use the target device node to process the target task.
[0031] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is used to execute the task processing method provided in the first aspect above.
[0032] Fifth aspect, an embodiment of the present application provides a device node, which includes: a processor; a memory for storing processor-executable instructions; and the processor is configured to read the executable instructions from the memory and execute the instructions to implement the task processing method provided in the first aspect above.
[0033] Sixth aspect, an embodiment of the present application provides a computer program product, and when the instructions in the computer program product are executed by a processor, the task processing method provided in the first aspect above is executed. Description of the Drawings
[0034] Figure 1 It is a schematic structural diagram of a computer cluster provided by an embodiment of the present application.
[0035] Figure 2 It is a schematic flowchart of a task processing method provided by an embodiment of the present application.
[0036] Figure 3 It is a schematic flowchart of another task processing method provided by an embodiment of the present application.
[0037] Figure 4 It is a schematic flowchart of yet another task processing method provided by an embodiment of the present application.
[0038] Figure 5 This is a schematic structural diagram of a task processing device provided by an embodiment of the present application.
[0039] Figure 6 This is a schematic structural diagram of another task processing device provided by an embodiment of the present application.
[0040] Figure 7 This is a schematic structural diagram of yet another task processing device provided by an embodiment of the present application.
[0041] Figure 8 This is a schematic structural diagram of a first device node provided by an embodiment of the present application. Detailed implementation manners
[0042] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. For the convenience of clearly describing the technical solutions in the embodiments of the present application, the first, second, etc. descriptions that appear in the embodiments of the present application are only for schematic and differentiating the described objects, without an order, and do not represent a special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation to the embodiments of the present application.
[0043] An embodiment of the present application provides a task processing method, which determines the current node status of each device node in a computer cluster in response to a target task; analyzes the target task to determine the task type of the target task; determines a target device node from each device node based on the task type of the target task and the current node status of each device node; and uses the target device node to process the target task. In this way, it can solve the problem in the prior art that device nodes (such as AIPC) cannot meet the high requirements for computing efficiency of AI applications due to insufficient computing power. Further, the task processing method in the embodiments of the present application can improve the computing power of the computer cluster system composed of device nodes, shorten the inference response time of tasks, and effectively improve the utilization efficiency of resources, providing strong support for large-scale AI applications.
[0044] Figure 1 This is a schematic structural diagram of a computer cluster provided by an embodiment of the present application. As Figure 1 shown, the computer cluster 100 includes at least a first device node 101 and other device nodes 102 other than the first device node. The first device node 101 and the other device nodes 102 are connected through a communication network 103.
[0045] The first device node 101 can be an artificial intelligence personal computer. The artificial intelligence personal computer (AIPC) can utilize local hardware resources such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), etc. to perform computational processing of AI tasks. Therefore, the first device node 101 in the embodiments of the present application may include hardware accelerators such as a CPU, a GPU, and an NPU. Moreover, the first device node 101 can independently run AI tasks and can also cooperate with other device nodes 102 to form a computing power cluster. The other device nodes 102 can also be artificial intelligence personal computers, and the other device nodes 102 can also independently run AI tasks.
[0046] The first device node 101 and the other device nodes 102 are connected through a communication network 103 to form a computer cluster. The communication protocol within the communication network 103 supports real-time data exchange and task scheduling between nodes. For example, the communication network 103 can be a local area network, and thus the interconnection and interoperability between nodes can be achieved through standard local area network protocols (such as Ethernet or Wi-Fi), ensuring low-latency and high-bandwidth data transmission.
[0047] In some examples, a scheduling system can be deployed on the first device node 101. The scheduling system can manage the AI tasks corresponding to the AI applications in the first device node 101. For example, it can dynamically allocate the AI task to multiple different device nodes in the computer cluster 100 according to the load conditions, computing capabilities of each device node (including the first device node 101 and the other device nodes 102) in the computer cluster 100, as well as the task type and task priority of the AI task. Specifically, the scheduling system can optimize the resource allocation in the computer cluster 100 and determine detailed scheduling strategies based on machine learning algorithms or traditional load balancing algorithms.
[0048] It should be noted that Figure 1 the number of the other device nodes 102 is only an example, and in actual applications, the number of the other devices 102 can be a natural number greater than or equal to 1. Moreover, the scheduling system described in the embodiments of the present application is for more clearly explaining the technical solutions of the embodiments of the present application and does not constitute a limitation on the technical solutions provided by the embodiments of the present application.
[0049] Figure 2 is a schematic flowchart of a task processing method provided by the embodiments of the present application. This embodiment can be applied to the first device node 101 in the computer cluster 100, such as Figure 2As shown, the method includes the following steps S201 - step S204.
[0050] Step S201, in response to a target task, determine the current node status of each device node in the computer cluster.
[0051] Exemplarily, as Figure 1 shown, at least one AI application may be installed on the first device node 101. The AI application can be various types of AI applications, such as AI applications for video processing, AI applications for image processing, inference-based AI applications, etc. If the user uses the AI application, the corresponding AI task will be triggered. Further, the scheduling system in the first device node 101 can further determine the current node status of each device node (including the first device node 101 and other device nodes 102) in the computer cluster 100 in response to the trigger of the target task. For example, the computer cluster includes AIPC 1 and multiple other AIPCs (AIPC 2 to AIPC 10). The user uses the inference-based AI application on AIPC 1 to process an AI inference task. At this time, the scheduling system on AIPC 1 determines the current node status of AIPC 1 to AIPC 10 in response to the trigger of the AI inference task.
[0052] In some examples, the node status of the device node includes at least one of the following: the load condition of the device node, the network status of the device node, the computing resources of the device node, and the response time of the device node. Among them, the load condition of the device node refers to the workload borne by the device node during operation, that is, the usage rate or working intensity of the device node. The network status of the device node refers to the operating conditions and performance of the device node in the communication network. The computing resources of the device node include but are not limited to CPU resources, GPU resources, NPU resources, memory resources, hard disk resources, etc. The response time of the device node refers to the time interval from when a request is sent by the user or the system to when the result is returned.
[0053] Step S202, analyze the target task to determine the task type of the target task.
[0054] Exemplarily, the scheduling system of the first device node analyzes the target task to determine the task type of the target task. For example, when a user submits an AI task, the task scheduling system first analyzes and classifies the AI task. Furthermore, the AI task can be divided into different types, such as the first type, the second type, and the third type, etc. Among them, the first type can be a task with high computing requirements and low real-time requirements. For example, the task type of a training task is generally the first type. The second type can be a task with low computing requirements and high real-time requirements. For example, the task type of an inference task is generally the second type. The third type can be a task with high processing priority and high real-time requirements. For example, the real-time obstacle detection task during autonomous driving is generally the third type. Of course, the task type to which the target task belongs needs to be analyzed based on the specific situation of the task.
[0055] It should be noted that the number of types of the above task types, as well as the type descriptions of the first type, the second type, and the third type, are only for exemplary illustration. Those skilled in the art can set them according to specific situations during actual application.
[0056] Step S203: Determine the target device node from each device node based on the task type of the target task and the current node status of each device node.
[0057] Exemplarily, the node status requirements corresponding to the target task can be determined based on the task type of the target task; furthermore, the current node status of each device node is compared with the node status requirements to obtain a comparison result; finally, the device nodes that meet the conditions of the comparison result are determined as the target device nodes. For example, if it is analyzed that the task type of the target task is the second type, that is, the target task belongs to a task that is sensitive to latency and has a small amount of computation; then the task scheduling system will select a device node with low latency and light load as the target device node to execute the target task. Another example is that if the analyzed target task is of the first type, that is, it is determined that the target task belongs to a task with a large amount of computation but low real-time requirements; then the task scheduling system will select a device node with more computing resources as the target device node to execute the target task.
[0058] It should be noted that the target device nodes determined from each device node in the embodiments of the present application may include the first device node or may not include the first device node. That is to say, in addition to having the task scheduling and allocation function, the first device node may also have the task execution function. For example, if the analyzed target task is of the first type, that is, it is determined that the target task belongs to a task with a large amount of computation but a low real-time requirement; and the current node state of the first device node is low load and rich computing resources, then the first device node can be determined as the target device node; conversely, if the first device node has a high load and scarce computing resources at this time, the determined target device node does not include the first device node.
[0059] Moreover, the number of target device nodes determined based on the task type of the target task and the current node states of each device node can be one or multiple, and the embodiments of the present application do not limit this.
[0060] Step S204: Use the target device node to process the target task.
[0061] Exemplarily, after determining the target device node, the task scheduling system will use at least one target device node to process the target task. In this way, through the cooperation between multiple device nodes, the efficient processing of AI tasks can be achieved, overcoming problems such as insufficient computing power, waste of computing resources, poor scalability of the cluster, and complexity of task scheduling existing in the existing device nodes.
[0062] In some examples, after the scheduling system in the first device node determines multiple target device nodes, it can disassemble the target task into multiple target subtasks according to the number of nodes included in the target device node and the size of the target task itself, and then use a preset task allocation strategy to allocate the multiple target subtasks to at least one target device node for processing. The preset task allocation strategy includes, but is not limited to: load balancing strategy, resource priority strategy, real-time task scheduling strategy, polling strategy, shortest job first strategy, minimum load first strategy, etc.
[0063] In some examples, if an AI inference task is triggered, the first device node can, in response to the trigger of the AI inference task, based on the computing requirements of the AI inference task and the resource status of each device node in the computer cluster, allocate the AI inference task to a suitable device node in the cluster. For example, if the GPU resources of a certain device node in the computer cluster are idle and the computing power is strong, then this device node will give priority to processing the computing part of the AI inference task. In some examples, the scheduling system of the first device node will monitor the load status and task progress of each device node in the computer cluster in real time to ensure that tasks can be evenly distributed and avoid overloading of some device nodes or idling of computing resources.
[0064] The task processing method provided by the embodiment of the present application determines the current node status of each device node in the computer cluster in response to a target task; analyzes the target task to determine the task type of the target task; based on the task type of the target task and the current node status of each device node, determines the target device node from each device node; and uses the target device node to process the target task. In this way, the computing power of the computer cluster system composed of device nodes can be improved, the inference response time of the task can be shortened, and the utilization efficiency of resources can be effectively improved, providing strong support for large-scale AI applications.
[0065] In some embodiments, step S202, analyzing the target task to determine the task type of the target task, includes: analyzing the target task based on at least one evaluation metric to determine the evaluation information of the target task under at least one evaluation metric; and determining the task type of the target task based on the evaluation information of the target task under at least one evaluation metric.
[0066] Exemplarily, some evaluation metrics can be preset in advance, and then the target task is analyzed based on at least one evaluation metric to determine the evaluation information of the target task under at least one evaluation metric. Furthermore, the task type of the target task is determined based on the evaluation information of the target task under at least one evaluation metric. The evaluation metric includes but is not limited to: the computing amount of the task, the real-time requirement of the task, etc. For example, the computing amount of the target task and the real-time requirement of the target task can be analyzed to obtain the evaluation information of the target task under these two evaluation metrics; furthermore, based on the evaluation information of the target task under these two evaluation metrics, the task type of the target task is determined. For example, if the computing amount of the target task is greater than the preset value and the real-time requirement of the target task is low, it is determined that the task type of the target task is the first type. In this way, the task type of the target task can be determined quickly and accurately, and then the target task is allocated to the appropriate device node based on this task type, thereby effectively reducing the task response time, realizing the sharing and cooperation of computing resources at the same time, improving the utilization efficiency of computing power, and reducing the waste of computing resources.
[0067] In some embodiments, step S203, determining the target device node from each device node based on the task type of the target task and the current node status of each device node, includes: determining the node status requirement corresponding to the target task based on the task type of the target task; comparing the current node status of each device node with the node status requirement to obtain a comparison result; and determining the target device node from each device node based on the comparison result; where the node status includes at least one of the following: the load condition of the device node, the network status of the device node, the computing resources of the device node, and the response time of the device node.
[0068] Exemplarily, the node status requirements corresponding to the target task can be determined based on the task type of the target task; then, the node status requirements are compared and matched with the current actual node status of each node, and based on the comparison and matching results, the device nodes with successful comparison and matching are used as the target device nodes. Furthermore, the target device nodes with successful comparison and matching form a working group corresponding to the target task to jointly process the target task. In this way, the topological structure of the working group for processing the target task can be accurately and quickly determined.
[0069] As Figure 3 shown, based on the above Figure 2 shown embodiment, step S204 may include the following steps S2041 - step S2043.
[0070] Step S2041, in the case of including multiple target device nodes, determine the number of nodes of the multiple target device nodes and the computing amount of the target task.
[0071] Exemplarily, in the embodiments of the present application, after determining that the working group corresponding to the target task includes multiple target device nodes, it is necessary to determine the number of nodes of the multiple target device nodes and the computing amount of the target task. Among them, the computing amount of the target task refers to the size of the data volume involved and the computing resources required when executing the target task. For example, there are 10 device nodes in a computer cluster, and the number of nodes of the target device nodes determined from the 10 device nodes is 5.
[0072] Step S2042, based on the number of nodes and the computing amount, divide the target task into multiple target subtasks.
[0073] Exemplarily, the target task can be divided into multiple target subtasks based on the number of nodes included in the target device nodes and the computing amount of the target task. For example, if the number of nodes of the target device nodes is 5 and the computing amount of the target task is large, the target task can be divided into 5 different target subtasks. Another example is that if the number of nodes of the target device nodes is 5 and the computing amount of the target task is small, the target task can be divided into 3 different target subtasks.
[0074] Step S2043, according to the preset task allocation strategy, allocate the multiple target subtasks to the multiple target device nodes for processing.
[0075] Exemplarily, in the embodiments of the present application, after determining the multiple target device nodes and the multiple target subtasks corresponding to the target task, the preset task allocation strategy can be used to determine how to allocate the multiple target subtasks to the multiple target device nodes in combination with the current status of each target device node (such as load, idle computing resources, response time, etc.).
[0076] In some examples, the task allocation strategies include, but are not limited to: load balancing strategy, resource priority strategy, real-time task scheduling strategy, polling strategy, shortest job first strategy, and minimum load first strategy. Among them, the load balancing strategy: dynamically allocates tasks according to the computing loads of each node to avoid overloading a single node. The resource priority strategy: allocates tasks with higher priorities to nodes with strong computing capabilities and idle resources to ensure the optimal use of computing resources. The real-time task scheduling strategy: for tasks with high real-time requirements (such as inference tasks), preferentially selects nodes with low network latency and moderate loads for computing to ensure the response time. The polling strategy: evenly distributes tasks to each node and is suitable for scenarios with relatively balanced computing amounts. The shortest job first strategy: preferentially allocates tasks with smaller computing requirements to nodes with idle resources to reduce the task queuing time. The minimum load first strategy: allocates tasks to the node with the minimum current load to prevent individual nodes from being overloaded.
[0077] The task processing method provided by the embodiments of this application divides a target task into multiple target subtasks according to the number of nodes of the target device node and the computing amount of the target task; then processes the multiple target subtasks by using multiple target device nodes according to a preset task allocation strategy, thereby completing the processing of the target task; in this way, it is possible to divide the target task into appropriate multiple target subtasks according to the target task itself and the number of nodes of the target device node, so as to process the target task by using multiple target device nodes in a computer cluster, and further improve the utilization rate of system resources.
[0078] In some embodiments, step S2043, allocating the multiple target subtasks to multiple target device nodes for processing according to a preset task allocation strategy, includes: determining the number of tasks of the multiple target subtasks; when the number of tasks is the same as the number of nodes, allocating the multiple target subtasks to each target device node among the multiple target device nodes for processing according to a preset task allocation strategy.
[0079] Exemplarily, if the computing amount of the target task is large, the target task can be divided into multiple target subtasks with the same number as the number of nodes of the target device node. For example, if the number of nodes of the target device node is 5 and the computing amount of the target task is large, the target task can be divided into 5 different target subtasks. At this time, the number of tasks of the target subtasks is the same as the number of nodes, and then according to a preset task allocation strategy, the multiple target subtasks are allocated to each target device node among the multiple target device nodes for processing. For example, the 5 different target subtasks are respectively allocated to 5 target device nodes, and one target device node runs one target subtask. In this way, when the computing amount of the target task is large, the multiple determined target device nodes can be used to process each target subtask respectively, thereby improving the task processing efficiency and the resource utilization rate.
[0080] In some embodiments, step S2043 of allocating a plurality of target subtasks to a plurality of target device nodes for processing according to a preset task allocation policy includes: determining the number of tasks of the plurality of target subtasks; and when the number of tasks is less than the number of nodes, allocating the plurality of target subtasks to some of the target device nodes among the plurality of target device nodes for processing according to the preset task allocation policy.
[0081] Exemplarily, although a plurality of determined target device nodes are included, if the computational load of the target task is small, the number of tasks of the divided target subtasks may be less than the number of nodes of the target device nodes. For example, if the number of nodes of the target device nodes is 5 and the computational load of the target task is small, the target task may be divided into 3 different target subtasks. At this time, the number of tasks of the target subtasks is less than the number of nodes, and then according to the preset task allocation policy, the plurality of target subtasks are allocated to some of the target device nodes among the plurality of target device nodes for processing. For example, 3 different target subtasks are allocated to a certain 3 of the 5 target device nodes for processing. In this way, when the computational load of the target task is small, some target device nodes can be used to process each target subtask respectively, thereby improving the processing efficiency of the task while saving computational resources.
[0082] In some embodiments, after the above step S2043, the following steps may further be included:
[0083] Step S21c: In response to receiving a first exit message sent by a second device node among the plurality of target device nodes, determining the target subtasks running in the second device node; wherein the first exit message is used to indicate that the second device node needs to exit the computer cluster.
[0084] Exemplarily, in the embodiment of the present application, after the scheduling system of the first device node assigns the target task to multiple target device nodes for processing, if the second device node among the multiple target device nodes needs to exit the computer cluster it belongs to at this time, the second device node will send a first exit message in the communication network where the computer cluster is located to broadcast that it needs to exit the computer cluster. Furthermore, after the first device node receives the first exit message, the scheduling system of the first device node will determine the target subtasks that are currently running in the second device node. For example, when an AIPC node among multiple AIPC target nodes needs to exit the computer cluster due to a fault or maintenance, it will notify other nodes in the computer cluster through the local area network where the computer cluster is located to perform resource transfer and task scheduling adjustment. Furthermore, the first device node can quickly detect the failure of the second device node and migrate the subtasks in the failed second device node to other available device nodes according to the fault recovery strategy to avoid a decline in cluster performance. After the second device node exits, its resources will also be automatically removed from the cluster to ensure that other device nodes can work properly.
[0085] In some examples, when a device node needs to exit the computer cluster (for example, when the node needs to be maintained or its hardware is replaced, it needs to exit the computer cluster), the device node will first notify other nodes in the computer cluster that it needs to exit the computer cluster. Before the device node exits, the scheduling system of the first device node will determine whether there are target subtasks assigned to the device node that needs to exit at this time and whether the target subtasks have been completed. If there are target subtasks and they have not been completed, the scheduling system of the first device node will migrate them to other nodes according to the requirements of the target subtasks to ensure that the cluster will not interrupt services due to the exit of a certain device node. After the device node exits, the topology of the computer cluster will be automatically adjusted, and new target tasks will be reallocated according to the computing power of the remaining nodes to ensure the efficient operation of the cluster.
[0086] Step S22c: Migrate the target subtasks running in the second device node.
[0087] Exemplarily, in the case where the second device node broadcasts its need to exit the computer cluster, if the scheduling system of the first device node determines that the target subtask running in the second device node is in the completion stage at this time, it obtains the completion result of the target subtask and deletes the second device node and the relevant node information of the second device node from its own computer cluster list. If the scheduling system of the first device node determines that the target subtask running in the second device node is in the execution stage at this time, it migrates the target subtask running in the second device node. For example, the target subtask running in the second device node can be migrated to other device nodes except the second device node among the multiple target device nodes. It can also be migrated to other device nodes except the multiple target device nodes in the computer cluster. The embodiments of the present application do not limit this, and those skilled in the art can migrate the target subtask according to the specific completion situation of the target subtask and the node status of the existing device nodes in the computer cluster.
[0088] In some examples, if a device node in the computer cluster cannot continue to process tasks due to a fault or high load, the task scheduling system of the first device node will detect this problem node in real time. If the problem node belongs to the multiple target device nodes, the task scheduling system of the first device node will automatically migrate the target subtask being executed on the problem node to other healthy nodes to ensure the uninterrupted execution of the task. And the scheduling system of the first device node will evaluate the load conditions of other healthy nodes and select the most suitable healthy node to take over the task to avoid task loss or delay. At the same time, the loads of other device nodes in the cluster will be adjusted accordingly to ensure the balanced utilization of computing resources.
[0089] The task processing method provided by the embodiments of the present application determines the target subtask running in the second device node by responding to the first exit message sent by the second device node among the multiple target device nodes; wherein, the first exit message is used to indicate that the second device node needs to exit the computer cluster; and migrates the target subtask running in the second device node; in this way, the system reliability can be enhanced. When a node fails or is overloaded, the above task migration and fault recovery mechanisms can ensure the uninterrupted execution of the task, thereby ensuring the high availability of the cluster.
[0090] In some implementation manners, when the number of target device nodes is 1, it further includes: responding to the second exit message sent by the target device node, and migrating the target task running in the target device node; wherein, the second exit message is used to indicate that the target device node needs to exit the computer cluster.
[0091] Exemplarily, when the determined target device node is 1, the target task is executed in this target device node. If at this time the target device broadcasts that it needs to exit the computer cluster, the scheduling system in the first device node will perform migration processing on the target task running in the target device node; in this way, when the target node device fails or is overloaded, the above task migration and fault recovery mechanism can ensure that the task is not interrupted, thus ensuring the high availability of the cluster is guaranteed.
[0092] In some embodiments, after the above step S2043, the following steps may further be included:
[0093] Step S21d: Determine the completion status of the target subtasks corresponding to each target device node among the multiple target device nodes.
[0094] Exemplarily, after the scheduling system of the first device node assigns each target subtask to each target device node for processing, the scheduling system of the first device node can determine the completion status of the target subtasks on each target device node among the multiple target device nodes. For example, a preset time can be set, and every interval of this preset time, the scheduling system of the first device node needs to determine the completion status of the target subtasks corresponding to each target device node.
[0095] Step S22d: Based on the completion status of each target subtask, determine the completion status of the target task.
[0096] Exemplarily, the target task includes each target subtask, and thus the completion status of the target task can be determined through the completion status of each target subtask. If the scheduling system of the first device node determines that the target task is in a completed state, the result will be notified to the user.
[0097] Step S23d: Notify the user of the completion status of the target task.
[0098] The task processing method provided by the embodiments of the present application determines the completion status of the entire target task by determining the completion status of the target subtasks corresponding to each target device node, and finally notifies the user of the completion status of the entire target task; in this way, manual intervention can be reduced, and the operation and maintenance work is greatly simplified.
[0099] In some embodiments, if there are target subtasks corresponding to multiple different target tasks running in a certain target device node, the target device node will process the target subtasks corresponding to different target tasks in sequence according to the distribution time of the target subtasks corresponding to different target tasks.
[0100] In some embodiments, when the number of nodes of the target device node is 1, it further includes: determining the completion status of the target task in the target device node; notifying the user of the completion status of the target task;
[0101] Exemplarily, when the determined target device node is one, the target task is executed in the target device node. The scheduling system in the first device node needs to determine the completion status of the target task in the target device node at regular intervals and notify the user of the completion status; in this way, the user experience can be improved while reducing manual intervention.
[0102] In some embodiments, the following method steps are further included:
[0103] Step S21a: In response to receiving a first joining message sent by a third device node, authenticate the third device node to obtain an authentication result; wherein, the first joining message is used to indicate that the third device node exists in the communication network corresponding to the computer cluster.
[0104] Exemplarily, in the case where the first device node has joined the computer cluster, whether the first device node is in the state of executing a task or in the state of allocating a task or in the idle state, the first device node can respond to receiving the first joining message sent by the third device node, and perform identity authentication with the third device node through the device discovery protocol to add the node information of the third device node to the computer cluster list.
[0105] In some examples, if the third device node has joined the communication network corresponding to the computer cluster, the third device node will send a first joining message through the communication network to broadcast its own existence in the communication network. Furthermore, after the first device node receives the first joining message sent by the third device node, it can perform identity authentication with the third device node through the device discovery protocol. For example, after the first device node receives the broadcast signal of the third device node, it will automatically perform identity authentication with it and add the third device node to the computer cluster after the authentication is passed.
[0106] Step S22a: In the case where the authentication result is authentication passed, obtain the node information of the third device node.
[0107] Exemplarily, if the first device node authenticates the third device node successfully, the scheduling system of the first device node can obtain the node information of the third device node. For example, if the identity authentication result of the third device node passes, the first device node and the third device node will exchange node information with each other, and the node information includes the hardware configuration of the node, the idle resources of the node (such as GPU and NPU computing capabilities), the network bandwidth occupied by the node, etc.
[0108] Step S23a: Save the node information of the third device node to add the third device node to the computer cluster.
[0109] Exemplarily, the scheduling system of the first device node may save the node information of the third device node and add the node information of the third device node to the computer cluster list, thereby adding the third device node to the computer cluster.
[0110] The task processing method provided by the embodiments of the present application authenticates the identity of the third device node by responding to the received first join message sent by the third device node to obtain an identity authentication result; wherein, the first join message is used to indicate that the third device node exists in the communication network corresponding to the computer cluster; in the case where the identity authentication result is authentication passed, obtain the node information of the third device node; save the node information of the third device node to add the third device node to the computer cluster; in this way, it is possible to automatically add new nodes or remove unnecessary nodes according to computing requirements, and support the system to flexibly expand according to the actual load.
[0111] As Figure 4 shown, on the basis of the above Figure 2 shown embodiment, before step S201, the following steps S21b-S22b are further included.
[0112] Step S21b, in response to the startup of the first device node, perform an access registration operation to join the communication network corresponding to the computer cluster.
[0113] Step S22b, automatically send a second join message through the communication network corresponding to the computer cluster to broadcast that the first device node is ready to join the computer cluster; wherein, the second join message is used to indicate that the first device node exists in the communication network corresponding to the computer cluster.
[0114] Exemplarily, if the first device node starts up, after detecting the network where the computer cluster is located, the first device node automatically performs an access registration operation to join the communication network corresponding to the computer cluster. After the first device node joins the communication network corresponding to the computer cluster, it will send a second join message through this communication network to broadcast that it is ready to join the computer cluster, thereby achieving the purpose of automatically joining the computer cluster. For example, when each AIPC device starts up in the local area network, it will automatically perform device discovery and broadcast its presence information to the network through a broadcast protocol (such as mdns or zeroconf).
[0115] In some examples, when the first device node (such as a certain AIPC in a computer cluster) starts up, the first device node will send a broadcast signal within the local area network to inform other device nodes that the first device node is online and ready to join the cluster. Other device nodes can detect the first device node by listening for these broadcast signals. Among them, the first device node and other device nodes can authenticate each other through a standard device discovery protocol (such as mdns, zeroconf, etc.) and exchange basic system information (such as hardware configuration, computing power, network bandwidth, etc.). If other device nodes confirm that the first device node is available, other device nodes will add the first device node to the cluster node list to achieve the automatic joining of the first device node.
[0116] The task processing method provided by the embodiments of this application performs a network access registration operation in response to the startup of the first device node to join the communication network corresponding to the computer cluster; and automatically sends a second join message through the communication network corresponding to the computer cluster to broadcast that the first device node is ready to join the computer cluster; wherein, the second join message is used to indicate the existence of the first device node in the communication network corresponding to the computer cluster; in this way, new nodes can automatically join the cluster, thereby realizing the flexible expansion of nodes in the computer cluster without manual intervention.
[0117] Corresponding to the embodiments of the foregoing task processing method, the present application also provides embodiments of a task processing device. Figure 5 A task processing device provided by the embodiments of this application, as Figure 5 shown, the task processing device 500 includes a current state determination module 501, a task type determination module 502, a target node determination module 503, and a task processing module 504.
[0118] The current state determination module 501 is configured to determine the current node state of each device node in the computer cluster in response to a target task;
[0119] The task type determination module 502 is configured to analyze the target task to determine the task type of the target task;
[0120] The target node determination module 503 is configured to determine a target device node from each device node based on the task type of the target task and the current node state of each device node;
[0121] The task processing module 504 is configured to use the target device node to process the target task.
[0122] In another possible implementation, the task type determination module 502 is configured to analyze the target task based on at least one evaluation metric to determine the evaluation information of the target task under at least one evaluation metric; and determine the task type of the target task based on the evaluation information of the target task under at least one evaluation metric.
[0123] In yet another possible implementation, the target node determination module 503 is configured to determine the node status requirements corresponding to the target task based on the task type of the target task; compare the current node status of each device node with the node status requirements to obtain a comparison result; and determine the target device node from each device node based on the comparison result; where the node status includes at least one of the following: the load condition of the device node, the network status of the device node, the computing resources of the device node, and the response time of the device node.
[0124] In yet another possible implementation, the task processing module 504 is configured to, in the case of including multiple target device nodes, determine the number of target device nodes and the computing amount of the target task; divide the target task into multiple target subtasks based on the number of nodes and the computing amount; and allocate the multiple target subtasks to the multiple target device nodes for processing according to a preset task allocation policy.
[0125] In yet another possible implementation, as Figure 6 shown, on the basis of the above Figure 5 the task processing apparatus 500 further includes a subtask determination module 505 and a task migration module 506.
[0126] The subtask determination module 505 is configured to determine the target subtask running in the second device node in response to receiving a first exit message sent by the second device node among the multiple target device nodes; where the first exit message is used to indicate that the second device node needs to exit the computer cluster.
[0127] The task migration module 506 is configured to perform migration processing on the target subtask running in the second device node.
[0128] In yet another possible implementation, the task processing apparatus 500 further includes a task progress determination module.
[0129] The task progress determination module is configured to determine the completion status of the target subtasks corresponding to each target device node among the multiple target device nodes; determine the completion status of the target task based on the completion status of each target subtask; and notify the user of the completion status of the target task.
[0130] In yet another possible implementation, the above task processing apparatus 500 further includes an identity authentication module, a node information acquisition module, and a node storage module.
[0131] An identity authentication module, configured to authenticate the third device node in response to receiving a first joining message sent by the third device node, and obtain an identity authentication result; wherein, the first joining message is used to indicate that the third device node exists in the communication network corresponding to the computer cluster;
[0132] A node information acquisition module, configured to acquire the node information of the third device node when the identity authentication result is authentication passed;
[0133] A node storage module, configured to store the node information of the third device node to add the third device node to the computer cluster.
[0134] In another possible implementation, as Figure 7 shown, based on the above Figure 5 , the task processing device 500 further includes an automatic network access module 51 and a message sending module 52.
[0135] The automatic network access module 51 is configured to perform network access registration operations in response to the startup of the first device node to join the communication network corresponding to the computer cluster;
[0136] The message sending module 52 is configured to automatically send a second joining message through the communication network corresponding to the computer cluster to broadcast that the first device node is ready to join the computer cluster; wherein, the second joining message is used to indicate that the first device node exists in the communication network corresponding to the computer cluster.
[0137] For the beneficial technical effects corresponding to the exemplary embodiments of the above task processing device 500, reference may be made to the corresponding beneficial technical effects in the above method embodiment section, which will not be elaborated here.
[0138] Corresponding to the embodiment of the foregoing task processing method, the present application further provides an embodiment of a computer cluster. As Figure 1 shown, the computer cluster 100 at least includes a first device node 101. Wherein,
[0139] The first device node 101 is configured to determine the current node status of each device node in the computer cluster in response to a target task; analyze the target task to determine the task type of the target task; based on the task type of the target task and the current node status of each device node, determine a target node from each device node; and use the target device node to process the target task.
[0140] It should be noted that for the beneficial technical effects corresponding to the exemplary embodiments of the above computer cluster 100, reference may be made to the corresponding beneficial technical effects in the above method embodiment section, which will not be elaborated here.
[0141] In some examples, corresponding to the embodiments of the foregoing task processing method, the present application further provides a first device node, which includes a processor and a memory for storing instructions executable by the processor; wherein, the processor is configured to read the executable instructions from the memory and execute the instructions to implement the steps in the embodiments of the foregoing task processing method.
[0142] It should be noted that for the beneficial technical effects corresponding to the exemplary embodiments of the foregoing first device node, reference may be made to the corresponding beneficial technical effects in the method embodiment section above, and details will not be elaborated herein.
[0143] Figure 8 FIG. is a schematic structural diagram of a first device node 800 provided in an embodiment of the present application. As Figure 8 shown, the hardware entities of the first device node 800 include: a processor 801, a communication interface 802, and a memory 803, where
[0144] The processor 801 generally controls the overall operation of the first device node 800.
[0145] The communication interface 802 enables the first device node 800 to communicate with other electronic devices or servers via a network.
[0146] The memory 803 is configured to store instructions and applications executable by the processor 801, and can also cache data to be processed or already processed by the processor 801 and each module in the first device node 800, and can be implemented by FLASH (flash memory) or RAM (Random Access Memory).
[0147] In addition to the foregoing methods and devices, an embodiment of the present application may further provide a computer program product, including computer program instructions, which, when run by a processor, cause the processor to execute the steps in the task processing methods of various embodiments of the present application described in the method embodiment section above.
[0148] The computer program product can be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0149] In addition, an embodiment of the present application may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps in the task processing method of various embodiments of the present application described in the method embodiment part above.
[0150] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium, for example but not limited to, includes a system, device or component of electricity, magnetism, light, electromagnetic, infrared ray, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0151] The basic principles of the present application have been described above in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that they are essential for each embodiment of the present application. In addition, the specific details of the above embodiments are only for the purposes of illustration and easy understanding, rather than limitations. The above details do not limit the present application to necessarily adopt the above specific details for implementation.
[0152] Those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.
[0153] Moreover, the above-described embodiments are only specific embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present application shall be included in the protection scope of the present application.
Claims
1. A task processing method, characterized in that, Applied to the first device node in a computer cluster, the method includes: In response to a target task, determining the current node status of each device node in the computer cluster; Analyzing the target task to determine the task type of the target task; Based on the task type of the target task and the current node status of each device node, determining a target device node from each of the device nodes; Using the target device node to process the target task.
2. The method according to claim 1, wherein The analyzing the target task to determine the task type of the target task includes: Analyzing the target task based on at least one evaluation metric to determine evaluation information of the target task under the at least one evaluation metric; Based on the evaluation information of the target task under the at least one evaluation metric, determining the task type of the target task.
3. The method according to claim 1, wherein The determining a target device node from each of the device nodes based on the task type of the target task and the current node status of each device node includes: Based on the task type of the target task, determining the node status requirements corresponding to the target task; Comparing the current node status of each device node with the node status requirements to obtain a comparison result; Based on the comparison result, determining the target device node from each of the device nodes; Wherein, the node status includes at least one of the following: the load condition of the device node, the network status of the device node, the computing resources of the device node, and the response time of the device node.
4. The method according to claim 1, wherein In the case of including multiple target device nodes, the using the target device node to process the target task includes: Determining the number of nodes of the multiple target device nodes and the computing amount of the target task; Based on the number of nodes and the computing amount, dividing the target task into multiple target subtasks; According to a preset task allocation policy, allocating the multiple target subtasks to the multiple target device nodes for processing.
5. The method according to claim 4, wherein The method further includes: In response to receiving a first exit message sent by a second device node among the multiple target device nodes, determining the target subtask running in the second device node; wherein, the first exit message is used to indicate that the second device node needs to exit the computer cluster; Performing a migration process on the target subtask running in the second device node.
6. The method according to claim 4, characterized in that, The method further includes: Determining the completion status of the target subtask corresponding to each target device node among the multiple target device nodes; Based on the completion status of each target subtask, determining the completion status of the target task; Notifying the user of the completion status of the target task.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: In response to receiving a first join message sent by a third device node, performing identity authentication on the third device node to obtain an identity authentication result; wherein, the first join message is used to indicate that the third device node exists in the communication network corresponding to the computer cluster; In the case where the identity authentication result is authentication passed, obtaining the node information of the third device node; Save the node information of the third device node to add the third device node to the computer cluster.
8. The method according to any one of claims 1-6, characterized in that The method further includes: In response to the startup of the first device node, perform an access registration operation to join the communication network corresponding to the computer cluster; Automatically send a second join message through the communication network corresponding to the computer cluster to broadcast that the first device node is ready to join the computer cluster; wherein, the second join message is used to indicate the existence of the first device node in the communication network corresponding to the computer cluster.
9. A computer cluster, characterized in that, The computer cluster at least includes: A first device node; The first device node is configured to, in response to a target task, determine the current node status of each device node in the computer cluster; analyze the target task to determine the task type of the target task; based on the task type of the target task and the current node status of each device node, determine a target node from each device node; and use the target device node to process the target task.
10. A first device node, characterized in that, The first device node includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the task processing method according to any one of claims 1 to 8 above.