Task scheduling method and device, equipment and storage medium
Through virtualization technology and resource scheduling of Kubernetes clusters, virtual resources are dynamically allocated, which solves the problem of unbalanced resource utilization in Ascend NPU resource management, improves resource utilization and task processing efficiency, and optimizes load balancing and equipment life.
Patent Information
- Application Number
- CN202510811989.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Traditional Ascend NPU resource management adopts static allocation method to have problems such as unbalanced resource utilization and difficulty in flexibly adjusting in large-scale AI batch task processing scenarios, resulting in some NPUs running overload while other NPUs are idle, unable to meet the real-time resource requirements of AI tasks.
Through virtualization technology, combining the resource scheduling and automatic expansion capabilities of Kubernetes clusters, virtual resources are dynamically allocated, and resource allocation is optimized using preset processor scoring rules to improve resource utilization and task processing efficiency.
The dynamic management of Ascend NPU resources is realized, the resource utilization rate and task processing efficiency are improved, resource waste is avoided, load balancing is optimized, equipment service life is extended, and maintenance costs are reduced.
Smart Images

Figure CN120336034A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of resource management, and particularly relates to a task scheduling method, device, equipment and storage medium. Background Art
[0002] Traditional Ascend NPU (Ascend Neural Processing Unit, a neural network processor for high-performance AI computing) resource management usually adopts a static allocation method, that is, fixed NPU resources are allocated to tasks according to predefined configurations. This method may perform well in the case of small task scales and stable resource requirements, but in large-scale AI batch task processing scenarios, there are significant limitations. First, static allocation easily leads to uneven resource utilization. Some Ascend NPUs may operate overloaded due to high task loads, while other NPUs are idle, resulting in resource waste. Second, with the dynamic changes in AI task requirements, it is difficult for the static allocation method to flexibly adjust resource allocation and cannot meet the immediate resource requirements of AI tasks.
[0003] To overcome these limitations, the industry has begun to explore resource dynamic scheduling methods based on cloud computing. As a popular container orchestration and management platform, Kubernetes provides powerful resource scheduling and auto-scaling capabilities, providing new ideas for the dynamic management of Ascend NPU resources. However, directly applying Kubernetes to the scheduling of Ascend NPU resources also faces many challenges. On the one hand, Ascend NPU resources have higher professionalism and complexity compared to traditional CPU (Central Processing Unit) and GPU (Graphic Processing Unit) resources, and more refined management and scheduling strategies are required. On the other hand, large-scale AI batch tasks usually contain multiple parallel subtasks, and there may be complex dependency relationships and resource competitions among these subtasks.
[0004] Therefore, how to efficiently schedule cloud computing tasks and improve resource utilization is a problem to be solved in this field. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a task scheduling method, device, equipment and storage medium, which can schedule and execute cloud computing tasks by means of virtualization technology through the virtual resource information corresponding to physical neural network processors, dynamically allocate virtual resources according to the running conditions of the cluster, and improve resource utilization and the processing efficiency of cloud computing tasks. The specific solutions are as follows:
[0006] In a first aspect, the present application provides a task scheduling method, which is applied to a Kubernetes cluster and includes:
[0007] Determine the current cloud computing task to be scheduled;
[0008] Based on the matching relationship between the virtual resource request corresponding to the current cloud computing task and the physical resource information of the physical neural network processors in the current cluster, determine the candidate physical neural network processors corresponding to the current cloud computing task;
[0009] Based on a preset processor scoring rule and in combination with the virtual resource request, determine the target physical neural network processor corresponding to the current cloud computing task and the target scheduling node corresponding to the target physical neural network processor from the candidate physical neural network processors;
[0010] Create virtual resource information corresponding to the virtual resource request through the target scheduling node, so as to execute the current cloud computing task by using the virtual resource information.
[0011] Optionally, the determining the current cloud computing task to be scheduled includes:
[0012] According to the number of virtual neural network processors and the video memory value in the virtual resource request corresponding to the cloud computing task, perform priority sorting on the subtasks in the obtained cloud computing batch tasks to obtain a priority queue corresponding to the cloud computing batch tasks;
[0013] Traverse the subtasks in the priority queue to determine the currently to-be-scheduled subtask as the current cloud computing task.
[0014] Optionally, the determining the candidate physical neural network processors corresponding to the current cloud computing task based on the matching relationship between the virtual resource request corresponding to the current cloud computing task and the physical resource information of the physical neural network processors in the current cluster includes:
[0015] According to the number of virtual neural network processors and the video memory value requested by the current cloud computing task, determine the matching virtual device type from the physical resource information of the physical neural network processors in the current cluster; the virtual device type is the device type corresponding to the virtual neural network processor after virtualization and segmentation of the physical neural network processor;
[0016] Determine the physical neural network processors corresponding to the virtual device type as the candidate physical neural network processors corresponding to the current cloud computing task.
[0017] Optionally, based on the preset processor scoring rules and in combination with the virtual resource request, determining a target physical neural network processor corresponding to the current cloud computing task and a target scheduling node corresponding to the target physical neural network processor from the candidate physical neural network processors includes:
[0018] Scoring the candidate physical neural network processors according to the preset processor scoring rules to obtain corresponding processor scores;
[0019] Based on the virtual resource request, the processor scores, and the number of virtual neural network processors corresponding to the candidate physical neural network processors, determining a target physical neural network processor corresponding to the current cloud computing task from the candidate physical neural network processors, and determining a target scheduling node corresponding to the target physical neural network processor.
[0020] Optionally, based on the virtual resource request, the processor scores, and the number of virtual neural network processors corresponding to the candidate physical neural network processors, determining a target physical neural network processor corresponding to the current cloud computing task from the candidate physical neural network processors includes:
[0021] Sorting the candidate physical neural network processors in descending order based on the processor scores to obtain a corresponding first list;
[0022] Traversing the first list according to the number of virtual neural network processors corresponding to the virtual resource request to determine a target physical neural network processor corresponding to the current cloud computing task from the first list.
[0023] Optionally, creating virtual resource information corresponding to the virtual resource request through the target scheduling node so as to execute the current cloud computing task by using the virtual resource information includes:
[0024] Creating a target virtual neural network processor corresponding to the virtual resource request through the target scheduling node and mounting the target virtual neural network processor to the current cloud computing task;
[0025] Executing the current cloud computing task by using the target virtual neural network processor.
[0026] Optionally, before determining the current cloud computing task to be scheduled, it further includes:
[0027] Obtaining a cloud computing batch task;
[0028] Determine whether the name of the virtual neural network processor requested by the cloud computing batch task exists in the physical resource information, and obtain a corresponding first determination result; the virtual neural network processor is a processor obtained by virtualizing and partitioning a physical neural network processor;
[0029] If the first determination result indicates that the name of the corresponding virtual neural network processor exists in the physical resource information, then determine whether the number of virtual neural network processors requested by the cloud computing batch task is greater than a preset number threshold, and obtain a corresponding second determination result;
[0030] If the second determination result indicates that the number of virtual neural network processors requested by the cloud computing batch task is not greater than the preset number threshold, then determine whether the video memory value corresponding to the virtual neural network processor requested by the cloud computing batch task is greater than a preset video memory threshold, and obtain a corresponding third determination result;
[0031] If the third determination result indicates that the video memory value corresponding to the virtual neural network processor requested by the cloud computing batch task is not greater than the preset video memory threshold, then determine the cloud computing batch task as a batch task to be scheduled, so as to determine the current cloud computing task to be scheduled from the batch tasks to be scheduled.
[0032] In a second aspect, the present application provides a task scheduling device, which is applied to a Kubernetes cluster and includes:
[0033] A task determination module, configured to determine the current cloud computing task to be scheduled;
[0034] A first processor determination module, configured to determine a candidate physical neural network processor corresponding to the current cloud computing task according to the matching relationship between the virtual resource request corresponding to the current cloud computing task and the physical resource information corresponding to the physical neural network processor in the current cluster;
[0035] A second processor determination module, configured to determine a target physical neural network processor corresponding to the current cloud computing task and a target scheduling node corresponding to the target physical neural network processor from the candidate physical neural network processors based on a preset processor scoring rule and in combination with the virtual resource request;
[0036] A task execution module, configured to create virtual resource information corresponding to the virtual resource request through the target scheduling node, so as to execute the current cloud computing task by using the virtual resource information.
[0037] In a third aspect, the present application provides an electronic device, including:
[0038] A memory, configured to store a computer program;
[0039] A processor for executing the computer program to implement the task scheduling method as described above.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, where when the computer program is executed by a processor, the task scheduling method as described above is implemented.
[0041] It can be seen that in the present application, the Kubernetes cluster first determines the current cloud computing task to be scheduled; then, according to the matching relationship between the virtual resource requests corresponding to the current cloud computing task and the physical resource information of the physical neural network processors in the current cluster, it determines the candidate physical neural network processors corresponding to the current cloud computing task; then, based on a preset processor scoring rule and in combination with the virtual resource requests, it determines the target physical neural network processor corresponding to the current cloud computing task and the target scheduling node corresponding to the target physical neural network processor from the candidate physical neural network processors; afterwards, it creates virtual resource information corresponding to the virtual resource requests through the target scheduling node, so as to execute the current cloud computing task using the virtual resource information. In this way, the present application can utilize virtualization technology to schedule and execute cloud computing tasks through the virtual resource information corresponding to physical neural network processors, which can improve the resource utilization rate of physical neural network processors; and in combination with the processor scoring rule, it can dynamically allocate virtual resources according to the running conditions of the cluster, which can improve the processing efficiency of cloud computing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0043] Figure 1 It is a flowchart of a task scheduling method disclosed in the present application;
[0044] Figure 2 It is a flowchart of a specific task scheduling method disclosed in the present application;
[0045] Figure 3 It is a schematic structural diagram of a task scheduling device disclosed in the present application;
[0046] Figure 4 It is a structural diagram of an electronic device disclosed in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] See Figure 1 As shown, an embodiment of the present invention discloses a task scheduling method applied to a Kubernetes cluster, including:
[0049] Step S11, determining the current cloud computing task to be scheduled.
[0050] In this application, the cloud computing task is submitted to the Kubernetes cluster. Usually, the user submits a cloud computing batch task to the cluster, and the batch task includes several subtasks and the specific information of each subtask; then, the current cloud computing task to be scheduled is determined according to the specific information of the batch task.
[0051] In a specific embodiment, before determining the current cloud computing task to be scheduled, it may further include: obtaining a cloud computing batch task; determining whether the name of the virtual neural network processor requested by the cloud computing batch task exists in the physical resource information to obtain a corresponding first determination result; the virtual neural network processor is a processor obtained by virtualizing and partitioning a physical neural network processor; if the first determination result indicates that the name of the corresponding virtual neural network processor exists in the physical resource information, then determining whether the number of virtual neural network processors requested by the cloud computing batch task is greater than a preset quantity threshold to obtain a corresponding second determination result; if the second determination result indicates that the number of virtual neural network processors requested by the cloud computing batch task is not greater than the preset quantity threshold, then determining whether the video memory value corresponding to the virtual neural network processor requested by the cloud computing batch task is greater than a preset video memory threshold to obtain a corresponding third determination result; if the third determination result indicates that the video memory value corresponding to the virtual neural network processor requested by the cloud computing batch task is not greater than the preset video memory threshold, then determining the cloud computing batch task as a batch task to be scheduled, so as to determine the current cloud computing task to be scheduled from the batch tasks to be scheduled. Specifically, after obtaining the cloud computing batch task submitted by the user, the specific information of the task can be checked. First, it is determined whether the physical resource information in the cluster contains the name of the virtual neural network processor requested by the cloud computing batch task. If it does not contain, there is no need to schedule the relevant cloud computing task; correspondingly, if the physical resource information of the cluster contains the name of the virtual neural network processor requested by the cloud computing batch task, then the number of virtual neural network processors requested by the cloud computing batch task can be further determined. If the requested number exceeds the corresponding limit, there is also no need to schedule the task. Further, if the number of virtual neural network processors requested by the cloud computing batch task does not exceed the preset quantity threshold, then the video memory value of the requested virtual neural network processor can be continuously determined. If the requested video memory value is greater than the preset video memory threshold, there is no need to schedule the task; otherwise, the cloud computing batch task that meets the conditions can be determined as a batch task to be scheduled. It can be understood that the batch tasks to be scheduled include several subtasks, that is, several cloud computing tasks. Subsequent individual cloud computing task scheduling needs to be carried out independently, and each cloud computing task can be processed in parallel.
[0052] In another specific embodiment, the determination of the current cloud computing task to be scheduled may include: according to the number of virtual neural network processors and the video memory value in the virtual resource request corresponding to the cloud computing task, performing a priority sorting on the subtasks in the obtained cloud computing batch task to obtain a priority queue corresponding to the cloud computing batch task; traversing the subtasks in the priority queue to determine the currently to-be-scheduled subtask as the current cloud computing task. Specifically, for the obtained cloud computing batch task, according to the number of virtual neural network processors and the video memory value in the virtual resource request corresponding to each cloud computing task therein, sorting each cloud computing task in the cloud computing batch task can obtain a priority queue corresponding to the cloud computing batch task; for example, sorting the tasks according to the magnitude of the product of the number of virtual neural network processors and the video memory value requested by the cloud computing task, moving the tasks that require a lower amount of resources forward, and finally obtaining a priority queue corresponding to the cloud computing batch task. Further, the subtasks in the priority queue can be traversed to determine the currently to-be-scheduled cloud computing task. In a specific embodiment, the priority flags of each subtask in the cloud computing batch task can also be set in advance, and then the task order can be adjusted in combination with the priority flags to obtain the final priority queue.
[0053] Step S12: Determine the candidate physical neural network processors corresponding to the current cloud computing task according to the matching relationship between the virtual resource request corresponding to the current cloud computing task and the physical resource information corresponding to the physical neural network processors in the current cluster.
[0054] In this application, through the above steps, the current cloud computing task to be scheduled and the virtual resource request corresponding to this task can be determined; then, it can be judged the resource matching relationship between the physical resource information corresponding to the physical neural network processors in the current cluster and the virtual resource request corresponding to the current cloud computing task. Based on this resource matching relationship, it can be further determined which physical neural network processors in the current cluster can be selected for the current cloud computing task, that is, the candidate physical neural network processors.
[0055] In a specific embodiment, determining the candidate physical neural network processors corresponding to the current cloud computing task according to the matching relationship between the virtual resource requests corresponding to the current cloud computing task and the physical resource information of the physical neural network processors in the current cluster may include: determining a matching virtual device type from the physical resource information of the physical neural network processors in the current cluster according to the number and video memory value of the virtual neural network processors requested by the current cloud computing task; the virtual device type is the device type corresponding to the virtual neural network processors obtained by virtualizing and partitioning the physical neural network processors; determining the physical neural network processors corresponding to the virtual device type as the candidate physical neural network processors corresponding to the current cloud computing task. Specifically, the virtual resource requests corresponding to the current cloud computing task include request information on the number and video memory value of the virtual neural network processors; while the physical resource information of the physical neural network processors in the current cluster includes the device types of the virtual neural network processors obtained by virtualizing and partitioning the physical neural network processors; according to the matching relationship between the virtual resource requests and the device types of the virtual neural network processors after corresponding partitioning of the physical resource information, several physical neural network processors suitable for the current cloud computing task can be determined from the physical resource information of the current cluster, that is, the candidate physical neural network processors corresponding to the current cloud computing task are obtained. It can be understood that the device type of the virtual neural network processors that can be created by a physical neural network processor is fixed, and the resources allocated to the virtual neural network processors are also fixed.
[0056] Step S13, based on a preset processor scoring rule, in combination with the virtual resource requests, determine the target physical neural network processor corresponding to the current cloud computing task and the target scheduling node corresponding to the target physical neural network processor from the candidate physical neural network processors.
[0057] In this application, through the above steps, candidate physical neural network processors matching the current cloud computing task can be determined from the physical resource information of the current cluster; then, the candidate physical neural network processors can be scored based on a preset processor scoring rule to determine the target physical neural network processor most suitable for the current cloud computing task according to the score, and moreover, the target physical neural network processor corresponds to a computing node, that is, finally, the target physical neural network processor and the target scheduling node corresponding to the current cloud computing task in the current cluster are determined.
[0058] In a specific embodiment, determining the target physical neural network processor corresponding to the current cloud computing task and the target scheduling node corresponding to the target physical neural network processor from the candidate physical neural network processors based on the preset processor scoring rule and in combination with the virtual resource request may include: scoring the candidate physical neural network processors through the preset processor scoring rule to obtain corresponding processor scores; determining the target physical neural network processor corresponding to the current cloud computing task from the candidate physical neural network processors based on the virtual resource request, the processor scores, and the number of virtual neural network processors corresponding to the candidate physical neural network processors, and determining the target scheduling node corresponding to the target physical neural network processor. Specifically, the candidate physical neural network processors corresponding to the current cloud computing task can be scored through the preset processor scoring rule to obtain corresponding processor scores; the scoring can consider status information such as the temperature, usage rate, and video memory of the processor; then, in combination with the number of virtual neural network processors that can be allocated corresponding to the candidate physical neural network processors, the relevant processor scores, and the virtual resource request of the current cloud computing task, the final target physical neural network processor and the corresponding target scheduling node are determined. In a specific embodiment, a part of the physical neural network processors can be screened first according to whether the virtual neural network processors are allocable, and then the target physical neural network processor corresponding to the current cloud computing task and the corresponding target scheduling node are determined in combination with the processor scores.
[0059] In another specific embodiment, determining the target physical neural network processor corresponding to the current cloud computing task from the candidate physical neural network processors based on the virtual resource request, the processor scores, and the number of virtual neural network processors corresponding to the candidate physical neural network processors may include: sorting the candidate physical neural network processors in descending order based on the processor scores to obtain a corresponding first list; traversing the first list according to the number of virtual neural network processors corresponding to the virtual resource request to determine the target physical neural network processor corresponding to the current cloud computing task from the first list. Specifically, the candidate physical neural network processors can be sorted according to the processor scores to obtain a first list, and this first list contains the number information of the virtual neural network processors supported by the physical neural network processors; then, according to the number of virtual neural network processors requested by the current cloud computing task, the target physical neural network processor for the current cloud computing task is determined. It can be understood that after allocating a physical neural network processor for the current cloud computing task, the node to which the physical neural network processor belongs can be used as the scheduling node for this cloud computing task.
[0060] Step S14: Create virtual resource information corresponding to the virtual resource request through the target scheduling node, so as to execute the current cloud computing task by using the virtual resource information.
[0061] In this application, through the above steps, the target physical neural network processor corresponding to the current cloud computing task and the corresponding target scheduling node can be determined from the physical resource information of the current cluster; then, the virtual resource information corresponding to the virtual resource request of the current cloud computing task can be created through the target scheduling node, and the task pod (cluster scheduling unit) corresponding to the current cloud computing task in the node can be processed by the created virtual resource information.
[0062] In a specific embodiment, the step of creating virtual resource information corresponding to the virtual resource request through the target scheduling node so as to execute the current cloud computing task by using the virtual resource information may include: creating a target virtual neural network processor corresponding to the virtual resource request through the target scheduling node, and mounting the target virtual neural network processor to the current cloud computing task; using the target virtual neural network processor to execute the current cloud computing task. Specifically, the target scheduling node can create virtual resource information corresponding to the virtual resource request of the current cloud computing task, that is, the target virtual neural network processor; in the Kubernetes cluster, the node can create an actual underlying virtual neural network processor (virtual device) through Ascend Docker Runtime (container runtime, used to manage and schedule NPU resources), and randomly generate a device number. The task pod corresponding to the current cloud computing task can mount the actual virtual neural network processor to its internal container through the device number to complete the binding of the virtual device and the physical device. It can be understood that subsequently, the Kubernetes cluster can use Ascend Docker Runtime to start the task pod and process the corresponding cloud computing task.
[0063] It can be seen that this application can utilize virtualization technology to schedule and execute cloud computing tasks through the virtual resource information corresponding to the physical neural network processor, which can improve the resource utilization rate of the physical neural network processor; and combined with the processor scoring rule, virtual resources can be dynamically allocated according to the running situation of the cluster, which can improve the processing efficiency of cloud computing tasks. And in the face of large-scale cloud computing batch tasks, through the resource scheduling and automatic scaling capabilities of the Kubernetes cluster, resources can be dynamically managed and adjusted immediately, and each sub-task can be scheduled in parallel to flexibly respond to task requirements, ensuring that the task can obtain the required computing resources in a timely manner, and improving the task processing efficiency.
[0064] As Figure 2 shown, the embodiment of this application discloses a task scheduling method, which specifically includes:
[0065] In this embodiment, the Kubernetes cluster involves an AI batch task management module, an AI batch task inspection module, an AI batch task scheduling module, and an AI batch task device allocation module. Specifically, a user can submit a batch task to the AI batch task management module. In the batch task, it is necessary to define the running image, start command, running environment variables, requested vNPU (Virtual Neural Processing Unit) device name, requested number of vNPU cards, and requested video memory of a single vNPU. After receiving the batch task, the AI batch task management module sends the batch task metadata information to the AI batch task processing module.
[0066] Then, the AI batch task inspection module further processes the batch task submitted by the user and checks the resource requests for vNPU in the batch task. First, it checks the requested vNPU device name. If the node resources do not include the vNPU device resources, the creation of the batch task is cancelled. Secondly, it checks the number of requested vNPU devices. If the total number of vNPU virtual devices requested by the batch task exceeds the maximum limit, the creation of the batch task is also cancelled. Finally, it checks the requested amount of vNPU resources. If the video memory request amount of each vNPU exceeds the maximum limit of a single card, the creation of the batch task is also cancelled. If the vNPU resource request check passes, the AI batch task processing module submits the user's batch task to the AI batch task scheduler module.
[0067] Furthermore, the AI batch task scheduling module further allocates Ascend NPU devices and scheduling nodes for the subtasks in the batch task according to the physical Ascend NPU resources of the entire cluster and the virtual device segmentation types supported by the NPU cards. Specifically, the AI batch task scheduling module includes an NPU scheduler and a Job (task) scheduler. Among them, the NPU scheduler allocates available vNPU virtual device types and physical NPU devices according to the vNPU video memory request value and the number of requested vNPU devices for each subtask in the current batch task. The vNPU virtual device is the underlying segmentation of the NPU device, and it can be allocated partial resources of the NPU physical device. The type of vNPU device that can be created by an NPU device is fixed, and the resources allocated by the vNPU device type are also fixed. The Job scheduler further determines the final scheduled NPU device and computing node for each subtask in the batch task according to the NPU device score situation output by the NPU scheduler.
[0068] Specifically, the NPU scheduler can calculate the usage priority of the NPU based on the utilization rate of the AICore (the core unit for performing AI calculations) of each NPU card, the AICPU utilization rate, the number of allocable vNPU virtual devices, the chip temperature, and the usage of the video memory. Through the interface provided by the NPU-smi function built into the system, the total amount and the used amount of the AICore of the NPU device i, the total amount and the used amount of the AICPU, the number of allocable vNPU virtual devices, the chip temperature, the total video memory value, and the used video memory value can be collected in real time, which are respectively represented as , , , , , , , . The NPU scheduler dynamically calculates the Core i of each NPU device, the CPU i , and the M i by perceiving the usage rate of the NPU node in real time. And by parsing the user request, the number of subtasks in the batch task and the vNPU requested video memory value of the subtasks are obtained, and the type of vNPU device required by the subtasks is calculated according to the vNPU requested video memory value. The number of allocable vNPU virtual devices VDT i on the physical NPU device is obtained according to the vNPU device type. In the scheduling, first, all NPU cards with non-allocable virtual devices on the nodes are filtered out, that is, VDT i is 0, and the NPU cards with allocable vNPU devices are retained.
[0069] Among the retained NPU devices, to better reflect the available situation of the NPU device i, each NPU device is scored in real time, and the comprehensive score S i is used to represent it, as shown in the following formula:
[0070] ;
[0071] Among them, the product of the utilization rates of the AICore, AICPU, and video memory of each NPU device is compared with the temperature of the NPU device to obtain the available score of each NPU device. The temperature of the NPU device can reflect its performance, and the lower its value, the higher the available score of the device.
[0072] After that, the ID (Identification, name), S i score, and VDT i virtual device quantity of the NPU devices of all nodes are saved to the list Q and provided for the Job scheduler to use.
[0073] Correspondingly, the Job scheduler can obtain the score list Q and the total amount of vNPU virtual devices required by the batch task, and then obtain four parameters for each subtask in the batch task: the subtask number, the subtask priority flag, the required vNPU virtual device request quantity, and the vNPU device video memory request quantity. The Job scheduler will sort the Q list according to the S i scores, and each time it will preferentially retrieve the record with the highest score, and obtain the VDT i of this record. The virtual device quantity, and subtract this quantity from the total amount of vNPU virtual devices required by the batch task, which represents that this NPU device can provide VDT i virtual devices for the batch task. Finally, record the NPU device ID, the node to which the NPU device belongs, and the VDT i virtual device quantity allocated to the batch task this time into list A. Repeat the above steps until the total amount of vNPU virtual devices required by the batch task is 0.
[0074] Next, the Job scheduler can create a subtask priority queue P, and preferentially allocate vNPU resources to subtasks with higher priorities. When the priorities of each subtask in the batch task are different and the requested vNPU resources are also different, the Job scheduler will allocate the NPU device and the node to which it belongs with the optimal score to the subtask with the highest priority. The priority sorting process is as follows: 1. Calculate the product of the vNPU virtual device quantity and the vNPU device video memory quantity requested by each subtask. When the product of the vNPU virtual device quantity and the vNPU device video memory quantity value of a certain subtask is lower than that of other subtasks, the subtask that requires a lower resource quantity is moved forward. Finally, obtain the temporary subtask priority queue F. 2. Sort the subtasks in the F queue according to the priority flag. If a certain subtask has a higher priority flag, move this subtask to the front end of F to wait for resource allocation. The final queue order obtained is the priority queue P. It can be understood that tasks with a smaller product of the vNPU device quantity and the video memory value are processed preferentially. Under the same conditions, subtasks with priority flags can be preferentially allocated resources.
[0075] It can be understood that each record in list A contains the maximum number of vNPU virtual devices that each NPU device can provide for the batch task. At this time, each subtask in the subtask priority queue P has not been assigned a specific NPU device ID and running node. Therefore, it is necessary to obtain all the subtasks in the subtask priority queue P, traverse list A, obtain each record, and bind the NPU device ID and the node to which the device belongs in this record according to the vNPU virtual device quantity required by each subtask. At the same time, match the VDT i virtual device quantity of this record with the vNPU device quantity required by this subtask. The allocation and matching process is as follows: 1. If the vNPU device quantity required by the task is greater than the VDT iIf the number of virtual devices, directly set the VDT of this record i Set the number of virtual devices to 0, record the NPU device ID and the node to which the device belongs for this task, and obtain the next record from list A to allocate VDT for it i The number of virtual devices until the number of vNPU virtual devices required by the task is 0. Then repeat the above steps, and obtain subsequent subtasks to match vNPU devices with the current record; 2. If the number of vNPU devices required by the task is less than or equal to the VDT of this record i If the number of virtual devices, directly set the VDT of this record i Subtract the number required by this subtask from the number of virtual devices set, record the NPU device ID and the node to which the device belongs for this task, and then obtain subsequent subtasks to continue vNPU device matching until the VDT of this record i The number of virtual devices is 0. Then repeat the above steps, continue to obtain the next record to allocate the physical NPU device ID and the node to which the device belongs for the current subtask. 3. If all subtasks have been allocated the physical NPU device ID and the node to which the device belongs, at this time, use the node to which the device belongs as the scheduling node of the subtask Pod. And provide the allocation relationship to the AI batch task device allocation module to complete the final creation of vNPU virtual devices and the start of the batch task Pod. It can be understood that if a single NPU device cannot meet the resource requirements of the task, a multi-device combination allocation strategy can be adopted until the task requirements are met or all available NPUs are traversed.
[0076] It should be noted that a resource buffer pool can be reserved in the Job scheduler to allocate resources for high-priority tasks; a resource allocation threshold can also be set in advance to trigger load balancing when the remaining resources are lower than the threshold, avoiding some NPUs from running overloaded due to high task loads, while reducing the idle time of other NPUs, so as to extend the service life of the devices and reduce the maintenance cost. And a resource preemption mechanism is supported between cloud computing tasks to optimize the resource allocation effect.
[0077] Furthermore, through the AI batch task scheduling module, each sub-task Pod in the batch task has been assigned a specific NPU physical device ID and a scheduling node. At this time, the Kubelet process on the scheduling node will send a vNPU device request to the AI batch task device allocation module. The AI batch task device allocation module will obtain all the sub-task Pods in the Pending state on the current node and process these sub-tasks one by one. The main process is as follows: 1. According to the physical NPU device, vNPU device type, and vNPU device request quantity set for the sub-task Pod by the scheduler module, convert the device type request into the virtual NPU device template name supported by the Ascend card according to the vNPU template, and pass the device template name and vNPU device request quantity to the Ascend Docker Runtime, which creates the actual underlying vNPU virtual device. The virtual device type is mediated. The virtual device will randomly generate a UUID (Universally Unique Identifier) device number, and the sub-task Pod mounts the actual VNPU virtual device to its internal container through this UUID. 2. After creating the vNPU virtual device, the AI batch task scheduler module will return the UUIDs of all virtual devices to the Kubelet. At this time, the Kubelet will use the Ascend DockerRuntime to start the sub-task Pod. After the sub-task Pod is in the Running state, the vNPU virtual device can be normally used inside it. 3. Wait for all the sub-task Pods in the batch task to be in the Running state, and then return the information that the batch task is successfully created. In this way, according to the actual requirements of the AI batch task, dynamic resource allocation can be performed through the virtual Ascend NPU device, enabling each Ascend NPU device to be fully utilized, thereby greatly improving the overall utilization rate of NPU resources.
[0078] It can be seen that in this solution, by introducing virtualization technology, the Ascend NPU device is finely divided to form multiple virtual Ascend NPU devices, and dynamic allocation is performed according to the actual requirements of AI batch tasks. This can effectively avoid the problems of resource idling and waste caused by static allocation, enabling each Ascend NPU device to be fully utilized, thereby greatly improving the overall utilization rate of NPU resources. It can be understood that in large-scale AI batch task processing scenarios, this solution can reasonably allocate virtual Ascend NPU resources according to the dependency relationships and resource requirements of tasks, ensuring that each subtask can run in parallel. This not only improves the efficiency of task processing but also shortens the task completion time, meeting the urgent needs of large-scale AI applications for high-performance computing NPU resources. It should be noted that compared with the traditional static allocation method, this solution can utilize the powerful resource scheduling and automatic scaling capabilities of the Kubernetes platform to achieve dynamic management and immediate adjustment of Ascend NPU resources. This can flexibly respond to the dynamic changes in AI task requirements, ensuring that tasks can immediately obtain the required computing resources, and improving the flexibility and response speed of resource scheduling. Further, this solution comprehensively considers the requirements of AI parallel tasks, the status of NPU resources, the priorities of AI tasks, and the dynamic changes of the cluster. The present invention realizes a more intelligent and efficient allocation of Ascend NPU resources and task scheduling. This helps to optimize the load balancing of the cluster, avoid some NPUs from overloading due to high task loads, and at the same time reduce the idle time of other NPUs, thereby extending the service life of the device and reducing the maintenance cost.
[0079] As Figure 3 shown, an embodiment of the present application discloses a task scheduling device applied to a Kubernetes cluster, including:
[0080] A task determination module 11, configured to determine a current cloud computing task to be scheduled;
[0081] A first processor determination module 12, configured to determine a candidate physical neural network processor corresponding to the current cloud computing task according to the matching relationship between the virtual resource request corresponding to the current cloud computing task and the physical resource information of the physical neural network processor in the current cluster;
[0082] A second processor determination module 13, configured to determine a target physical neural network processor corresponding to the current cloud computing task and a target scheduling node corresponding to the target physical neural network processor from the candidate physical neural network processors based on a preset processor scoring rule in combination with the virtual resource request;
[0083] The task execution module 14 is configured to create virtual resource information corresponding to the virtual resource request through the target scheduling node, so as to execute the current cloud computing task by using the virtual resource information.
[0084] It can be seen that the present application can utilize virtualization technology to schedule and execute cloud computing tasks by corresponding virtual resource information to a physical neural network processor, which can improve the resource utilization rate of the physical neural network processor; and combined with the processor scoring rule, virtual resources can be dynamically allocated according to the operation status of the cluster, which can improve the processing efficiency of cloud computing tasks.
[0085] In a specific embodiment, the task determination module 11 may include:
[0086] The task sorting unit is configured to sort the subtasks in the obtained cloud computing batch task according to the number of virtual neural network processors and the video memory value in the virtual resource request corresponding to the cloud computing task, so as to obtain a priority queue corresponding to the cloud computing batch task;
[0087] The task determination unit is configured to traverse the subtasks in the priority queue to determine the currently to-be-scheduled subtask as the current cloud computing task.
[0088] In a specific embodiment, the first processor determination module 12 may include:
[0089] The virtual device type determination unit is configured to determine a matching virtual device type from the physical resource information corresponding to the physical neural network processors in the current cluster according to the number of virtual neural network processors and the video memory value requested by the current cloud computing task; the virtual device type is the device type corresponding to the virtual neural network processor obtained by virtualizing and partitioning the physical neural network processor;
[0090] The first processor determination unit is configured to determine the physical neural network processor corresponding to the virtual device type as the to-be-selected physical neural network processor corresponding to the current cloud computing task.
[0091] In a specific embodiment, the second processor determination module 13 may include:
[0092] The processor scoring unit is configured to score the to-be-selected physical neural network processor through a preset processor scoring rule to obtain corresponding processor scores;
[0093] A second processor determination sub-module, configured to determine a target physical neural network processor corresponding to the current cloud computing task from the candidate physical neural network processors based on the virtual resource request, the processor score, and the number of virtual neural network processors corresponding to the candidate physical neural network processors, and determine a target scheduling node corresponding to the target physical neural network processor.
[0094] In another specific embodiment, the second processor determination sub-module may include:
[0095] A processor sorting unit, configured to sort the candidate physical neural network processors based on the principle of descending processor score to obtain a corresponding first list;
[0096] A second processor determination unit, configured to traverse the first list according to the number of virtual neural network processors corresponding to the virtual resource request, so as to determine a target physical neural network processor corresponding to the current cloud computing task from the first list.
[0097] In one specific embodiment, the task execution module 14 may include:
[0098] A task mounting unit, configured to create a target virtual neural network processor corresponding to the virtual resource request through the target scheduling node, and mount the target virtual neural network processor to the current cloud computing task;
[0099] A task execution unit, configured to execute the current cloud computing task by using the target virtual neural network processor.
[0100] In one specific embodiment, the device may further include:
[0101] A task acquisition module, configured to acquire a cloud computing batch task;
[0102] A first judgment module, configured to judge whether there is a name of a virtual neural network processor requested by the cloud computing batch task in the physical resource information to obtain a corresponding first judgment result; the virtual neural network processor is a processor obtained by virtualizing and partitioning a physical neural network processor;
[0103] A second judgment module, configured to judge whether the number of virtual neural network processors requested by the cloud computing batch task is greater than a preset number threshold when the first judgment result indicates that there is a name of a corresponding virtual neural network processor in the physical resource information, so as to obtain a corresponding second judgment result;
[0104] A third judgment module, configured to, when the second judgment result indicates that the number of virtual neural network processors of the cloud computing batch task request is not greater than a preset number threshold, judge whether the video memory value corresponding to the virtual neural network processor of the cloud computing batch task request is greater than a preset video memory threshold, so as to obtain a corresponding third judgment result;
[0105] A batch task determination module, configured to, when the third judgment result indicates that the video memory value corresponding to the virtual neural network processor of the cloud computing batch task request is not greater than a preset video memory threshold, determine the cloud computing batch task as a batch task to be scheduled, so as to determine a current cloud computing task to be scheduled from the batch tasks to be scheduled.
[0106] Further, an embodiment of the present application also discloses an electronic device, Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure cannot be regarded as any limitation on the scope of use of the present application.
[0107] Figure 4 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the task scheduling method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0108] In this embodiment, the power supply 23 is used to provide working voltages for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0109] In addition, the memory 22, as a carrier for resource storage, may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be short-term storage or permanent storage.
[0110] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the task scheduling method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program that can be used to complete other specific tasks.
[0111] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing disclosed task scheduling method is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated herein.
[0112] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0113] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0114] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0115] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0116] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A task scheduling method, characterized in that Applied to a Kubernetes cluster, including: Determine the current cloud computing task to be scheduled; Determine the candidate physical neural network processors corresponding to the current cloud computing task according to the matching relationship between the virtual resource requests corresponding to the current cloud computing task and the physical resource information of the physical neural network processors in the current cluster; Based on a preset processor scoring rule, in combination with the virtual resource requests, determine the target physical neural network processor corresponding to the current cloud computing task and the target scheduling node corresponding to the target physical neural network processor from the candidate physical neural network processors; Create virtual resource information corresponding to the virtual resource requests through the target scheduling node, so as to execute the current cloud computing task using the virtual resource information.
2. The task scheduling method according to claim 1, characterized in that The determination of the current cloud computing task to be scheduled includes: According to the number and video memory value of the virtual neural network processors in the virtual resource requests corresponding to the cloud computing tasks, perform priority sorting on the subtasks in the obtained cloud computing batch tasks to obtain the priority queue corresponding to the cloud computing batch tasks; Traverse the subtasks in the priority queue to determine the current subtask to be scheduled as the current cloud computing task.
3. The task scheduling method according to claim 1, wherein The determination of the candidate physical neural network processors corresponding to the current cloud computing task according to the matching relationship between the virtual resource requests corresponding to the current cloud computing task and the physical resource information of the physical neural network processors in the current cluster includes: Determine the matching virtual device type from the physical resource information of the physical neural network processors in the current cluster according to the number and video memory value of the virtual neural network processors requested by the current cloud computing task; the virtual device type is the device type corresponding to the virtual neural network processors obtained by virtualizing and partitioning the physical neural network processors; Determine the physical neural network processors corresponding to the virtual device type as the candidate physical neural network processors corresponding to the current cloud computing task.
4. The task scheduling method according to claim 1, wherein The determination of the target physical neural network processor corresponding to the current cloud computing task and the target scheduling node corresponding to the target physical neural network processor from the candidate physical neural network processors based on a preset processor scoring rule, in combination with the virtual resource requests, includes: Score the candidate physical neural network processors through a preset processor scoring rule to obtain the corresponding processor scores; Based on the virtual resource requests, the processor scores, and the number of virtual neural network processors corresponding to the candidate physical neural network processors, determine the target physical neural network processor corresponding to the current cloud computing task from the candidate physical neural network processors, and determine the target scheduling node corresponding to the target physical neural network processor.
5. The task scheduling method according to claim 4, wherein The determination of the target physical neural network processor corresponding to the current cloud computing task from the candidate physical neural network processors based on the virtual resource requests, the processor scores, and the number of virtual neural network processors corresponding to the candidate physical neural network processors includes: Sort the candidate physical neural network processors according to the principle of descending order of the processor scores to obtain a corresponding first list; Traverse the first list according to the number of virtual neural network processors corresponding to the virtual resource request, so as to determine the target physical neural network processor corresponding to the current cloud computing task from the first list.
6. The task scheduling method according to claim 1, wherein The creating, by the target scheduling node, of virtual resource information corresponding to the virtual resource request, so as to execute the current cloud computing task by using the virtual resource information, includes: Creating, by the target scheduling node, a target virtual neural network processor corresponding to the virtual resource request, and mounting the target virtual neural network processor to the current cloud computing task; Executing the current cloud computing task by using the target virtual neural network processor.
7. The task scheduling method according to any one of claims 1 to 6, characterized in that, Before determining the current cloud computing task to be scheduled, it further includes: Obtaining a cloud computing batch task; Judging whether there is a name of a virtual neural network processor requested by the cloud computing batch task in the physical resource information to obtain a corresponding first judgment result; the virtual neural network processor is a processor obtained by virtualizing and partitioning a physical neural network processor; If the first judgment result indicates that there is a name of a corresponding virtual neural network processor in the physical resource information, judging whether the number of virtual neural network processors requested by the cloud computing batch task is greater than a preset number threshold to obtain a corresponding second judgment result; If the second judgment result indicates that the number of virtual neural network processors requested by the cloud computing batch task is not greater than the preset number threshold, judging whether the video memory value corresponding to the virtual neural network processor requested by the cloud computing batch task is greater than a preset video memory threshold to obtain a corresponding third judgment result; If the third judgment result indicates that the video memory value corresponding to the virtual neural network processor requested by the cloud computing batch task is not greater than the preset video memory threshold, determining the cloud computing batch task as a batch task to be scheduled, so as to determine the current cloud computing task to be scheduled from the batch task to be scheduled.
8. A task scheduling device, characterized in that, Applied to a Kubernetes cluster, it includes: A task determination module, configured to determine the current cloud computing task to be scheduled; A first processor determination module, configured to determine candidate physical neural network processors corresponding to the current cloud computing task according to the matching relationship between the virtual resource request corresponding to the current cloud computing task and the physical resource information corresponding to the physical neural network processors in the current cluster; A second processor determination module, configured to determine, based on a preset processor scoring rule and in combination with the virtual resource request, a target physical neural network processor corresponding to the current cloud computing task and a target scheduling node corresponding to the target physical neural network processor from the candidate physical neural network processors; A task execution module, configured to create, by the target scheduling node, virtual resource information corresponding to the virtual resource request, so as to execute the current cloud computing task by using the virtual resource information.
9. An electronic device, characterized in that, It includes: A memory, configured to store a computer program; A processor for executing the computer program to implement the task scheduling method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program which, when executed by a processor, implements the task scheduling method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Resource scheduling method and device, electronic equipment and storage medium
CN111966500A
Resource scheduling method based on graph neural network
CN112598112A
Task processing method, device and equipment and readable storage medium
CN117453404A
Data processing method and system based on cloud service intelligent deployment
CN117492934A
Ship-based neural network algorithm processing topology generation and dynamic maintenance method and device
CN117787363A