Task scheduling method and device, storage medium and electronic equipment

By building a task model and network environment information filtering, selecting computing nodes and optimizing resource allocation, the problem of inefficient scheduling of deep learning training tasks is solved, and more efficient resource utilization and task execution are achieved.

CN120407125AActive Publication Date: 2025-08-01CHINA TOWER CO LTD

Patent Information

Application Number
CN202510888583.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-01
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

In the prior art, deep learning training task scheduling is inefficient and cannot effectively cope with the real-time state of tasks and resources, resulting in waste of resources and delayed tasks, affecting the overall training efficiency and user experience.

Method used

By building a task model, predict the performance results of the training task on multiple computing nodes, combine network environment information to filter and select computing nodes that meet preset requirements, perform training tasks, and optimize resource allocation and scheduling.

Benefits of technology

It improves the scheduling efficiency of deep learning training tasks and the utilization rate of computing resources, reduces task execution time and resource waste, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407125A_ABST
    Figure CN120407125A_ABST
Patent Text Reader

Abstract

The invention discloses a task scheduling method and device, a storage medium and electronic equipment, and relates to the field of cloud computing. The method comprises the steps that a task model is constructed according to a target training task, and the task model is used for representing attributes and resource requirements of the target training task; on the basis of the task model, predicting performance results when the target training task is executed on the N computational nodes to obtain N performance results, N being an integer greater than 1, and the performance results comprising execution efficiency and resource utilization rate of the target training task on the computational nodes; filtering the N computing nodes according to the N performance results and the network environment information to obtain S first computing nodes, and determining a target node from the S first computing nodes; and executing the target training task through the target node. The technical problem of low deep learning training task scheduling efficiency in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing, and in particular, to a task scheduling method, device, storage medium, and electronic device. Background Art

[0002] In the context of current cloud computing and big data processing, the efficient scheduling of deep learning training tasks has become a key bottleneck restricting the rapid iteration and deployment of AI models. Traditional cluster job scheduling algorithms, such as First-Come-First-Served (FCFS), Shortest Job First (SJF), or priority-based scheduling, although performing well in general computing tasks, are inefficient in processing deep learning tasks. The main reason is that these algorithms fail to fully consider the dynamicity and heterogeneity of the resource requirements of deep learning tasks, as well as the complexity of inter-task communication. For example, large model training often requires high computing power and large video memory space, while real-time lightweight inference tasks emphasize low latency and network stability. Existing technologies often adopt static strategies in resource allocation and cannot respond to the real-time status of tasks and resources, resulting in resource waste and task delays, seriously affecting the overall training efficiency and user experience.

[0003] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] This application provides a task scheduling method, device, storage medium, and electronic device to at least solve the technical problem of low efficiency in scheduling deep learning training tasks in the prior art.

[0005] According to one aspect of this application, a task scheduling method is provided, including: constructing a task model based on a target training task, where the task model is used to characterize the attributes and resource requirements of the target training task; predicting performance results of the target training task when executed on N computing nodes based on the task model to obtain N performance results, where N is an integer greater than 1, and the performance results include the execution efficiency and resource utilization rate of the target training task on the computing nodes; filtering the N computing nodes according to the N performance results and network environment information to obtain S first computing nodes, and determining a target node from the S first computing nodes, where S is an integer greater than or equal to 1 and less than or equal to N, the first computing nodes are used to characterize the computing nodes that meet the network environment information, and the target node is used to characterize the first computing node that meets the preset requirements, where the preset requirements are used to constrain the execution duration of the computing node for the target training task; and executing the target training task through the target node.

[0006] Optionally, based on the task model, predict the performance results when the target training task is executed on N computing nodes, obtaining N performance results, including: predicting the first value of the target training task on each computing node based on the task model, obtaining N first values, where the first value is used to characterize the computing resource consumption of the graphics processor in each computing node when executing the target training task; predicting the second value of the target training task on each computing node based on the task model, obtaining N second values, where the second value is used to characterize the transmission time generated by data exchange between computing nodes; predicting the third value of the target training task on each computing node based on the task model, obtaining N third values, where the third value is used to characterize the video memory space required by the graphics processor in each computing node when executing the target training task; determine the N performance results according to the N first values, N second values, and N third values.

[0007] Optionally, the task model includes the target neural network for training the target training task, the number of iterations, and the batch size, where the batch size is used to characterize the amount of data processed by the target neural network in each round of training. Predicting the first value of the target training task on each computing node based on the task model, obtaining N first values, includes: when there are M network layers in the target neural network, determining the computational amount of each network layer according to the target information of each network layer in the M network layers and the batch size, obtaining M computational amounts, where the target information is used to characterize the parameter information required by each network layer when executing the target training task; determining the computational amount of the target neural network according to the M computational amounts and the number of iterations; determining the first value corresponding to each computing node according to the computational amount of the target neural network and the number of graphics processors corresponding to each computing node, obtaining N first values.

[0008] Optionally, the task model further includes a parallel mode, where the parallel mode is used to characterize the strategy of resource allocation and computing task division during the training process of the target neural network. Predicting the second value of the target training task on each computing node based on the task model, obtaining N second values, includes: when the parallel mode is data parallel, determining the second value corresponding to each computing node according to the number of graphics processors corresponding to each computing node, the number of parameters of each network layer in the M network layers, and the number of iterations, obtaining N second values; when the parallel mode is tensor parallel, determining the second value corresponding to each computing node according to the number of graphics processors corresponding to each computing node, the feature tensors output by each network layer in the M network layers, the batch size, and the number of iterations, obtaining N second values.

[0009] Optionally, based on the task model, predict the third value of the target training task on each computing node to obtain N third values, including: determining an optimizer, where the optimizer is used to adjust and store the model parameters of the target neural network; determining the first impact value of the optimizer on the video memory in the graphics processor, where the first impact value is used to quantify the video memory space occupied by the optimizer to store data; determining the second impact value according to the task type of the target training task, where the second impact value is used to quantify the occupancy of the intermediate feature map output by any network layer in the target neural network in the video memory of the graphics processor; determining the third value corresponding to each computing node according to the target information, the first impact value, the second impact value, and the batch size of any network layer in the M network layers to obtain N third values.

[0010] Optionally, filter the N computing nodes according to the N performance results and the network environment information to obtain S first computing nodes, and determine the target node from the S first computing nodes, including: filtering the N computing nodes according to the N third values corresponding to the N computing nodes in the N performance results and the network environment information to obtain S first computing nodes; calculating the fourth value of each first computing node according to the first value, the second value, and the average utilization rate of the graphics processor corresponding to each first computing node in the S first computing nodes to obtain S fourth values, where the fourth value is used to represent the total time consumption of each first computing node to execute the target training task; using the first computing node corresponding to the fourth value less than the preset threshold among the S fourth values as the target node.

[0011] Optionally, filter the N computing nodes according to the N third values corresponding to the N computing nodes in the N performance results and the network environment information to obtain S first computing nodes, including: detecting whether the third value corresponding to each computing node in the N performance results is greater than or equal to the video memory capacity corresponding to the computing node in the network environment information; filtering the computing nodes whose N third values corresponding to the N computing nodes are greater than or equal to the video memory capacity corresponding to the computing node in the network environment information; using the computing nodes whose N third values corresponding to the N computing nodes are less than the video memory capacity corresponding to the computing node in the network environment information as the first computing nodes to obtain S first computing nodes.

[0012] According to another aspect of the present application, there is also provided a task scheduling device, including: a construction unit configured to construct a task model according to a target training task, where the task model is used to characterize the attributes and resource requirements of the target training task; a prediction unit configured to predict performance results when the target training task is executed on N computing nodes based on the task model, obtaining N performance results, where N is an integer greater than 1, and the performance results include the execution efficiency and resource utilization rate of the target training task on the computing nodes; a determination unit configured to filter the N computing nodes according to the N performance results and network environment information to obtain S first computing nodes, and determine a target node from the S first computing nodes, where S is an integer greater than or equal to 1 and less than or equal to N, the first computing nodes are used to characterize the computing nodes that meet the network environment information, and the target node is used to characterize the first computing node that meets the preset requirements, where the preset requirements are used to restrict the duration of the target training task executed by the computing node; and an execution unit configured to execute the target training task through the target node.

[0013] According to another aspect of the present application, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes an executable program stored therein, and when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the above task scheduling method.

[0014] According to another aspect of the present application, there is also provided an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above task scheduling method.

[0015] In this application, first, a task model is constructed based on the target training task, where the task model is used to represent the attributes and resource requirements of the target training task. Secondly, based on the task model, the performance results of the target training task when executed on N computing nodes are predicted to obtain N performance results, where N is an integer greater than 1, and the performance results include the execution efficiency and resource utilization rate of the target training task on the computing nodes. Then, according to the N performance results and the network environment information, the N computing nodes are filtered to obtain S first computing nodes, and a target node is determined from the S first computing nodes, where S is an integer greater than or equal to 1 and less than or equal to N. The first computing nodes are used to represent the computing nodes that meet the network environment information, and the target node is used to represent the first computing node that meets the preset requirements, where the preset requirements are used to constrain the duration of the computing node to execute the target training task. Finally, the target training task is executed through the target node, that is, by constructing the task model, predicting the performance, and filtering and selecting the computing nodes based on the network environment information, the purpose of accurately matching the computing resources with the needs of the deep learning training task is achieved, thereby realizing the technical effect of improving the scheduling efficiency of the deep learning training task and the optimized utilization of computing resources, and further solving the technical problem of low scheduling efficiency of the deep learning training task in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0017] Figure 1 is a flowchart of an optional task scheduling method according to an embodiment of the present application;

[0018] Figure 2 is a structural diagram of an optional task scheduling method according to an embodiment of the present application;

[0019] Figure 3 is a schematic diagram of an optional task scheduling method according to an embodiment of the present application;

[0020] Figure 4 is a schematic diagram of an optional task scheduling device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0023] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set between this system and relevant users or institutions to provide corresponding operation entrances for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, the expert decision-making process will be entered.

[0024] According to the embodiments of this application, a method embodiment of a task scheduling method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described here can be executed in a different order than here.

[0025] It should be noted that an intelligent scheduling system can be the execution subject of the task scheduling method in the embodiments of this application. It can be understood that the task scheduling method provided by the embodiments of this application can also be executed by other systems or devices as the execution subject, and the embodiments of this application do not make specific limitations on this.

[0026] Figure 1 is a flowchart of an optional task scheduling method according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0027] Step S101: Construct a task model based on the target training task.

[0028] In step S101, the task model is used to characterize the attributes and resource requirements of the target training task.

[0029] Optionally, the target training task refers to a deep learning training task submitted by a user, which has a specific model structure, algorithm hyperparameters, dataset size, and expected training goal.

[0030] Optionally, the intelligent scheduling system receives and parses the task request information submitted by the user, including the model structure, batch size, number of iterations, neural network structure, and parallel mode, etc. These information are compiled and modeled into a formatted model that describes the task attributes and resource requirements, that is, the task model.

[0031] Optionally, the task model is the basis for subsequent resource requirement evaluation and task scheduling. It ensures that the intelligent scheduling system can accurately understand the task attributes and resource requirements and provides accurate data input for the subsequent steps.

[0032] Step S102: Predict the performance results of the target training task when executed on N computing nodes based on the task model, and obtain N performance results.

[0033] In step S102, N is an integer greater than 1, and the performance results include the execution efficiency and resource utilization rate of the target training task on the computing node.

[0034] Optionally, the N computing nodes here N represents the number of all available nodes in the intelligent scheduling system, including various GPU (Graphics Processing Unit) acceleration nodes, and each node has different computing capabilities, video memory sizes, and network bandwidths.

[0035] Optionally, the performance results refer to the expected execution efficiency (such as the amount of computation of the task, the communication overhead of the task) and resource utilization rate (such as video memory usage rate, etc.) of the deep learning training task when executed on each computing node.

[0036] Optionally, through prediction, the intelligent scheduling system can estimate the expected performance of the task on different nodes in advance, which helps to filter out inappropriate nodes and provides a decision basis for subsequent task scheduling.

[0037] Step S103: Filter the N computing nodes based on the N performance results and network environment information to obtain S first computing nodes, and determine a target node from the S first computing nodes.

[0038] In step S103, S is an integer greater than or equal to 1 and less than or equal to N. The first computing nodes are used to represent the computing nodes that meet the network environment information, and the target node is used to represent the first computing node that meets the preset requirements.

[0039] In step S103, the preset requirements are used to restrict the duration for the computing nodes to execute the target training task.

[0040] Optionally, the network environment information refers to a series of metric data related to the network status of the intelligent scheduling system and the status of computing nodes, which is collected and monitored in real time by the GPU cluster management module. This information plays a crucial role in the process of deep learning training task scheduling and resource allocation, ensuring the optimization of task performance and the efficient utilization of cluster resources.

[0041] Optionally, the network environment information includes but is not limited to GPU resource status (including the utilization rate, idle status, video memory usage, computing performance, etc. of each GPU, which is used to evaluate the ability and availability of the GPU to execute tasks), CPU resource status (CPU utilization rate, number of idle cores, memory usage, etc. Although deep learning tasks mainly rely on GPUs, the CPU resource status also affects the task preparation and data processing speed), network status (involving key performance indicators such as network bandwidth, latency, and packet loss rate between computing nodes. For distributed deep learning tasks, the network status directly affects the synchronization efficiency of model parameters and data transmission speed), storage status (including the read / write speed and remaining space of storage devices. For data-intensive deep learning tasks, storage performance directly affects data loading and storage efficiency), and node health status (the online status, fault records, temperature, power status, etc. of nodes, ensuring that the computing nodes are in a normal working state and avoiding task scheduling to faulty or unstable nodes).

[0042] Optionally, the S first computing nodes are nodes that the system filters out by comparing the performance results of the computing nodes with the network environment information, and can effectively execute deep learning tasks and meet the network condition requirements.

[0043] Optionally, the target node refers to the node that the intelligent scheduling system further filters out from the S first computing nodes and meets the preset requirements, that is, the first computing node that can execute the target training task with the minimum time cost.

[0044] Optionally, the intelligent scheduling system further narrows down the node selection range based on the network environment information and performance prediction results, ensuring that the selected nodes not only have sufficient computing resources but also excellent network conditions, which can reduce the overall time overhead of task execution.

[0045] Step S104: Execute the target training task through the target node.

[0046] Optionally, the target training task is assigned to the target node for execution. The intelligent scheduling system distributes resources such as task configuration files and data sets to the target node through the task scheduling module, starts the container environment, and executes the deep learning task, and monitors the task status and resource usage in real time to ensure the smooth operation of the task.

[0047] From the content of steps S101 to S104, it can be seen that in this application, first, a task model is constructed based on the target training task, where the task model is used to characterize the attributes and resource requirements of the target training task. Secondly, based on the task model, the performance results of the target training task when executed on N computing nodes are predicted, and N performance results are obtained, where N is an integer greater than 1, and the performance results include the execution efficiency and resource utilization rate of the target training task on the computing node. Then, according to the N performance results and network environment information, the N computing nodes are filtered to obtain S first computing nodes, and the target node is determined from the S first computing nodes, where S is an integer greater than or equal to 1 and less than or equal to N, and the first computing node is used to characterize the computing node that meets the network environment information, and the target node is used to characterize the first computing node that meets the preset requirements, where the preset requirements are used to restrict the duration of the computing node to execute the target training task. Finally, the target training task is executed through the target node, that is, by means of task model construction, performance prediction, and computing node filtering and selection based on network environment information, the purpose of accurately matching the computing resources with the needs of deep learning training tasks is achieved, thereby realizing the technical effect of improving the scheduling efficiency of deep learning training tasks and optimizing the utilization of computing resources, and further solving the technical problem of low scheduling efficiency of deep learning training tasks in the prior art.

[0048] In an alternative embodiment, the intelligent scheduling system predicts a first value of the target training task on each computing node based on the task model, obtaining N first values, where the first value is used to characterize the computing resource consumption of the graphics processor in each computing node when executing the target training task. The intelligent scheduling system predicts a second value of the target training task on each computing node based on the task model, obtaining N second values, where the second value is used to characterize the transmission time generated by data exchange between computing nodes. The intelligent scheduling system predicts a third value of the target training task on each computing node based on the task model, obtaining N third values, where the third value is used to characterize the video memory space required by the graphics processor in each computing node when executing the target training task. Then, N performance results are determined according to the N first values, the N second values, and the N third values.

[0049] Optionally, the graphics processor refers to a GPU accelerator, which is the core processing unit for parallel computing of deep learning tasks and has high parallel computing capabilities and a large amount of video memory.

[0050] Optionally, the first value refers to the prediction of the computing resource consumption of the GPU in each computing node when executing the target training task, that is, it reflects the computing overhead of the task on different nodes.

[0051] Optionally, the second value refers to the prediction of the transmission time generated by data exchange (such as gradient synchronization) between computing nodes, which is used to evaluate the communication efficiency of the task, that is, the communication overhead of the task.

[0052] Optionally, the third value refers to the prediction of the video memory space required by the GPU in each computing node when executing the task, ensuring the reasonable allocation of video memory resources to avoid overflow.

[0053] Optionally, first, the intelligent scheduling system constructs a task model based on the deep learning task description uploaded by the user, including the model structure (such as convolutional layer, fully connected layer, etc.), algorithm parameters (such as batch size, number of iterations), parallel strategy (data parallel or tensor parallel), etc. Then the system calls the API instructions provided by the Pytorch framework (a machine learning library), such as torch.summary, to analyze the model structure and parameter configuration, calculate the number of parameters, computational volume, and size of the intermediate feature map of each layer of the model, etc., and then predict the computational resource consumption when the GPU in each computing node executes the task, obtaining N first numerical values. Based on the Ring-AllReduce architecture (a communication mode for distributed deep learning training), the system calculates the synchronous communication overhead between each GPU node under different parallel strategies (data parallel or tensor parallel), including the communication volume and transmission time within and between layers, obtaining N second numerical values. Also based on the task model, the system evaluates the demand for video memory space by the model parameters, intermediate feature maps, and optimizer selection (such as SGD (Stochastic Gradient Descent), Adam (Adaptive Moment Estimation)), considering the impact of different task types (training or inference) on video memory occupancy, and predicts the video memory space required when the GPU in each computing node executes the task, obtaining N third numerical values.

[0054] Optionally, the intelligent scheduling system combines the N first numerical values, N second numerical values, and N third numerical values, comprehensively analyzes the task execution situation on each computing node, and calculates N performance results through an algorithm. These results not only consider the computational overhead but also take into account the communication latency and video memory occupancy, thus comprehensively reflecting the expected performance of the task on different nodes.

[0055] As can be seen from the above, through the task model, the intelligent scheduling system can perform fine-grained prediction on the resource requirements of deep learning tasks, accurate to computational overhead, communication time, and video memory occupancy. And it can achieve a comprehensive evaluation of task performance and dynamic optimal allocation of resources, ultimately significantly improving the execution efficiency and resource utilization level of deep learning tasks.

[0056] In an optional embodiment, when there are M network layers in the target neural network, the intelligent scheduling system determines the computational volume of each of the M network layers according to the target information of each network layer in the M network layers and the batch size, obtaining M computational volumes. The target information is used to characterize the parameter information required for each network layer to execute the target training task. Then, according to the M computational volumes and the number of iterations, it determines the computational volume of the target neural network. Then, according to the computational volume of the target neural network and the number of graphics processors corresponding to each computing node, it determines the first numerical value corresponding to each computing node, obtaining N first numerical values.

[0057] Optionally, the target neural network is a neural network model submitted by the user for deep learning training, and its structure and parameters will be used as the basis for scheduling decisions; the M network layers refer to the set of all layers included in the target neural network, where M is the number of network layers, such as convolutional layers and fully connected layers; target information: describes the parameter information required for each network layer to execute the target training task. For example, for a convolutional layer, it includes, but is not limited to, the number of input channels, the height, width, and number of channels of the output feature map, the width of the convolutional kernel, etc.; the computational workload quantifies the computational requirements of the computational task for the graphics processor, including the number of multiplications, additions, and the number of gradient updates for convolutional layers and fully connected layers; N first values: N is the total number of computing nodes, and each node corresponds to a first value, reflecting the computational overhead of its GPU when processing deep learning tasks.

[0058] Optionally, the intelligent scheduling system first analyzes the structure of the target neural network, extracts the target information of each network layer, including the number of parameters of the layer, the size of the input and output features, etc., laying the foundation for subsequent computational workload prediction. Based on the target information and the batch size, the system predicts the computational workload of each of the M network layers. For example, for a convolutional layer, the computational workload of each layer is calculated according to the number of input channels, the height, width, and number of channels of the output feature map, and the width of the convolutional kernel. The fully connected layer predicts the computational workload based on the number of input neurons, the number of output neurons, and the batch size. M computational workloads are obtained, each corresponding to a layer in the network. The system further determines the total computational workload of the entire target neural network according to the M computational workloads and the number of iterations. The total computational workload is equal to the sum of the computational workloads of each layer multiplied by the number of iterations, reflecting the computational resources required to complete one training cycle.

[0059] Optionally, considering that the number of GPUs in each computing node is different in distributed training, the system predicts the computational resource consumption (i.e., the first value) of each GPU in each computing node when executing deep learning tasks according to the computational workload of the target neural network and the number of GPUs in each computing node. The parallel processing ability of the GPU is considered during the calculation to ensure that the prediction result is consistent with the actual situation. The system performs the above prediction process for all N computing nodes, and finally obtains N first values, each clearly reflecting the computational resource consumption of the GPU on the corresponding computing node when executing deep learning tasks.

[0060] As can be seen from the above, by extracting the target information of the network layer layer by layer and combining the batch size and the number of iterations, the intelligent scheduling system can accurately predict the computational volume of the target neural network at each computational layer, and then accurately evaluate the computational resource consumption corresponding to each computing node. Moreover, by predicting the computational resource consumption corresponding to the computing node, the execution efficiency of the deep learning task can be significantly improved, while optimizing the allocation of GPU resources, avoiding resource waste, and improving the resource utilization rate of the entire computing cluster.

[0061] In an alternative embodiment, the task model further includes a parallel mode, where the parallel mode is used to characterize the strategy of resource allocation and computational task division during the training process of the target neural network. Based on the task model, the second values of the target training task on each computing node are predicted, obtaining N second values, including: when the parallel mode is data parallelism, the intelligent scheduling system determines the second value corresponding to each computing node according to the number of graphics processors corresponding to each computing node, the number of parameters of each of the M network layers, and the number of iterations, obtaining N second values; when the parallel mode is tensor parallelism, the intelligent scheduling system determines the second value corresponding to each computing node according to the number of graphics processors corresponding to each computing node, the feature tensors output by each of the M network layers, the batch size, and the number of iterations, obtaining N second values.

[0062] Optionally, the parallel mode is used to describe the strategy of resource allocation and computational task division during the training process of the target neural network, which is mainly divided into two types: data parallelism and tensor parallelism.

[0063] Optionally, data parallelism means dividing the training data set into several parts, each part independently performing the same model calculation on different computing nodes, and finally aggregating the calculation results of each node. This method is often used for training tasks with relatively small model parameters and large data volumes.

[0064] Optionally, tensor parallelism means dividing the model weight parameters or feature maps into multiple tensors, and calculating them on the GPUs of different computing nodes respectively, and finally updating the model by communicating and aggregating the tensor results. This method is suitable for scenarios with relatively large model parameters and heavy computational loads.

[0065] Optionally, the intelligent scheduling system first determines whether the target training task adopts a data parallelism or tensor parallelism strategy based on the parallel mode field in the task model.

[0066] Optionally, if the parallel mode is data parallelism, the intelligent scheduling system will predict the time required for data exchange based on the number of GPUs in each computing node, the number of parameters in each network layer, and the number of iterations. Specifically, when each network layer synchronizes gradients among GPUs through the Ring-AllReduce mechanism, there is a fixed amount of communication, which depends on the number of model parameters, the number of GPUs, and the number of iterations. The intelligent scheduling system calculates the data exchange transmission time between each computing node based on these factors, that is, the second value corresponding to each node, and finally obtains N second values.

[0067] Optionally, if tensor parallelism is adopted, the intelligent scheduling system will predict the communication time based on the number of GPUs in each computing node, the size of the feature tensors output by each network layer, the batch size, and the number of iterations. In tensor parallelism, the calculation of communication time involves how to effectively split and recombine model parameters or intermediate feature maps. The intelligent scheduling system combines this data with the network bandwidth and the number of GPUs to predict the transmission time required for each computing node to synchronize feature tensors, and also obtains N second values.

[0068] As can be seen from the above, the intelligent scheduling system can refine the prediction of the second value of communication overhead in the parallel training process according to the specific parallel mode (data parallelism or tensor parallelism), ensuring the accuracy of the prediction results. This helps to plan and optimize the data transmission strategy in advance. By predicting the communication delay between each computing node, the system can better coordinate the parallel strategy and resource allocation, ensure that data transmission will not become a bottleneck, and improve the overall execution efficiency of deep learning tasks. Accurately predicting the communication time of parallel training can help the intelligent scheduling system optimize the task placement strategy, reduce unnecessary communication delays, thus significantly improving the performance of distributed deep learning training tasks and shortening the task completion time.

[0069] In an optional embodiment, based on the task model, the third value of the target training task on each computing node is predicted, obtaining N third values, including: First, the intelligent scheduling system determines the optimizer, where the optimizer is used to adjust and store the model parameters of the target neural network. Then, the first influence value of the optimizer on the video memory in the graphics processor is determined, where the first influence value is used to quantify the video memory space occupied by the optimizer to store data. Then, the second influence value is determined according to the task type of the target training task, where the second influence value is used to quantify the occupancy of the intermediate feature map output by any network layer in the target neural network in the video memory of the graphics processor. Finally, according to the target information, the first influence value, the second influence value, and the batch size of any network layer among the M network layers, the third value corresponding to each computing node is determined, obtaining N third values.

[0070] Optionally, an optimizer refers to an algorithm used to adjust model parameters in deep learning training, such as SGD, Adam, etc. Different optimizers have a direct impact on the memory usage of the video memory.

[0071] Optionally, the first impact value is used to quantify the occupancy of the GPU video memory space by the data stored by the optimizer (such as gradients, momentum, etc.). Different optimizer selections will result in different first impact values; the second impact value is used to quantify the occupancy of the intermediate feature maps output by any network layer in the target neural network in the GPU video memory, and is used to evaluate the dynamic demand for video memory during task execution.

[0072] Optionally, the task type includes training tasks and inference tasks. Since the storage requirements for intermediate feature maps are different during the execution of the two, the task type directly affects the determination of the second impact value.

[0073] Optionally, the intelligent scheduling system first determines the type of optimizer used for the target training task, such as SGD, Adam, etc. Then, it calculates the first impact value of the optimizer on the GPU video memory according to the type of optimizer. For example, SGD usually needs to save gradient values equivalent to the number of parameters, while the Adam optimizer needs to save the first and second order momenta additionally, so the occupied video memory space is 4 times the number of parameters. The system quantifies the contribution of the optimizer to video memory occupancy through this calculation.

[0074] Optionally, the intelligent scheduling system determines the second impact value according to the type of the target training task (training or inference). Intermediate feature maps need to be saved during the execution of training tasks, so its second impact value (the size of the intermediate feature maps occupying the video memory) is usually 2 times the size of the feature maps; while inference tasks do not need to save these intermediate results, so its second impact value is relatively low, usually 1 times the size of the feature maps. This difference reflects the specific impact of the task type on video memory requirements.

[0075] Optionally, the intelligent scheduling system further comprehensively calculates the total demand for video memory when the GPU on each computing node executes the target training task according to the target information of each network layer in the M network layers, the first impact value of the optimizer, the second impact value of the task type, and the batch size. This includes the video memory occupancy of model parameters, optimizer data, and intermediate feature maps, and takes into account the different impacts of data parallelism and tensor parallelism on video memory requirements in distributed training. Finally, the system obtains N third numerical values, that is, the video memory demand predictions of N computing nodes. Through the above steps, the intelligent scheduling system performs video memory demand prediction on all N computing nodes and obtains N third numerical values, which reflect the video memory occupancy of different computing nodes during the execution of deep learning training tasks.

[0076] As can be seen from the above, the intelligent scheduling system can accurately predict the video memory space requirements of the target training task on each computing node according to the optimizer type, task type, parameter information of the network layer, and batch size, so as to avoid over-allocation or shortage of video memory resources and ensure the successful execution of the task. And by accurately predicting the third value, the intelligent scheduling system can dynamically adjust the resource allocation strategy according to the video memory requirements of the training task, select the most suitable computing node for each task, reduce the risk of video memory overflow, and improve the utilization efficiency of the GPU video memory. In addition, the accurate prediction and reasonable allocation of video memory requirements significantly improve the execution reliability of deep learning tasks, while reducing unnecessary waiting time and resource waste.

[0077] In an optional embodiment, the intelligent scheduling system filters the N computing nodes according to the N third values corresponding to the N computing nodes in the N performance results and the network environment information to obtain S first computing nodes, and then calculates the fourth value of each first computing node according to the first value, the second value corresponding to each first computing node in the S first computing nodes and the average utilization rate of the graphics processor, obtaining S fourth values, where the fourth value is used to represent the total time consumption of each first computing node for executing the target training task. After that, the first computing nodes corresponding to the fourth values less than the preset threshold in the S fourth values are used as the target nodes.

[0078] Optionally, the intelligent scheduling system filters out the computing nodes that can meet the video memory requirements of the target training task and have excellent network conditions according to the video memory requirements (third values) of the N computing nodes and the network environment information, obtaining S first computing nodes. The system then further evaluates the S first computing nodes. For each first computing node, the intelligent scheduling system combines the first value, the second value with the average utilization rate of the current graphics processor to calculate the total expected time consumption for executing the target training task, that is, the fourth value. This value reflects the total execution time of the task on a specific node, including computing time, communication time, and waiting time.

[0079] Optionally, the system sets a preset threshold to limit the maximum time consumption of task execution. The first computing nodes corresponding to the fourth values in the S fourth values that are lower than or equal to the preset threshold are used as the final target nodes. These nodes are considered to be the most suitable for executing the target training task because they can complete the task within a reasonable time while maintaining the effective utilization of resources.

[0080] As can be seen from the above, through the preliminary filtering of video memory and network conditions, and the comprehensive evaluation of computing resource consumption, data transmission time, GPU utilization, etc., the intelligent scheduling system can ensure the effective utilization of resources and significantly improve the execution efficiency of deep learning tasks. The introduction of preset thresholds enables the system to accurately screen out the computing nodes that can complete tasks with the optimal time consumption, avoiding resource waste and excessive task execution time. Through a series of screening and evaluation processes, it is ensured that deep learning training tasks can be intelligently assigned to the most suitable computing nodes for execution, not only reducing the task completion time, but also improving the utilization efficiency of computing power resources, thereby providing users with faster and more stable deep learning training services and optimizing the overall operation efficiency of the cluster.

[0081] In an optional embodiment, the intelligent scheduling system detects whether the third value corresponding to each computing node in the N performance results is greater than or equal to the video memory capacity of the computing node in the network environment information, filters out the computing nodes among the N third values corresponding to the N computing nodes that are greater than or equal to the video memory capacity of the computing node in the network environment information, and takes the computing nodes among the N third values corresponding to the N computing nodes that are less than the video memory capacity of the computing node in the network environment information as the first computing nodes, obtaining S first computing nodes.

[0082] Optionally, the video memory capacity refers to the total video memory of the GPU in the computing node, which is used to store model parameters, intermediate calculation results, etc.

[0083] Optionally, the intelligent scheduling system first obtains the N performance results generated by the resource demand evaluation module, and each result includes the synchronous communication time overhead, computing time overhead, and GPU video memory occupancy requirement (third value) of the target training task on different computing nodes. The system checks the third value in the N performance results one by one, that is, the video memory space required by the GPU when executing the deep learning task on each computing node, and compares it with the video memory capacity of the same computing node recorded in the network environment information. For those computing nodes whose third value (video memory demand) is greater than or equal to the video memory capacity in the network environment information, the intelligent scheduling system will exclude them from the candidate list to ensure that the system does not assign tasks to nodes with insufficient video memory resources and avoid task failures caused by video memory overflow. After filtering, the system retains those computing nodes whose third value (video memory demand) is less than the video memory capacity in the network environment information, that is, the GPU video memory capacity of these nodes is sufficient to support the execution of deep learning tasks, and marks these nodes as the first computing nodes, a total of S first computing nodes.

[0084] Optionally, by performing a fine-grained analysis of the user's deep learning tasks, the intelligent scheduling system is equivalent to constructing a task scheduling framework with time priority, realizing online prediction of the computing power resource requirements of different deep learning tasks and the task performance changes under different placement strategies, so as to improve the task execution efficiency and cluster resource utilization rate of the deep learning R & D platform.

[0085] As can be seen from the above, the intelligent scheduling system can automatically compare the video memory requirements of tasks with the video memory capacity of computing nodes, ensure that tasks are assigned to nodes with sufficient video memory resources for execution, avoid task failures or delays caused by insufficient video memory, and improve the accuracy and efficiency of resource scheduling. By filtering out computing nodes with insufficient video memory capacity, the system can select the most suitable node for executing deep learning tasks, that is, the first computing node. This not only reduces the invalid options in the scheduling process, but also optimizes the task execution environment and avoids resource waste. And ensuring the matching of the video memory resources of the computing node with the task requirements can improve the execution efficiency of deep learning tasks, reduce the waiting time during task execution, and at the same time improve the overall resource utilization rate and throughput of the system, providing strong support for the rapid completion of deep learning tasks.

[0086] In an alternative embodiment, Figure 2 is a structural diagram of an alternative task scheduling method according to an embodiment of the present application, as Figure 2 shown, including: a training task management module, a GPU cluster management module, a training task resource requirement evaluation module, and a training task request scheduling module. Among them, the training task management module is used to calculate task requests, including model structure, algorithm hyperparameters, input data sets, etc. Specifically, it is mainly responsible for managing the upload, parsing, and model information construction of deep learning training tasks. It receives the deep learning training task request uploaded by the user through the API interface, and constructs a user task request information model according to the task description file (such as a yaml file) and uploads it to the training task resource requirement evaluation module. Modeling the user task request information as where, represents the model used in the deep learning training task, respectively represent the batch size (B), the number of iterations (I), the neural network structure (D), and the parallel mode ( ); represents the data volume of the user task input data set. Among them, the standardization of the user's submitted task and its description information is crucial for the scheduling efficiency and accuracy of the system. The model used in the task can be selected from the packet in the artificial intelligence framework, and the user is also allowed to customize it, but it needs to conform to the artificial intelligence framework specification and the containerized Docker packaging standard to facilitate the system's rapid parsing and processing of the model, and reduce scheduling failures or resource waste caused by format errors or unclear descriptions.

[0087] The GPU cluster management module is used to obtain system monitoring metrics, such as graphics card configuration, video memory, network bandwidth, etc. Specifically, it is mainly responsible for the monitoring and management of cluster resources. By deploying monitoring software, it can detect the resource and link status of each node in real time, providing computing power awareness and network awareness parameters for the algorithm. The monitoring software is mainly responsible for metric collection, data persistence, and supports flexible metric queries using PromQL. In actual applications, the scraping addresses and related parameters of different types of Exporters need to be added to the monitoring software configuration file. For example, node_exporter is deployed to monitor the basic metrics of the server (CPU, memory, network interface, etc.), and snmp_exporter is deployed to collect network device metrics exposed through the SNMP (Simple Network Management Protocol) protocol, such as the bandwidth utilization rate and latency of routers and switches. Since distributed training involves high-performance computing and large-scale parameter transmission between cards, the GPU cluster management module models the resource information of the GPU-accelerated nodes managed by the cluster as and uploads it to the deep learning task resource requirement assessment module and the deep learning task request scheduling module. Among them, represents the resource information of the GPU-accelerated node, respectively represent the number of GPUs and the floating-point computing performance on the i-th computing node; represents the available video memory that can be allocated for each GPU graphics card in the current computing node i, that is, the actual remaining video memory minus the video memory reserved for other tasks; represents the PCI-E bus bandwidth within the computing node i, represents the utilization rate of the current GPU.

[0088] The training task resource requirement assessment module is used to parse task requirements, including computing volume, video memory occupancy, etc. Specifically, it is responsible for making a fine-grained estimate of the computing power and storage resource requirements of the deep learning training tasks uploaded by users according to the structure of the neural network model and algorithm hyperparameters, combined with the real-time status information of the computing power nodes in the cluster, and inputting the calculation results into the training task request scheduling module.

[0089] The training task request scheduling module will filter out the nodes with insufficient video memory resources based on the resource requirements of the training tasks and the real-time status information of the cluster. Then, it will further execute a scheduling algorithm based on the principle of minimizing the model training time to determine the target nodes, and at the same time complete the activation of computing power and network resources.

[0090] In an alternative embodiment, the Pytorch framework is adopted as the deep learning training task platform. First, the training task resource requirement evaluation module will call the API instruction torch.summary provided in the extension library of Pytorch, and input the task model structure and parameter configuration to obtain the basic information of the model structure. The function call format is: torch.summary(model, input_size), and the output results are: the number of parameters in each layer of the model , the size of the intermediate feature map , the computational cost of one forward propagation of the model , the video memory space requirement of the model parameters and the video memory space requirement of one forward propagation of the model , where represents which layer in the model, respectively represent the number of channels, height, and width of the intermediate feature map. According to the output of the extended torch.summary function, predict the total video memory space requirement, computational cost, and communication cost introduced by synchronization updates between multiple GPUs when executing user tasks under different placement strategies. The specific calculation methods are as follows:

[0091] 1) Computational cost of task requests:

[0092] Most of the deep learning models are concentrated in the convolutional layer and the fully connected layer. Only some simple floating-point calculations or no calculations are involved in other layers, and their computational costs are negligible compared with those of the convolutional layer and the fully connected layer. For the convolutional layer, convolution is implemented through a convolution kernel. Assume is the number of input channels, , and are the height, width, and number of channels of the output feature map, is the width of the convolution kernel (assuming the convolution kernel is symmetric). The computational cost of a feature point in the output feature map should include the number of multiplications and additions in the convolution operation (including the bias term), and the size of the output feature map is , so the total computational cost of a convolutional layer is . For the fully connected layer, assume the number of input neurons is , and the number of output neurons is . The fully connected layer can also be regarded as a special case of a convolutional layer with a convolution kernel size of . Then its number of input channels is equal to the number of input neurons , and the number of convolution kernels is equal to the number of output neurons .

[0093] Therefore, the total computational cost of a single-layer neural network is shown in formula (1):

[0094] (1)

[0095] Among them, is the batch size of the training task; is the computational load difference caused by different task types. For the training task , for the inference task .

[0096] The total computational load of the entire neural network model is equal to the sum of the computational loads of each layer multiplied by the total number of iterations . After adopting distributed parallelism, each GPU graphics card will evenly divide the total computational load. Let represent the computational load of one forward propagation of the model. Then the computational overhead on each GPU in the computing node i is as shown in formula (2):

[0097] (2)

[0098] Among them, represents the computational load when the model performs one forward propagation on the i-th computing node, represents which layer in the model, respectively represent the number of GPUs and the floating-point computing performance on the i-th computing node.

[0099] 2) Synchronous communication overhead of task requests:

[0100] Based on the distributed parallel computing of the Ring-AllReduce architecture, each GPU accelerator will be connected in a ring topology. Gradient aggregation is achieved by continuously sending and receiving data, and the model parameters are updated locally. The synchronous communication overhead of distributed tasks under the Ring-AllReduce mechanism is divided into intra-layer and inter-layer communication overheads, and is related to the parallelization method. The following will analyze them in turn:

[0101] For the intra-layer communication overhead, it is generated by the model performing one Ring-AllReduce operation for a single layer and is used for the aggregation of the model gradients. The total communication volume of a single card is , where is the number of GPUs.

[0102] For data parallelism, , is the number of parameters of this layer of the model ( );

[0103] For tensor parallelism, , is the product of the dimensions of the output results corresponding to a single sample.

[0104] Therefore, the intra-layer communication overhead of the computing node i corresponding to data parallelism is calculated as shown in Equation (3):

[0105] (3)

[0106] where represents the number of parameters of the layer of the model (i.e., ), is the total number of iterations.

[0107] The intra-layer communication overhead of the computing node i corresponding to tensor parallelism is calculated as shown in Equation (4):

[0108] (4)

[0109] where , are the dimensions of the feature map ( represents the number of channels of the feature map, represents the height and width of the feature map), i.e., and .

[0110] The inter-layer communication overhead comes from the data processing during the transfer of intermediate calculation results between layers. The data parallelism method does not involve the inter-layer communication cost, while the per-GPU communication volume of the inter-layer communication corresponding to the tensor parallelism method is , where . Therefore, the inter-layer communication overhead corresponding to tensor parallelism is calculated as shown in Equation (5):

[0111] (5)

[0112] By summing up the intra-layer and inter-layer communication overheads, the synchronous communication overheads corresponding to the data parallelism and tensor parallelism methods are obtained:

[0113] Then, the communication overhead of the computing node i corresponding to data parallelism is as shown in Equation (6):

[0114] (6)

[0115] The communication overhead of the computing node i corresponding to tensor parallelism is as shown in Equation (7):

[0116] (7)

[0117] 3) Memory occupancy requirement for task requests:

[0118] The video memory occupied by a neural network model mainly stores the model's own weight parameters, gradient values, and intermediate result outputs of the model. Common hierarchical types containing weight parameters include convolutional layers, fully connected layers, BatchNorm layers, Embedding layers, etc. Non-linear activation layers, pooling layers, Dropout layers, etc. only have hyperparameters and no weight parameters, so they do not occupy video memory. Suppose the total number of all parameters of a certain convolutional layer is , where the total number of convolutional weights is , and the total number of bias parameters is . The number of input channels of this layer is , the width of the convolutional kernel is , and the number is (the number of convolutional kernels is equal to the number of output channels of this layer ). According to the calculation characteristics of the convolutional layer, the total number of parameters of the entire convolutional layer can be calculated as shown in formulas (8)-(10):

[0119] (8)

[0120] (9)

[0121] (10)

[0122] Similarly, for a fully connected layer, the number of output neurons is , and the fully connected layer can also be regarded as a special case of a convolutional layer with a convolutional kernel size of . Its number of input channels is equal to the number of input neurons , and the number of convolutional kernels is equal to the number of output neurons , then the total number of parameters of a certain fully connected layer is as shown in formula (11):

[0123] (11)

[0124] Because the gradients and momenta during the model calculation process also need to be saved in the video memory, the choice of optimizer also has a certain impact on the video memory occupancy. For example, for the typical Stochastic Gradient Descent algorithm SGD, the amount of gradients to be saved is the same as the number of weight parameters, and the total video memory occupancy is twice the video memory space occupied by the parameters. For SGD with momentum (Momentum-SGD), the total video memory occupancy is 3 times the size of the weight parameters. For more complex optimizers, such as the Adam optimizer, the total video memory occupancy is 4 times the size of the weight parameters.

[0125] During the calculation of deep learning tasks, for each input data sample, intermediate calculation results between layers, i.e., intermediate feature maps, are often cached during the forward inference process. The size of the intermediate output tensor stored in the convolutional layer is equal to the size of the output feature map of that layer (where H is the height of the output feature map of that layer, W is the width, and C is the number of channels). The output tensor size of the fully connected layer is a one-dimensional vector, i.e., , and the number of output channels is equal to the number of output neurons , and it can be calculated by substituting into the above formula.

[0126] Let represent the model weight parameters, represent the video memory occupancy of one forward propagation of the model. Then the formula for calculating the total video memory occupancy of a single-layer network on computing node i is shown in formula (12):

[0127] (12) <00 series of calculations and explanations continue in a similar vein, with some lines being placeholders for specific equations or notations. For example,

[0128] and are likely part of an equation sequence. The text then goes on to discuss how different factors affect the video memory occupancy in different parallel strategies. In data parallelism, the complete model copy is loaded onto each GPU, and training data is split into batches for each GPU to process separately. Since each GPU stores a complete model copy, the space occupied by weight parameters in the video memory occupancy on each GPU Among them, represents the impact of optimizer selection on the video memory occupied by model parameters, represents the impact of different task types on the video memory occupied by intermediate feature maps. For training tasks , for inference tasks . And represents the storage space required for different numerical types. When using single-precision floating-point numbers , that is, each single-precision floating-point number occupies 4 bytes of video memory, represents the weight parameters of the i-th computing node of the model, represents the video memory occupancy of one forward propagation of the model on the i-th computing node.

[0129] In the data parallel strategy, a complete copy of the model is loaded onto each GPU, and the training data is split into small batches (batches) and distributed to each GPU for separate processing. Since each GPU stores a complete copy of the model, the space occupied by the weight parameters in the video memory occupancy on each GPU remains unchanged, while the space occupied by the intermediate feature maps is evenly divided; in the tensor parallel mode, the complete data set is loaded, but the model itself is split into multiple parts, and each part runs on a different GPU. Therefore, the space occupied by the weight parameters on each GPU is evenly divided, and the space occupied by the intermediate feature maps is the same as that of the original model.

[0130] Therefore, the video memory occupancy requirement for computing node i in the data parallel mode is shown in formula (13):

[0131] (13)

[0132] The memory occupation requirement of computing node i in the tensor parallel mode is shown in formula (14):

[0133] (14)

[0134] The evaluation of the resource requirements of a task can be summarized as . fun represents the resource requirement evaluation algorithm for training tasks. The input is the user request information service and the node resource information node, and the output is the synchronous communication time overhead, the calculation time overhead, and the required model memory space requirement on the single GPU under the corresponding placement strategy of the training task on this node.

[0135] The training task request scheduling module is mainly divided into two stages: screening and scheduling. In the screening stage, nodes with insufficient video memory resources or hardware resource failures will be filtered out according to the resource requirements of the training task and the cluster status information to ensure that the selected nodes have enough video memory to support the computing requirements of the model. When it is confirmed that the node meets the video memory requirements, the scheduling module will calculate the total time overhead for executing the user task on this node (node i) As shown in formula (15):

[0136] (15)

[0137] Among them, represents that the startup time of the worker node can also be ignored because the time required to start a container / virtual machine is very short (within a few seconds). Because the GPU utilization refers to the proportion of time that the GPU is actually used for executing computing tasks. And the computing and communication within the GPU node cannot be parallel, and the next layer of computing can only start after the synchronous communication of this layer is completed. So the total time overhead is the sum of the two. After traversing all optional nodes, the nodes are sorted in ascending order according to the time overhead, and the worker node with the shortest task execution time is selected for task scheduling. The computing power scheduling platform completes the activation of computing power resources and the connection of the network.

[0138] In an alternative embodiment, Figure 3 is a schematic diagram of an alternative task scheduling method according to an embodiment of the present application. As Figure 3 shown, first, the user submits a deep learning training task request (AI training task request) through the API interface. After receiving the user task request, the training task management module will parse the task description file submitted by the user, extract the key requirement indicators of the task, and construct a user task request information model according to the task description file Among them, the batch size, number of iterations, neural network structure, parallel mode, and the data volume of the input data set are uploaded to the training task resource requirement evaluation module. Then, the GPU cluster management module is responsible for the real-time monitoring and management of cluster resources. When a user task request arrives, the cluster management module uploads the monitored information on the node resource status and network status to the training task resource requirement evaluation module. Then, the scheduling engine sub-module receives the user task request information model uploaded by the training task management module, obtains the computing network metrics of the cluster and the real-time status information of the nodes from the GPU cluster management module, and then executes the resource requirement evaluation algorithm , estimates the resource requirements, calculates the synchronous communication time overhead of this user computing task under different placement strategies, calculates the time overhead, and the video memory space requirements of the model required on a single GPU graphics card. Then, according to the resource requirement evaluation results and the resource metric information uploaded by the GPU cluster management module, the training task request scheduling module filters out the nodes that do not meet the resource constraint conditions, calculates the expected total execution time corresponding to the remaining optional nodes, and selects the node with the smallest total execution time as the target node. After obtaining the target node, the cluster master node sends the task configuration file, etc. to the target node, and routes the task request to the selected target node to execute the deep learning training task on the target node. During the task execution process, the GPU cluster management module will monitor the task execution situation in real time, collect task execution data (such as resource usage, execution time, etc.), and feedback the task execution status and results to the user. Among them, the nodes are represented as Node1...Nodex.

[0139] The embodiment of the present application also provides a task scheduling device. It should be noted that the task scheduling device of the embodiment of the present application can be used to execute the task scheduling method provided by the embodiment of the present application. The following introduces the task scheduling device provided by the embodiment of the present application.

[0140] According to the embodiment of the present application, there is also provided a device for implementing the above task scheduling method, Figure 4 which is a schematic diagram of an optional task scheduling device according to the embodiment of the present application, as Figure 4 shown, including: a construction unit 401, a prediction unit 402, a determination unit 403, and an execution unit 404.

[0141] Optionally, a construction unit 401 is configured to construct a task model according to a target training task, where the task model is used to represent the attributes and resource requirements of the target training task; a prediction unit 402 is configured to predict performance results of the target training task when executed on N computing nodes based on the task model, obtaining N performance results, where N is an integer greater than 1, and the performance results include the execution efficiency and resource utilization rate of the target training task on the computing nodes; a determination unit 403 is configured to filter the N computing nodes according to the N performance results and network environment information to obtain S first computing nodes, and determine a target node from the S first computing nodes, where S is an integer greater than or equal to 1 and less than or equal to N, the first computing nodes are used to represent the computing nodes that meet the network environment information, and the target node is used to represent the first computing node that meets the preset requirements, where the preset requirements are used to constrain the duration of the target training task executed by the computing node; an execution unit 404 is configured to execute the target training task through the target node.

[0142] Optionally, the prediction unit 402 includes: a first prediction subunit, a second prediction subunit, a third prediction subunit, and a first determination subunit. Among them, the first prediction subunit is configured to predict a first value of the target training task on each computing node based on the task model, obtaining N first values, where the first value is used to represent the computing resource consumption of the graphics processor in each computing node when executing the target training task; the second prediction subunit is configured to predict a second value of the target training task on each computing node based on the task model, obtaining N second values, where the second value is used to represent the transmission time generated by data exchange between the computing nodes; the third prediction subunit is configured to predict a third value of the target training task on each computing node based on the task model, obtaining N third values, where the third value is used to represent the video memory space required by the graphics processor in each computing node when executing the target training task; the first determination subunit is configured to determine the N performance results according to the N first values, the N second values, and the N third values.

[0143] Optionally, the first prediction subunit includes: a first determination module, a second determination module, and a third determination module. Among them, the first determination module is configured to, when there are M network layers in the target neural network, determine the computation amount of each network layer according to the target information and batch size of each network layer in the M network layers, obtaining M computation amounts, where the target information is used to represent the parameter information required for each network layer to execute the target training task; the second determination module is configured to determine the computation amount of the target neural network according to the M computation amounts and the number of iterations; the third determination module is configured to determine the first value corresponding to each computing node according to the computation amount of the target neural network and the number of graphics processors corresponding to each computing node, obtaining N first values.

[0144] Optionally, the second prediction subunit includes: a fourth determination module and a fifth determination module. The fourth determination module is configured to determine, when the parallel mode is data parallel, a second value corresponding to each computing node according to the number of graphics processors corresponding to each computing node, the number of parameters of each of the M network layers, and the number of iterations, so as to obtain N second values. The fifth determination module is configured to determine, when the parallel mode is tensor parallel, a second value corresponding to each computing node according to the number of graphics processors corresponding to each computing node, the feature tensors output by each of the M network layers, the batch size, and the number of iterations, so as to obtain N second values.

[0145] Optionally, the third prediction subunit includes: a sixth determination module, a seventh determination module, an eighth determination module, and a ninth determination module. The sixth determination module is configured to determine an optimizer, where the optimizer is used to adjust and store the model parameters of the target neural network. The seventh determination module is configured to determine a first influence value of the optimizer on the video memory in the graphics processor, where the first influence value is used to quantify the video memory space occupied by the optimizer for storing data. The eighth determination module is configured to determine a second influence value according to the task type of the target training task, where the second influence value is used to quantify the occupancy of the intermediate feature maps output by any network layer in the target neural network in the video memory of the graphics processor. The ninth determination module is configured to determine a third value corresponding to each computing node according to the target information of any one of the M network layers, the first influence value, the second influence value, and the batch size, so as to obtain N third values.

[0146] Optionally, the determination unit 403 includes: a first filtering subunit, a first computing subunit, and a second determination subunit. The first filtering subunit is configured to filter the N computing nodes according to the N third values corresponding to the N computing nodes in the N performance results and the network environment information, so as to obtain S first computing nodes. The first computing subunit is configured to calculate a fourth value of each of the S first computing nodes according to the first value, the second value, and the average utilization rate of the graphics processor corresponding to each of the S first computing nodes, so as to obtain S fourth values, where the fourth value is used to characterize the total time consumption of each of the S first computing nodes for executing the target training task. The second determination subunit is configured to use the first computing nodes corresponding to the fourth values less than the preset threshold among the S fourth values as the target nodes.

[0147] Optionally, the determination unit 403 includes: a first detection subunit, a second filtering subunit, and a third determination subunit. Among them, the first detection subunit is configured to detect whether the third value corresponding to each computing node in the N performance results is greater than or equal to the video memory capacity corresponding to the computing node in the network environment information; the second filtering subunit is configured to filter out the computing nodes among the N third values corresponding to the N computing nodes that are greater than or equal to the video memory capacity corresponding to the computing node in the network environment information; the third determination subunit is configured to use the computing nodes among the N third values corresponding to the N computing nodes that are less than the video memory capacity corresponding to the computing node in the network environment information as the first computing nodes, to obtain S first computing nodes.

[0148] According to another aspect of the present application, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored executable program, and wherein, when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the above task scheduling method.

[0149] According to another aspect of the present application, there is also provided an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs, and wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above task scheduling method.

[0150] The serial numbers of the above embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0151] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0152] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in an electrical or other form.

[0153] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0154] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0155] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0156] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A task scheduling method, characterized in that, Including: Construct a task model according to a target training task, where the task model is used to characterize the attributes and resource requirements of the target training task; Based on the task model, predict the performance results when the target training task is executed on N computing nodes, and obtain N performance results, where N is an integer greater than 1, and the performance results include the execution efficiency and resource utilization rate of the target training task on the computing nodes; Filter the N computing nodes according to the N performance results and network environment information to obtain S first computing nodes, and determine a target node from the S first computing nodes, where S is an integer greater than or equal to 1 and less than or equal to N, the first computing node is used to characterize the computing node that meets the network environment information, and the target node is used to characterize the first computing node that meets the preset requirements, where the preset requirements are used to constrain the duration of the computing node to execute the target training task; Execute the target training task through the target node.

2. The task scheduling method according to claim 1, wherein Based on the task model, predict the performance results when the target training task is executed on N computing nodes, and obtain N performance results, including: Predict the first value of the target training task on each of the computing nodes based on the task model, and obtain N first values, where the first value is used to characterize the computing resource consumption of the graphics processor in each computing node when executing the target training task; Predict the second value of the target training task on each of the computing nodes based on the task model, and obtain N second values, where the second value is used to characterize the transmission time generated by data exchange between computing nodes; Predict the third value of the target training task on each of the computing nodes based on the task model, and obtain N third values, where the third value is used to characterize the video memory space required by the graphics processor in each computing node when executing the target training task; Determine the N performance results according to the N first values, the N second values, and the N third values.

3. The task scheduling method according to claim 2, wherein The task model includes a target neural network for training the target training task, the number of iterations, and the batch size, where the batch size is used to characterize the amount of data processed by the target neural network in each round of training. Predict the first value of the target training task on each of the computing nodes based on the task model, and obtain N first values, including: When there are M network layers in the target neural network, determine the amount of computation for each of the M network layers according to the target information of each of the M network layers and the batch size, and obtain M amounts of computation, where the target information is used to characterize the parameter information required for each network layer to execute the target training task; Determine the amount of computation of the target neural network according to the M amounts of computation and the number of iterations; Determine the first value corresponding to each computing node according to the amount of computation of the target neural network and the number of graphics processors corresponding to each computing node, and obtain N first values.

4. The task scheduling method according to claim 3, wherein The task model further includes a parallel mode, wherein the parallel mode is used to characterize the strategy of resource allocation and computing task division in the training process of the target neural network. Based on the task model, the second values of the target training task on each of the computing nodes are predicted, and N second values are obtained, including: When the parallel mode is data parallel, the second values corresponding to each computing node are determined according to the number of graphics processors corresponding to each computing node, the number of parameters of each of the M network layers, and the number of iterations, and N second values are obtained; When the parallel mode is tensor parallel, the second values corresponding to each computing node are determined according to the number of graphics processors corresponding to each computing node, the feature tensors output by each of the M network layers, the batch size, and the number of iterations, and N second values are obtained.

5. The task scheduling method according to claim 3, wherein Based on the task model, the third values of the target training task on each of the computing nodes are predicted, and N third values are obtained, including: Determine an optimizer, wherein the optimizer is used to adjust and store the model parameters of the target neural network; Determine the first influence value of the optimizer on the video memory in the graphics processor, wherein the first influence value is used to quantify the video memory space occupied by the optimizer to store data; Determine the second influence value according to the task type of the target training task, wherein the second influence value is used to quantify the occupancy of the intermediate feature maps output by any one of the network layers in the target neural network in the video memory of the graphics processor; According to the target information of any one of the M network layers, the first influence value, the second influence value, and the batch size, the third values corresponding to each computing node are determined, and N third values are obtained.

6. The task scheduling method according to claim 2, wherein Filter the N computing nodes according to the N performance results and network environment information to obtain S first computing nodes, and determine the target node from the S first computing nodes, including: Filter the N computing nodes according to the N third values corresponding to the N computing nodes in the N performance results and network environment information to obtain S first computing nodes; Calculate the fourth values of each of the S first computing nodes according to the first values, second values, and average utilization rate of the graphics processors corresponding to each of the S first computing nodes, and S fourth values are obtained, wherein the fourth value is used to characterize the total time consumption of each of the S first computing nodes to execute the target training task; Use the first computing nodes corresponding to the fourth values less than the preset threshold among the S fourth values as the target nodes.

7. The task scheduling method according to claim 6, wherein Filter the N computing nodes according to the N third values corresponding to the N computing nodes in the N performance results and network environment information to obtain S first computing nodes, including: Detect whether the third value corresponding to each of the N computing nodes in the N performance results is greater than or equal to the video memory capacity corresponding to the computing node in the network environment information; Filter the computing nodes among the N third values corresponding to the N computing nodes that are greater than or equal to the video memory capacity corresponding to the computing node in the network environment information; Use the computing nodes among the N third values corresponding to the N computing nodes that are less than the video memory capacity corresponding to the computing node in the network environment information as the first computing nodes, and obtain S first computing nodes.

8. A task scheduling device, characterized in that Comprising: A construction unit, configured to construct a task model according to a target training task, where the task model is used to characterize the attributes and resource requirements of the target training task; A prediction unit, configured to predict performance results when the target training task is executed on N computing nodes based on the task model, and obtain N performance results, where N is an integer greater than 1, and the performance results include the execution efficiency and resource utilization rate of the target training task on the computing node; A determination unit, configured to filter the N computing nodes according to the N performance results and network environment information to obtain S first computing nodes, and determine a target node from the S first computing nodes, where S is an integer greater than or equal to 1 and less than or equal to N, the first computing node is used to characterize the computing node that meets the network environment information, and the target node is used to characterize the first computing node that meets the preset requirements, where the preset requirements are used to constrain the duration of the computing node to execute the target training task; An execution unit, configured to execute the target training task through the target node.

9. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program runs, the device where the computer-readable storage medium is located executes the task scheduling method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, Comprising one or more processors and a memory, the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors execute the task scheduling method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Task allocation method, device and system

    CN116089051A

  • Heterogeneous parallel computing system and distributed training method

    CN118796402A

  • Distributed computing system training method and device, program product and medium

    CN119149254A

  • Heterogeneous computing system and training time consumption prediction method, equipment, medium and product thereof

    CN119204360A

  • Model training task scheduling method and apparatus, and electronic device

    WO2024041400A1

Cited By

  • Task scheduling method and device for training cluster, electronic equipment, computer readable storage medium and computer program product

    CN120872552A

  • Task scheduling method and device for training cluster, electronic device, computer readable storage medium and computer program product

    CN120872552B

  • Communication energy efficiency optimization method and device, equipment and storage medium

    CN121151301A