Machine learning training task communication scheduling method and device, equipment and storage medium

By constructing communication topology diagrams and computing diagrams in distributed deep learning training, refine communication operation workflows, and dynamically schedule communication resources, the problem of inefficient network communication is solved, and resource utilization and training efficiency are improved.

CN120256056APending Publication Date: 2025-07-04PENG CHENG LAB
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510375705.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the distributed deep learning training, the existing technology has problems such as low network communication efficiency, low resource utilization, insufficient overlap of computing and communication time, resulting in limited overall performance improvement.

Method used

By obtaining training task queues, extracting training features, building communication topology diagrams and calculation diagrams, refine the communication operation workflow, determine priority based on the urgency and calculation intensity, dynamically schedule communication resources, and optimize network resource allocation.

Benefits of technology

It improves the resource utilization and overall performance of distributed training, reduces communication latency and resource waste, and improves training efficiency. It is suitable for large-scale model training and hybrid deployment scenarios of heterogeneous tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256056A_ABST
    Figure CN120256056A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of machine learning, and discloses a machine learning training task communication scheduling method and device, equipment and a storage medium. The method comprises the steps of obtaining a to-be-trained task in a task queue; training feature extraction is carried out on the to-be-trained task, and a communication operation workflow and the priority of each communication operation in the workflow are obtained according to the training features obtained through analysis; and obtaining a communication scheduling scheme of the to-be-trained task according to the priority of each communication operation, and executing communication scheduling of the whole training task through the scheduling scheme. By means of the mode, accurate prediction of task computing power load and communication requirements and fine-grained modeling of communication scheduling between the nodes are achieved, scheduling planning can be conducted on the communication behavior of the level of communication between the nodes, efficient overlapping of calculation and communication is achieved, the training iteration time is shortened, and the scheduling efficiency is improved. The method is suitable for scenes of large-scale model training and multi-task parallel training, and has remarkable economic benefits and application values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning technology, and in particular to a method, apparatus, device and storage medium for communication scheduling of machine learning training tasks. Background Art

[0002] With the rapid development of artificial intelligence technology, deep learning has become its core driving force. However, the training requirements of large-scale data and complex models far exceed the capabilities of a single machine, and distributed training has become an inevitable choice. In distributed training, multiple computing nodes work together to synchronize gradients or parameters through data parallelism, but the problem of network communication efficiency has become a major bottleneck. Existing communication scheduling algorithms are often based on task priority, ignoring the different urgency of communication operations within the task, resulting in low utilization of communication resources.

[0003] In practical applications, especially in multi-task, multi-tenant distributed training environments, there is insufficient overlap between the computing and communication behaviors of computing nodes, and the computing and communication time of each node are not fully parallelized, affecting the overall performance. At the same time, the competition for communication resources in training tasks may cause communication delays and data packet loss, further reducing training efficiency. These problems mean that although computing resources are utilized to a certain extent, due to the underlying principles of training between models in machine learning, the low communication efficiency will in turn restrict the overall performance improvement of distributed training. The research on efficient communication scheduling algorithms has become the key to optimizing distributed training. At present, a more refined and dynamic communication scheduling algorithm is urgently needed to solve the above problems, so as to optimize the allocation of network resources, improve the utilization of GPUs, reduce training time, and improve the overall performance of distributed training systems.

[0004] Therefore, how to optimize network communication scheduling in distributed deep learning training to improve resource utilization and training efficiency is a technical problem that needs to be solved urgently in this field. Summary of the invention

[0005] The main purpose of this application is to provide a method, device, equipment and storage medium for communication scheduling of machine learning training tasks, aiming to solve the technical problem of how to optimize network communication scheduling in distributed deep learning training in the prior art to improve resource utilization and training efficiency.

[0006] To achieve the above objectives, the present application proposes a method for communication scheduling of machine learning training tasks, the method comprising: Get the tasks to be trained in the training task queue; Extracting training features of the task to be trained to obtain training features of the task to be trained; According to the training characteristics, obtaining the communication operation workflow of the task to be trained and the priority of each communication operation in the communication operation workflow; According to the priorities of the respective communication operations, obtain the communication scheduling scheme for the task to be trained, and perform the communication scheduling for the task to be trained according to the communication scheduling scheme.

[0007] In one embodiment, obtaining the priorities of the respective communication operations in the communication operation workflow according to the training features includes: According to the training features, obtain the communication topology graph and the computation graph of the task to be trained; According to the computation graph and the communication load rehearsal result, obtain the urgency score of the communication operation; According to the communication topology graph and the computing power load rehearsal result, obtain the computing intensity of the computing node corresponding to the communication operation; According to the urgency score of the communication operation and the computing intensity of the computing node corresponding to the communication operation, obtain the priorities of the respective communication operations in the communication operation workflow.

[0008] In one embodiment, obtaining the urgency score of the communication operation according to the computation graph and the communication load rehearsal result includes: According to the computation graph, determine the computing nodes corresponding to the respective communication operations in the communication operation workflow and the dependent nodes corresponding to the computing nodes; According to the computing node, the dependent node of the communication operation, and the communication load rehearsal result, obtain the communication transmission time and the communication node idle time of the communication operation; According to the communication transmission time and the communication node idle time of the communication operation, obtain the urgency score of the communication operation.

[0009] In one embodiment, obtaining the computing intensity of the computing node corresponding to the communication operation according to the communication topology graph and the computing power load rehearsal result includes: According to the communication topology graph, determine the computing node corresponding to the communication operation; According to the computing resource requirement information, obtain the computing workload of the computing node corresponding to the communication operation; According to the computing workload and the standard iteration duration, obtain the computing intensity of the computing node corresponding to the communication operation.

[0010] In one embodiment, extracting the training features of the task to be trained to obtain the training features of the task to be trained includes: According to the task to be trained, obtain the configuration information of the model to be trained, where the configuration information of the model to be trained includes model structure information, model parameter scale, and training data processing complexity; Determine the computing node allocation result and training parallel strategy according to the configuration information of the model to be trained; Based on the computing node allocation result and training parallel strategy, perform a model training preview on the task to be trained, and obtain a training preview result; Extract training features from the training preview result to obtain the training features of the task to be trained.

[0011] In one embodiment, the performing a model training preview on the task to be trained based on the computing node allocation result and training parallel strategy to obtain a training preview result includes: According to the computing node allocation result and training parallel strategy, decompose the task to be trained into sub-tasks according to computing nodes to obtain a computing sub-task queue; Based on the computing node parameters corresponding to each computing sub-task in the computing sub-task queue, perform a model training preview on the computing sub-task to obtain the computing workload, communication requirements, and communication data traffic of the computing sub-task on the corresponding computing node; Summarize the computing workload, communication requirements, and communication data traffic of each computing sub-task on the corresponding computing node to obtain the training preview result, where the training preview result includes a computing power load preview result and a communication load preview result.

[0012] In one embodiment, the obtaining the communication operation workflow of the task to be trained according to the training features includes: Obtain the computation graph of the task to be trained according to the training features; According to the computation graph, construct a communication topology graph for each computing node, where the communication topology graph includes at least all communication links required in a complete iteration process; According to the communication topology graph, obtain the communication operation workflow of the task to be trained, where the granularity of each communication operation in the communication operation workflow is the communication granularity between GPUs.

[0013] In addition, to achieve the above object, the present application also proposes a communication scheduling device for machine learning training tasks, and the communication scheduling device for machine learning training tasks includes: A task queue management module, configured to obtain a task to be trained in a training task queue; A training feature extraction module, configured to extract training features from the task to be trained to obtain the training features of the task to be trained; A communication operation planning module, configured to obtain the communication operation workflow of the task to be trained and the priority of each communication operation in the communication operation workflow according to the training features; A communication scheduling execution module, configured to obtain a communication scheduling plan for the to-be-trained task according to the priorities of the respective communication operations, and execute the communication scheduling for the to-be-trained task according to the communication scheduling plan.

[0014] In addition, to achieve the above object, the present application also provides a machine learning training task communication scheduling device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the machine learning training task communication scheduling method as described above.

[0015] In addition, to achieve the above object, the present application also provides a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the machine learning training task communication scheduling method as described above.

[0016] The technical solution of the present application includes: obtaining a to-be-trained task in a training task queue; extracting training features of the to-be-trained task to obtain training features of the to-be-trained task; obtaining a communication operation workflow of the to-be-trained task and priorities of the respective communication operations in the communication operation workflow according to the training features; obtaining a communication scheduling plan for the to-be-trained task according to the priorities of the respective communication operations, and executing the communication scheduling for the to-be-trained task according to the communication scheduling plan.

[0017] One or more technical solutions proposed by the present application have at least the following technical effects: By dynamically perceiving the computing characteristics and communication dependencies of training tasks, the efficiency and resource utilization rate of distributed training can be significantly improved. First, through training feature extraction and rehearsal, accurate prediction of task computing load and communication requirements is achieved, avoiding resource waste in static scheduling. Second, based on the fine-grained modeling of the communication topology graph and the computing graph, the selection of communication links can be optimized, reducing the delay and congestion of communication between nodes and improving the utilization rate of high-bandwidth links. In addition, by combining the urgency of communication operations with the computing intensity of computing nodes, a priority scheduling strategy is dynamically generated, realizing efficient overlap of computing and communication, and further compressing the training iteration time. Finally, this solution can effectively reduce the global synchronization delay in distributed training, improve the overall performance of training tasks, be applicable to scenarios of large-scale model training and heterogeneous task hybrid deployment, and has significant economic benefits and application value. Description of the Drawings

[0018] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It is a schematic flowchart provided for the first embodiment of the communication scheduling method for machine learning training tasks in the present application; Figure 2 It is a schematic flowchart provided for the second embodiment of the communication scheduling method for machine learning training tasks in the present application; Figure 3 It is a schematic flowchart provided for the third embodiment of the communication scheduling method for machine learning training tasks in the present application; Figure 4 It is a schematic module structure diagram of the communication scheduling device for machine learning training tasks in the embodiments of the present application; Figure 5 It is a schematic device structure diagram of the hardware operating environment involved in the communication scheduling method for machine learning training tasks in the embodiments of the present application.

[0021] The implementation, functional features, and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments

[0022] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0023] To better understand the technical solutions of the present application, the following will be described in detail in combination with the accompanying drawings of the specification and specific embodiments.

[0024] The main solution of the embodiments of the present application is: obtaining the to-be-trained tasks in the training task queue; extracting training features from the to-be-trained tasks to obtain the training features of the to-be-trained tasks; obtaining the communication operation workflow of the to-be-trained tasks and the priorities of each communication operation in the communication operation workflow according to the training features; obtaining the communication scheduling scheme of the to-be-trained tasks according to the priorities of each communication operation, and performing communication scheduling on the to-be-trained tasks according to the communication scheduling scheme.

[0025] In this embodiment, for the convenience of description, the following will be described with a machine learning training task communication scheduling device as the execution subject.

[0026] With the rapid development of artificial intelligence technology, deep learning, as the core driving force of artificial intelligence, plays a crucial role. By constructing complex neural network models, deep learning can automatically learn features and patterns from massive data, thus achieving efficient processing and accurate prediction of various tasks. However, with the continuous expansion of data scale and the continuous increase of model complexity, single-machine training has become difficult to meet the training needs of large-scale models, and distributed training has therefore become the inevitable choice for most training tasks.

[0027] In distributed training, multiple computing nodes work together to complete the model training task. Common parallel strategies include data parallelism, model parallelism, pipeline parallelism, and expert parallelism, etc. Taking data parallelism as an example, each node holds a copy of the model and performs calculations locally, and then synchronizes gradients or parameters among nodes through communication operations. This distributed training method can significantly accelerate the model training process, improve training efficiency, and make it possible to process large-scale data and complex models. However, in practical applications, especially in multi-tenant training clusters, when multiple training tasks are carried out simultaneously, the efficiency problem of network communication gradually emerges and becomes one of the key factors restricting the performance of distributed training.

[0028] In a multi-task and multi-tenant distributed training environment, the efficiency problem of network communication is mainly manifested in the following aspects. First, the scheduling granularity problem of communication operations. Most existing communication scheduling algorithms are scheduled based on task granularity, that is, communication tasks are arranged according to the priorities of tasks. This method ignores the different urgencies of each communication operation within the task, resulting in inflexible scheduling and unable to maximize resource utilization. Second, the existing scheduling schemes fail to fully consider the close relationship between computing and communication operations, resulting in less overlap between GPU computing time and communication time, thus affecting the overall performance. Finally, the network resource competition problem during multi-task parallelism makes it that in large-scale distributed training, although computing resources are utilized to a certain extent, it severely restricts the improvement of the overall training efficiency. There is an urgent need for a more refined and dynamic communication scheduling algorithm to solve the above problems, so as to optimize the allocation of network resources, improve the utilization rate of GPUs, reduce training time, and enhance the overall performance of the distributed training system.

[0029] This application provides a solution, aiming to solve the technical problem that there is a lack of a resource scheduling mechanism with a global perspective in the existing technology, which can dynamically perceive task characteristics, coordinate computing and communication resources, and thus achieve the optimization of overall performance.

[0030] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or a machine learning training task communication scheduling device that can implement the above functions. Hereinafter, taking the machine learning training task communication scheduling device as the execution subject as an example, this embodiment and the following embodiments will be described.

[0031] Based on this, the embodiments of the present application provide a machine learning training task communication scheduling method. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the machine learning training task communication scheduling method of the present application.

[0032] In this embodiment, the machine learning training task communication scheduling method includes steps S10 to S40: Step S10: Obtain the to-be-trained tasks in the training task queue.

[0033] It should be noted that the training task queue refers to an ordered container for storing to-be-executed training tasks. However, it should be emphasized that the tasks in the queue are usually complete model tasks, rather than hierarchical model tasks. A complete model task refers to a training task based on a complete iterative training step or a training task that completes multiple rounds of complete iterative training. These tasks are put into the queue as a whole and wait to be executed.

[0034] It can be understood that even if the execution order of the iterative tasks defined above changes, under the premise that the training resources remain unchanged, the overall time consumption of a single training task will not change. This is because the execution time of each task is mainly determined by its computational volume, communication volume, and resource allocation, and is not affected by other tasks. The advantage of doing this is that it simplifies the scheduling logic in multi-task concurrent training, avoids the additional computational and communication overhead introduced by consuming a large amount of computational resources for frequent hierarchical operation scheduling, and improves resource utilization and training efficiency.

[0035] Step S20: Extract training features from the to-be-trained tasks to obtain the training features of the to-be-trained tasks.

[0036] It should be noted that for any training task, since the training task itself is a multi-round complete iteration process, and the communication overhead and computing power consumption generated in each iteration process are relatively fixed. Among them, the consumption of computing power resources is mainly related to the scale of the model structure, the scale of the input data, and the type of computing operations. The overhead of communication operations is mainly related to the communication frequency, communication architecture, training parallel strategy, and the amount of communication parameters. Specifically, the training features are multi-dimensional descriptions of the training task to be processed, covering multiple aspects such as the model structure, data features, computing power consumption, communication duration features, hardware and system features, and training strategy features.

[0037] It can be understood that the scale of the model structure refers to the number of layers of the model, the number of neurons in each layer, and these parameter quantities directly affect the amount of computation in each layer and the overall amount of computation. The batch size and feature dimension of the input data will also affect the amount of computation. In addition, the type of operations in the model (such as convolution, matrix multiplication, activation function) also has a significant impact on the amount of computation. For example, the computational complexity of convolution operations is usually higher than that of other types of computational operations. On the other hand, the scale of the model parameters (such as the size of the weight matrix) directly affects the amount of data in the communication process. For example, large neural networks usually have a larger number of parameters and higher communication overhead. Secondly, in synchronous training, gradients need to be synchronized after each iteration, and the communication frequency is relatively high. However, in asynchronous training, the communication frequency is relatively low, but it may introduce additional delays. In addition, the network topology structure of the distributed system (such as ring, star, tree) will affect the communication efficiency. Finally, choosing different parallel strategies for different types of models will also result in different communication patterns and data amounts in each iteration. In summary, through the extraction of training features, the computing power consumption and communication overhead of the training task to be processed can be comprehensively evaluated, so as to provide a scientific basis for subsequent task scheduling and resource allocation.

[0038] It should be understood that for the vast majority of models, during the training process, the structure of the model (such as the number of layers, the number of neurons in each layer, the connection method, etc.) is usually fixed. This makes the amount of computation in each iteration (such as the scale of matrix multiplication, the number of calculations of activation functions) remain consistent. In addition, since the allocation of computing nodes is directly configured by the manager of the distributed training cluster, on the premise that the task content is determined, the allocation results of computing resources and communication resources can be obtained in advance according to the rules and regulations, which means that the network bandwidth, delay, and topology structure between the allocated computing nodes are usually fixed, which will make the communication overhead remain consistent in each iteration. Therefore, for any training task, the communication overhead and computing power consumption generated in each iteration process are relatively fixed.

[0039] In a feasible implementation manner, step S20 may include steps A11 to A13: Step A11: Obtain the configuration information of the model to be trained according to the task to be trained.

[0040] It should be noted that the configuration information of the model to be trained is obtained through comprehensive analysis of the task objective and training data, and should include the core attributes of the model, data processing requirements, and its potential resource consumption characteristics. These information are the basis for subsequent resource allocation and strategy formulation.

[0041] It can be understood that the model structure information includes the number of layers of the model, the type of each layer (such as convolutional layer, attention layer), activation function, connection method (such as residual connection), etc. These parameters directly affect the amount of computation and the feasibility of parallelization. The scale of model parameters refers to the quantization level of the number of parameters set in this field (such as billions of parameters), which represents the video memory occupancy of the model, the amount of communication data, the allocation result of computing nodes, and the selection result of parallel strategies during training. The complexity of training data processing includes data loading speed and preprocessing time consumption, and these parameters will affect the efficiency of the task's computing pipeline.

[0042] It should be understood that the configuration information of the model to be trained includes at least model structure information, model parameter scale, and training data processing complexity. These contents, as the basic attributes of the training task, combined with the computing power resource scheduling rules formulated by the distributed training cluster manager, can obtain the allocation result of computing nodes and the training parallel strategy allocated for this task.

[0043] Step A12: Determine the allocation result of computing nodes and the training parallel strategy according to the configuration information of the model to be trained.

[0044] It should be noted that node allocation and parallel strategy need to be jointly optimized according to the computing characteristics (computing power requirements) of the model, communication requirements (number of parameters and synchronization frequency), and hardware resources (number of GPUs, network topology). The goal is to maximize resource utilization and training speed.

[0045] It can be understood that the training parallel strategy includes but is not limited to: data parallelism, which is suitable for tasks with a small number of parameters but a large amount of data, and gradient synchronization is required regularly; model parallelism, which is suitable for scenarios where the single-card video memory of ultra-large-scale models (such as large language models) is insufficient, and the model needs to be split into different nodes; pipeline parallelism, which is for models with deep computational graph dependencies (such as Transformer), and is executed in stages in a pipeline to improve throughput; hybrid parallelism, which refers to mixing multiple parallel methods to balance communication overhead and computing efficiency as needed.

[0046] It should be understood that the choice of parallel strategy needs to balance communication cost and computational efficiency. For example, when bandwidth is limited, gradient synchronization in data parallelism may become a bottleneck; while when video memory is limited, hybrid parallelism (such as ZeRO optimization) may be the only feasible solution. Here, it should be noted that ZeRO optimization refers to offloading the optimizer state and gradients to CPU memory or NVMe storage, significantly releasing GPU video memory, which is applicable to scenarios where the video memory of a single card is extremely limited (such as training medium-sized models with consumer-grade GPUs). In theory, it is possible to train a model with trillions of parameters on a single RTX 3090 (with 24GB of video memory).

[0047] Step A13: Based on the computing node allocation result and the training parallel strategy, perform model rehearsal on the to-be-trained task to obtain a training rehearsal result.

[0048] It should be noted that the node allocation result refers to the number of computing nodes, the GPU model and video memory capacity of each node; the training parallel strategy refers to the parallel method and parameter distribution rule assigned to this training task; the communication network configuration refers to the bandwidth (Gbps), latency (ms), and topology (such as the All-Reduce cluster structure) between nodes.

[0049] It can be understood that the rehearsal content mainly includes the estimation of the computational time related to training in each connection layer of this training task. For example, use a representative data batch to simulate the forward propagation and backward propagation time consumption, and then perform simulated communication operations to obtain the simulated communication overhead generated. Combine the communication network performance (bandwidth, latency) and synchronization frequency (such as the number of All-Reduce times) to estimate the total communication time, so as to be able to estimate the training time consumption of each layer and the communication time consumption for this layer to complete communication operations during the entire iterative training process.

[0050] It should be understood that through this process, the computational time and communication time consumption of each level in the distributed training task can be accurately estimated, forming a chain-dependent structure, which is used to locate performance bottlenecks (such as excessive computational or communication latency in a specific layer) and for subsequent communication optimization resource allocation, avoiding resource waste or failure caused by insufficient planning during actual training.

[0051] In a feasible implementation, step A13 may include steps B131 to B133: Step B131: According to the computing node allocation result and the training parallel strategy, decompose the to-be-trained task according to computing nodes to obtain a computing sub-task queue.

[0052] It should be noted that for some training tasks with multiple rounds of iteration, the specific decomposition logic is to decompose the tasks according to the allocated computing nodes and task rounds. For example, in a specific training scenario, assuming that there are a total of N layers in the model and M rounds of training are required, and the system allocates P computing nodes to this node. On the premise of taking into account computing power and communication requirements, the total task volume allocated to each node is: average computing volume per training layer * N * M / P. Since the video memory of the computing node restricts the single task, for any computing unit, the total number of tasks in its sub-task queue needs to be further divided according to the video memory bottleneck and the number of layers. Specifically, assuming that due to the existence of the video memory bottleneck, each node can perform at most 3 layers of training at a time, then P nodes will be allocated with 3 layers as a sub-task in one iteration of training, resulting in a total of N * M / P * 3 sub-tasks, that is, each sub-task is the training of a certain number of consecutive hidden layers.

[0053] Step B132: Based on the computing node parameters corresponding to each computing sub-task in the computing sub-task queue, perform a pre-training of the model for the computing sub-task to obtain the computing workload, communication requirements, and communication data traffic of the computing sub-task on the corresponding computing node.

[0054] It should be noted that since the computing power of the computing node is related to the hardware devices used (such as GPU model, number of CPU cores, memory size), the computing power can be quantified. According to the specific sub-task content allocated, both the computing duration required to complete the sub-task can be roughly estimated, and the amount of data to be transmitted for the communication operations to be performed can be predicted, that is, the communication requirements and communication data traffic required after this computing operation. In addition, it should be emphasized that if a single GPU is allocated to train multiple consecutive hidden layers, the data transmission between these hidden layers is equivalent to the data transmission within the node and usually does not pass through the communication channels within the cluster, so it does not occupy the public communication resources. At this time, the communication requirements can be regarded as only including the communication requirements between GPUs, that is, the communication requirements and communication data traffic only count the requirements and demand traffic generated by the communication behavior between nodes.

[0055] It can be understood that the computing sub-task queue contains a task splitting list (task type, parallel mode, parameter distribution rule) generated by the task scheduler. The computing node parameters include hardware specifications such as GPU model (computing power size, video memory capacity), number of CPU cores (parallel thread number), and memory bandwidth. In addition, it also includes network performance such as the bandwidth and latency of the inter-node connection.

[0056] Step B133: Summarize the computing workload, communication requirements, and communication data traffic on the computing nodes corresponding to the respective computing subtasks to obtain the training rehearsal result, where the training rehearsal result includes a computing power load rehearsal result and a communication load rehearsal result.

[0057] It should be noted that by summarizing the computing and communication load data of each subtask in a distributed task to generate a computing power load rehearsal result and a communication load rehearsal result from a global perspective, the computing saturation, communication bottleneck, and critical path of the overall task can be clarified, providing an accurate basis for communication operation scheduling and parallel strategy optimization.

[0058] It can be understood that through the dependency relationship between each computing node and the rehearsed computing time consumption, the actual communication resource consumption change curve graph during the iteration process can be obtained. Based on this graph, the location of the communication bottleneck can be determined, and in the subsequent process, the "peak shaving and valley filling" strategy can be adopted to optimize the overall communication scheduling. For example, for some communication operations, the entire communication operation can be split twice, divided into communication operations that are dependent on other nodes and communication operations without dependency relationships. Especially in the backpropagation process of data parallelism, when synchronizing data from the subsequent hidden layer to the previous hidden layer, it should be ensured that the synchronized data is first transmitted to the nearest previous hidden layer so that it can start the backpropagation calculation of this layer as early as possible, while the synchronization process of the more distant end can be appropriately slowed down according to the real-time communication pressure.

[0059] Step A14: Extract training features from the training rehearsal result to obtain the training features of the task to be trained.

[0060] It should be noted that based on the training rehearsal result (computing power load and communication load data) generated in the previous steps, the core features of the task to be trained are extracted to quantify indicators such as its global computing power load distribution, communication link hotspots, critical path time consumption, and video memory occupancy curve, and associated with the model type and parallel strategy, providing a precise dimension for subsequent resource elastic scheduling and algorithm optimization.

[0061] It can be understood that the training features include features in multiple different aspects. For example: computing power allocation features, which are used to describe the dispersion degree of the computing time consumption differences of each node after task splitting. The specific quantification indicators include the load difference coefficient and the resource elasticity index; communication bottleneck intensity features, which are used to quantify the proportion of communication time consumption in the total time consumption of a single iteration. The specific quantification indicators include the link pressure coefficient and the bandwidth saturation; critical path concentration, which is used to describe the proportion of the critical path time consumption in the overall time consumption. According to the actual results, it is judged whether the critical path needs to be split into multi-node parallelism to reduce the communication pressure.

[0062] It should be understood that the training features determine the sensitivity of the task to the hardware topology. Through rehearsal, potential risks can be identified before actual training, and preventive scheduling of communication or computing resources can be carried out in advance. For example, tasks with a high critical path concentration need to be preferentially deployed in a low-latency network cluster, while tasks with high video memory fluctuations are more suitable to be paired with heterogeneous nodes with large memory CPUs.

[0063] Step S30: According to the training features, obtain the communication operation workflow of the to-be-trained task and the priorities of each communication operation in the communication operation workflow.

[0064] It should be noted that the communication operation workflow refers to an ordered execution sequence (workflow) of communication operations at the task level. Initially, the workflow is only arranged according to the task execution order, and dynamic priorities are assigned to it later. The assignment is carried out on the premise of minimizing communication waiting to avoid resource contention and achieve high-throughput distributed training.

[0065] It can be understood that each communication operation in the communication operation workflow should be a communication operation of the same fine-grained level. Assuming that the fine-grained level is defined as a single GPU-to-GPU communication as one communication operation, the communication operations in the entire workflow should be refined to consider each GPU-to-GPU communication in every iteration of the entire training task as one communication operation. If there are N GPU-to-GPU operations in one iteration process and the training task iterates M times, then there will be N*M communication operations in the communication operation workflow. Similarly, when the fine-grained level of communication scheduling changes, for example, when the communication operation is divided with inter-cluster communication as the fine-grained level (one cluster contains multiple computing nodes), then the communication operations between clusters are counted in the workflow and arranged in order as the communication operation workflow.

[0066] Step S40: According to the priorities of each communication operation, obtain the communication scheduling scheme for the to-be-trained task, and execute the communication scheduling for the to-be-trained task according to the communication scheduling scheme.

[0067] It should be noted that by comprehensively considering the priorities of each communication operation, a reasonable communication scheduling scheme is formulated to optimize the allocation and use of communication resources, ensure that high-priority communication operations can be executed in a timely manner, and thus improve the efficiency and performance of the entire training task.

[0068] It can be understood that, in combination with the preview data obtained in the foregoing process, various methods can be adopted for determining the priority, such as the weighted average method, the analytic hierarchy process, etc. Select a suitable method according to the specific application scenario and requirements, and then generate a communication scheduling scheme according to the priorities of each communication operation. This scheme clarifies the execution order, time arrangement, resource allocation, etc. of each communication operation in the workflow. For example, the execution of communication operations can be arranged in descending order of priority, and more communication resources and bandwidth can be allocated to high-priority operations.

[0069] It should be understood that during the multi-round iterative training process, in each iteration, the computational workload generated by each computing node is roughly the same as the communication overhead, and there are only minor fluctuations in the transmission volume. Therefore, the actual communication scheduling scheme will be repeatedly executed during each iteration process, which can simplify the resource overhead in communication scheduling. According to the generated communication scheduling scheme, the scheduling of communication operations is actually executed. During the execution process, it is necessary to monitor the execution status of communication operations and resource usage in real time, and promptly handle possible communication conflicts, delays and other problems to ensure the effective execution of the scheduling scheme.

[0070] In this embodiment, the to-be-trained tasks are obtained from the training task queue. These tasks are complete model tasks and include multi-round iterative training steps. Through training feature extraction, multi-dimensional features of the tasks are obtained, including model structure, data features, computing power consumption, etc. Based on these features, a communication topology graph and a computing graph are constructed, a communication operation workflow with a fine granularity of GPU-to-GPU communication is generated, and the order and dependency relationship of each communication operation are clarified. At the same time, in combination with the preview results of communication load and the preview results of computing power load, the urgency scores of each communication operation and the computing intensity of the corresponding computing nodes are calculated. Finally, the priorities of communication operations are determined according to these metrics, a communication scheduling scheme is formulated, and communication scheduling is executed according to this scheme to optimize the use of communication resources and improve training efficiency.

[0071] In summary, through refined training feature extraction and preview, this technical solution can accurately evaluate the urgency of each communication operation and the load conditions of computing nodes, so as to formulate a reasonable communication scheduling scheme. This not only helps to optimize the allocation of communication resources, reduce communication delays, but also improves the utilization rate of computing resources and avoids resource waste. In multi-task concurrent training, this scheme simplifies the scheduling logic and reduces the additional computing and communication overhead generated by frequent scheduling hierarchical operations, further improving the overall performance and resource utilization rate of the system.

[0072] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2, in step S30 of the machine learning training task communication scheduling method, steps S301 to S304 are further included: Step S301: Obtain the communication topology graph and computation graph of the task to be trained according to the training features.

[0073] It should be noted that the communication topology graph describes the physical or logical connection relationship between computing nodes, including attributes such as link bandwidth, latency, and protocol, covering all communication links that may be activated within a complete iteration cycle. Herein, a node refers to a physical device (such as a GPU or CPU cluster) or a logical unit (such as a virtual computing group), an edge represents a communication link, and the maximum bandwidth, current load rate, and historical congestion times can be marked in the graph.

[0074] It can be understood that the computation graph describes the execution logic and dependency relationship of the task, usually represented by a directed acyclic graph (DAG). The nodes are subtasks (such as matrix multiplication and gradient calculation), and the edges represent dependency relationships (such as having to wait for the forward propagation of a certain layer to complete before starting the backpropagation).

[0075] In a feasible implementation manner, after step S301, steps C11 to C13 are further included: Step C11: Obtain the computation graph of the task to be trained according to the training features.

[0076] It should be noted that according to the model structure, parallel strategy, and training parameters (batch size, optimizer type, gradient aggregation frequency) in the training features, the computation graph of the task to be trained can be obtained.

[0077] It can be understood that first, the task is split and the computing node devices are allocated according to the parallel strategy and the layer structure of the model, then the dependency relationship between nodes is constructed according to the dependency relationship (forward propagation and backpropagation) between the layer structures, and the communication requirements are marked in the graph, thus constructing the computation graph of the task to be trained.

[0078] Step C12: Construct the communication topology graph of each computing node according to the computation graph.

[0079] It should be noted that by using the computation graph and the hardware configuration table of each node (GPU connection method, bandwidth, latency) as inputs for physical link discovery, the physical link links between the computing nodes corresponding to each GPU can be obtained, and at the same time, the connection method is marked, such as direct NVLink, PCIe switch path, or network path.

[0080] It should be understood that the communication topology diagram contains at least all the communication links required for a complete iteration process, so that the scheduler can make decisions within the physically feasible range (such as avoiding assigning tasks to unconnected devices). In addition, the scheduler can only make globally optimal decisions when all links are known. If there is no complete link information when scheduling multiple communication operations, it may not be possible to effectively avoid bandwidth contention or find the optimal path, resulting in insufficient resource utilization.

[0081] Step C13: According to the communication topology diagram, obtain the communication operation workflow of the task to be trained.

[0082] It should be noted that the purpose of this step is to generate a fine-grained communication operation sequence for inter-GPU communication based on the communication topology graph. Expand the N inter-GPU operations in one iteration into N communication tasks, and obtain N×M operations for M iterations. The order of operations is inherited from the dependencies of the tasks in the computation graph, thus forming a complete workflow. Assume that there are two inter-GPU communication operations in one iteration, namely, GPU0 sends data to GPU1 and GPU1 sends data to GPU2, and these two operations have a dependency relationship (that is, GPU0 must send data to GPU1 before GPU1 sends data to GPU2). According to the dependency relationship of the communication topology graph and the computation graph, these two operations are sequentially expanded into two communication tasks. For M iterations, 2×M communication operations will be generated, forming a communication operation workflow.

[0083] It can be understood that the communication operation workflow is a concretization and refinement of the communication part in the computation graph. It is closely related to the computation graph but different from it. The computation graph mainly describes the logical structure and dependencies of the computation task, while the communication operation workflow focuses on the execution details and sequence of the communication part. The two together guide the execution of the training task.

[0084] It should be understood that in this embodiment, the fine-grainedness of each communication operation in the communication operation workflow is preferably set to the fine-grainedness of inter-GPU communication. This is because the GPU is the smallest hardware unit of the computing unit, and the fine-grained division of inter-GPU communication operations can more accurately control and optimize the communication process, improve communication efficiency, reduce communication delays, and make full use of the parallelism of hardware resources, thereby improving the performance of the entire training task. In addition, as the model scale continues to expand and the computational complexity increases, the optimization of inter-GPU communication helps to better manage and schedule communication tasks, avoid the emergence of communication bottlenecks, and ensure the efficiency and stability of the training process.

[0085] Step S302: Obtain an urgency score of the communication operation according to the calculation graph and the communication load preview result.

[0086] It should be noted that by analyzing the task dependency relationships and communication load rehearsal results in the computational graph, the urgency of communication operations is quantified, providing a basis for subsequent communication scheduling.

[0087] It can be understood that by determining the positions and dependency relationships of each communication operation in the workflow within the computational graph, understanding its impact on subsequent tasks, and then using data such as the communication transmission time obtained from the rehearsal, the execution efficiency of the communication operation is evaluated. Finally, according to the characteristics and load conditions of the communication operation, an urgency score is calculated according to a certain formula or method, reflecting its priority in the overall communication task.

[0088] It should be understood that the urgency score reflects the importance and urgency of the communication operation in the overall task. The higher the score, the more this operation needs to be processed preferentially to avoid causing delays to subsequent tasks.

[0089] Step S303: According to the communication topology graph and the computational power load rehearsal results, obtain the computational intensity of the computing nodes corresponding to the communication operation.

[0090] It should be noted that by analyzing the communication topology graph and the computational power load rehearsal results, the computational intensity of the computing nodes corresponding to the communication operation is evaluated, providing a reference for resource allocation and load balancing.

[0091] It can be understood that the rehearsal results should include data in multiple aspects such as computational workload, processing time, and resource utilization rate to comprehensively reflect the load conditions of the computing nodes. Based on the foregoing, the computing nodes involved in the communication operation and their connection relationships can be determined through the communication topology graph, and then the load conditions of the computing nodes are evaluated using data such as the computational workload and processing time obtained from the rehearsal. Finally, according to the computational workload and the minimum iteration time of a single training iteration, the computational intensity of the computing nodes corresponding to the communication operation is calculated.

[0092] It should be understood that the computational intensity reflects the computational load conditions of the computing nodes per unit time. A higher computational intensity means that the node needs to process more computational tasks and may require more resource support.

[0093] Step S304: According to the urgency score of the communication operation and the computational intensity of the computing nodes corresponding to the communication operation, obtain the priorities of each communication operation in the communication operation workflow.

[0094] It should be noted that the urgency score can only quantify the urgency of each communication operation. Relying solely on this indicator for priority allocation is insufficient because it does not consider the computational characteristics of the job. By combining the computational intensity, that is, the ratio of the computational workload of the job to the minimum iteration time, the impacts of both the computational operation and the communication operation on the final scheduling scheme can be balanced.

[0095] It is understandable that, by comprehensively considering the urgency of communication operations and the computing intensity of computing nodes, the priorities of each operation in the communication operation workflow can be determined. Specifically, taking the calculation result of the product of the two as the index data for priority arrangement, the priority arrangement order of all communication operations can be obtained, providing the final decision basis for communication scheduling.

[0096] In this embodiment, through training feature extraction, multi-dimensional features of the task are obtained, including model structure, data features, computing power consumption, etc. Based on these features, a communication topology graph and a computing graph are constructed, a communication operation workflow with a fine granularity of inter-GPU communication is generated, and the order and dependency relationship of each communication operation are clarified. At the same time, combining the results of communication load rehearsal and computing power load rehearsal, the urgency scores of each communication operation and the computing intensity of the corresponding computing nodes are calculated. Finally, based on these metrics, the priorities of each operation in the communication operation workflow are determined, a communication scheduling scheme is formulated, and communication scheduling is performed according to this scheme to optimize the use of communication resources and improve training efficiency.

[0097] In summary, through refined training feature extraction and rehearsal, this technical solution can accurately evaluate the urgency of each communication operation and the load situation of computing nodes, so as to formulate a reasonable communication scheduling scheme. In multi-task concurrent training, this scheme simplifies the scheduling logic, reduces the additional computing and communication overhead caused by frequent scheduling hierarchical operations, and further improves the overall performance and resource utilization rate of the system.

[0098] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar content as in the above-described second embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , in step S302 of the communication scheduling method for machine learning training tasks, it includes steps D11 to D13: Step D11: According to the computing graph, determine the computing nodes corresponding to each communication operation in the communication operation workflow and the dependency nodes corresponding to the computing nodes.

[0099] It is understandable that for each communication operation in the communication operation workflow, according to the node definition and connection relationship in the computing graph, the computing nodes corresponding to this communication operation are determined. For example, if a certain communication operation is to send data from node A to node B, then node A and node B are the computing nodes corresponding to this communication operation.

[0100] It should be understood that through the computational graph, the dependency relationships between various computational nodes can be analyzed to determine the other nodes on which each computational node depends. For example, if the computation of node B may depend on the output of node A, then node A is the dependency node of node B. Such dependency relationships will affect the execution order of communication operations and the direction of data flow. Generally speaking, after the task allocation for computational nodes is completed in the pre-step, each computational node has only a uniquely corresponding dependency node, and only a small number of nodes will have overlapping identities (input layer nodes and output layer nodes).

[0101] Step D12: Obtain the communication transmission time and communication node idle time of the communication operation based on the computational nodes, dependency nodes, and communication load rehearsal results of the communication operation.

[0102] It should be noted that the communication transmission time refers to the transmission duration of a certain communication operation from the start of execution to the complete receipt of the communication content by the corresponding node. On this basis, the communication node idle time refers to the idle time of the communication node after the completion of this communication operation until the start of execution of the dependent computational node.

[0103] It can be understood that, for example, there is a communication operation to send 100MB of data from node A to node B, the bandwidth of the communication link is 1000MB / s, and the delay is 1ms. According to these communication load rehearsal results, the communication transmission time is approximately 100MB / 1000MB / s = 0.1s. If the dependent computational node B needs to wait for 2s to start the next computational task after receiving the data, then the idle time of communication node B is 2s - 0.1s = 1.9s.

[0104] Step D13: Obtain the urgency score of the communication operation based on the communication transmission time and communication node idle time of the communication operation.

[0105] It should be noted that the urgency score is the ratio of the transmission time t c of the communication operation to the sum of this transmission time t c and the idle time t i The formula is as follows:

[0106] It can be understood that this score reflects the relative urgency of the communication operation. The higher the score, the more the communication operation needs to be executed preferentially.

[0107] It should be understood that the urgency score is a relative indicator used to measure the priority of a communication operation in the overall communication scheduling. It comprehensively considers the time required for the communication operation itself and the idle time of the communication node after completing the operation, thus more comprehensively reflecting the impact of the operation on the entire system.

[0108] In a feasible implementation manner, step S303 includes steps E11 to E13: Step E11: Determine the computing node corresponding to the communication operation according to the communication topology graph.

[0109] It should be noted that the nodes and links in the communication topology graph, combined with the definition and requirements of the communication operation, determine which computing nodes are involved in the communication operation. For example, in point-to-point communication, determine the sending node and the receiving node; in collective communication, determine all participating nodes.

[0110] It can be understood that assume that the communication topology graph contains four computing nodes A, B, C, and D, and there are communication links A - B, B - C, C - D, and D - A. For a communication operation that sends data from node A to node B and then from node B to node C, according to the communication topology graph, it can be determined that the computing nodes involved in this communication operation are A, B, and C, and the corresponding dependent nodes are B, C, and D respectively.

[0111] It should be understood that the communication operation only considers point-to-point communication operations, that is, communication operations between nodes, and broadcast communication operations should be regarded as multiple single operations in which a single node sends to other nodes sequentially or simultaneously.

[0112] Step E12: Obtain the computing workload of the computing node corresponding to the communication operation according to the computing resource requirement information.

[0113] It should be noted that through the rehearsal process in the foregoing steps, an approximate value of the computing workload of each computing node when performing the corresponding task can be obtained, and this value is used as a premise to provide a basis for subsequent communication resource allocation and load balancing.

[0114] Step E13: Obtain the computing intensity of the computing node corresponding to the communication operation according to the computing workload and the standard iteration duration.

[0115] It should be noted that the computing intensity is the workload of the computing task corresponding to this communication operation and the minimum iteration time of a single training iteration. Among them, the minimum iteration time refers to the shortest time required to complete one iteration in the entire training task under the premise of no communication competition, corresponding to the standard iteration duration here, which refers to the average time required to complete one iteration under normal circumstances, and it can be determined according to historical data, empirical models, or system design requirements.

[0116] It is understandable that the computing intensity I i has the following calculation formula, where Wi represents the computing workload and di represents the standard iteration duration.

[0117]

[0118] It should be understood that the computing intensity refers to the ratio of the workload of the computing tasks of the corresponding computing nodes to the minimum iteration time of a single training iteration in a communication operation. It reflects the computing load of the computing nodes per unit time and is an indicator for evaluating the computing resource requirements of the nodes.

[0119] In this embodiment, by comparing the transmission time and the idle time of each communication operation, the urgency of the communication operation can be quantified. The urgency score mainly reflects the priority of the communication operation from the perspectives of the communication transmission time and the idle time, while the computing intensity reflects the load situation of the computing nodes from the perspective of the computing resource requirements. Combining the two can comprehensively evaluate the priority of the communication operation, taking into account both the timeliness of communication and the bearing capacity of the computing resources.

[0120] In summary, by comprehensively considering the urgency score and the computing intensity, this technical solution can more accurately determine the execution order and resource allocation of communication operations, thereby reducing the waiting time of communication operations. Communication operations with high urgency can be processed in a timely manner, and nodes with high computing intensity can obtain sufficient resource support, making the task execution of the entire system more smooth and improving the response speed and throughput of the system.

[0121] It should be noted that the above examples are only for understanding this application and do not constitute a limitation to the communication scheduling method for machine learning training tasks of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0122] This application also provides a communication scheduling device for machine learning training tasks. Please refer to Figure 4 The communication scheduling device for machine learning training tasks includes: A task queue management module 10, configured to obtain the to-be-trained tasks in the training task queue; A training feature extraction module 20, configured to extract training features from the to-be-trained tasks to obtain the training features of the to-be-trained tasks; A communication operation planning module 30, configured to obtain the communication operation workflow of the to-be-trained tasks and the priorities of the communication operations in the communication operation workflow according to the training features; The communication scheduling execution module 40 is used to obtain a communication scheduling scheme for the task to be trained according to the priority of each communication operation, and execute communication scheduling for the task to be trained according to the communication scheduling scheme.

[0123] In one embodiment, the training feature extraction module 20 is also used to obtain configuration information of the model to be trained according to the task to be trained, and the configuration information of the model to be trained includes model structure information, model parameter scale and training data processing complexity; determine the computing node allocation result and training parallel strategy according to the model configuration information to be trained; based on the computing node allocation result and training parallel strategy, perform model preview on the task to be trained to obtain training preview result; perform training feature extraction on the training preview result to obtain the training feature of the task to be trained.

[0124] In one embodiment, the training feature extraction module 20 is also used to decompose the task to be trained according to the computing node allocation result and the training parallel strategy to obtain a computing subtask queue; based on the computing node parameters corresponding to each computing subtask in the computing subtask queue, perform model training preview on the computing subtask to obtain the computing workload, communication requirements and communication data flow of the computing subtask on the corresponding computing node; summarize the computing workload, communication requirements and communication data flow on the computing node corresponding to each computing subtask to obtain the training preview result, wherein the training preview result includes a computing load preview result and a communication load preview result.

[0125] In one embodiment, the communication operation planning module 30 is further used to obtain a computational graph of the task to be trained based on the training features; construct a communication topology graph of each computing node based on the computational graph, wherein the communication topology graph at least includes all communication links required in a complete iteration process; obtain a communication operation workflow of the task to be trained based on the communication topology graph, wherein each communication operation granularity in the communication operation workflow is an inter-GPU communication granularity.

[0126] In one embodiment, the communication operation planning module 30 is further used to obtain the communication topology diagram and the calculation diagram of the task to be trained according to the training characteristics; obtain the urgency score of the communication operation according to the calculation diagram and the communication load preview result; obtain the computing intensity of the computing node corresponding to the communication operation according to the communication topology diagram and the computing load preview result; obtain the priority of each communication operation in the communication operation workflow according to the urgency score of the communication operation and the computing intensity of the computing node corresponding to the communication operation.

[0127] In one embodiment, the communication operation planning module 30 is further configured to determine, according to the computational graph, the computing nodes corresponding to the respective communication operations in the communication operation workflow and the dependent nodes corresponding to the computing nodes; obtain the communication transmission time and the communication node idle time of the communication operation according to the computing nodes, dependent nodes of the communication operation, and the communication load preview result; and obtain the urgency score of the communication operation according to the communication transmission time and the communication node idle time of the communication operation.

[0128] In one embodiment, the communication operation planning module 30 is further configured to determine the computing nodes corresponding to the communication operation according to the communication topology graph; obtain the computing workload of the computing nodes corresponding to the communication operation according to the computing resource requirement information; and obtain the computing intensity of the computing nodes corresponding to the communication operation according to the computing workload and the standard iteration duration.

[0129] The machine learning training task communication scheduling device provided in this application adopts the machine learning training task communication scheduling method in the above embodiment, and can solve the technical problem of how to optimize network communication scheduling in distributed deep learning training to improve resource utilization and training efficiency in the prior art. Compared with the prior art, the beneficial effects of the machine learning training task communication scheduling device provided in this application are the same as those of the machine learning training task communication scheduling method provided in the above embodiment, and other technical features in the machine learning training task communication scheduling device are the same as the features disclosed in the method of the above embodiment, and will not be elaborated herein.

[0130] This application provides a machine learning training task communication scheduling device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the machine learning training task communication scheduling method in the first embodiment above.

[0131] Next, refer to Figure 5, which shows a schematic structural diagram of a machine learning training task communication scheduling device suitable for implementing the embodiments of the present application. The machine learning training task communication scheduling device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The shown machine learning training task communication scheduling device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0132] As Figure 5 shown, the machine learning training task communication scheduling device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the machine learning training task communication scheduling device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the machine learning training task communication scheduling device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a machine learning training task communication scheduling device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.

[0133] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0134] The machine learning training task communication scheduling device provided by the present application adopts the machine learning training task communication scheduling method in the above embodiments, and can solve the technical problem of how to optimize network communication scheduling in distributed deep learning training to improve resource utilization and training efficiency in the prior art. Compared with the prior art, the beneficial effects of the machine learning training task communication scheduling device provided by the present application are the same as those of the machine learning training task communication scheduling method provided by the above embodiments, and other technical features in the machine learning training task communication scheduling device are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.

[0135] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0136] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0137] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the machine learning training task communication scheduling method in the above embodiments.

[0138] The computer-readable storage medium provided by the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0139] The above computer-readable storage medium may be included in the machine learning training task communication scheduling device; or it may exist separately and not be assembled into the machine learning training task communication scheduling device.

[0140] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the machine learning training task communication scheduling device, the machine learning training task communication scheduling device is caused to: obtain the to-be-trained tasks in the training task queue; perform training feature extraction on the to-be-trained tasks to obtain the training features of the to-be-trained tasks; obtain the communication operation workflow of the to-be-trained tasks and the priorities of each communication operation in the communication operation workflow according to the training features; obtain the communication scheduling scheme of the to-be-trained tasks according to the priorities of each communication operation, and execute the communication scheduling of the to-be-trained tasks according to the communication scheduling scheme.

[0141] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by connecting through an Internet service provider via the Internet).

[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0143] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0144] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned machine learning training task communication scheduling method, and can solve the technical problem of how to optimize network communication scheduling in distributed deep learning training to improve resource utilization and training efficiency in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the machine learning training task communication scheduling method provided in the above embodiments, and will not be elaborated here.

[0145] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the machine learning training task communication scheduling method as described above are implemented.

[0146] The computer program product provided by the present application can solve the technical problem of how to optimize network communication scheduling in distributed deep learning training in the prior art to improve resource utilization rate and training efficiency. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the machine learning training task communication scheduling method provided by the above embodiments, and will not be elaborated here.

[0147] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A communication scheduling method for machine learning training tasks, characterized in that The machine learning training task communication scheduling method includes: Obtain the training tasks to be trained in the training task queue; Extract training features from the training tasks to be trained to obtain the training features of the training tasks to be trained; Based on the training features, obtain the communication operation workflow of the training tasks to be trained and the priorities of each communication operation in the communication operation workflow; Based on the priorities of each communication operation, obtain the communication scheduling scheme for the training tasks to be trained, and execute the communication scheduling for the training tasks to be trained according to the communication scheduling scheme.

2. The communication scheduling method for machine learning training tasks according to claim 1, wherein The obtaining the priorities of each communication operation in the communication operation workflow based on the training features includes: Based on the training features, obtain the communication topology graph and the computation graph of the training tasks to be trained; Based on the computation graph and the communication load rehearsal result, obtain the urgency score of the communication operation; Based on the communication topology graph and the computing power load rehearsal result, obtain the computing intensity of the computing node corresponding to the communication operation; Based on the urgency score of the communication operation and the computing intensity of the computing node corresponding to the communication operation, obtain the priorities of each communication operation in the communication operation workflow.

3. The machine learning training task communication scheduling method according to claim 2, wherein The obtaining the urgency score of the communication operation based on the computation graph and the communication load rehearsal result includes: Based on the computation graph, determine the computing nodes corresponding to each communication operation in the communication operation workflow and the dependent nodes corresponding to the computing nodes; Based on the computing node, dependent node of the communication operation and the communication load rehearsal result, obtain the communication transmission time and the communication node idle time of the communication operation; Based on the communication transmission time and the communication node idle time of the communication operation, obtain the urgency score of the communication operation.

4. The machine learning training task communication scheduling method according to claim 2, wherein, The obtaining the computing intensity of the computing node corresponding to the communication operation based on the communication topology graph and the computing power load rehearsal result includes: Based on the communication topology graph, determine the computing node corresponding to the communication operation; Based on the computing resource requirement information, obtain the computing workload of the computing node corresponding to the communication operation; Based on the computing workload and the standard iteration duration, obtain the computing intensity of the computing node corresponding to the communication operation.

5. The communication scheduling method for machine learning training tasks according to claim 1, wherein The extracting training features from the training tasks to be trained to obtain the training features of the training tasks to be trained includes: Based on the training tasks to be trained, obtain the configuration information of the model to be trained, where the configuration information of the model to be trained includes model structure information, model parameter scale, and training data processing complexity; Based on the configuration information of the model to be trained, determine the computing node allocation result and the training parallel strategy; Based on the computing node allocation result and the training parallel strategy, perform a model training rehearsal on the training tasks to be trained to obtain a training rehearsal result; Extract training features from the training rehearsal result to obtain the training features of the training tasks to be trained.

6. The machine learning training task communication scheduling method according to claim 5, wherein The performing a model training rehearsal on the training tasks to be trained based on the computing node allocation result and the training parallel strategy to obtain a training rehearsal result includes: According to the computing node allocation result and the training parallel strategy, decompose the to-be-trained task according to computing nodes to obtain a computing subtask queue; Based on the computing node parameters corresponding to each computing subtask in the computing subtask queue, perform a model training rehearsal on the computing subtasks to obtain the computing workload, communication requirements, and communication data traffic of the computing subtasks on the corresponding computing nodes; Summarize the computing workload, communication requirements, and communication data traffic of each computing subtask on the corresponding computing node to obtain the training rehearsal result, where the training rehearsal result includes a computing power load rehearsal result and a communication load rehearsal result.

7. The communication scheduling method for machine learning training tasks according to claim 1, wherein The obtaining of the communication operation workflow of the to-be-trained task according to the training features includes: Obtain the computation graph of the to-be-trained task according to the training features; According to the computation graph, construct a communication topology graph for each computing node, where the communication topology graph includes at least all communication links required in a complete iteration process; According to the communication topology graph, obtain the communication operation workflow of the to-be-trained task, where the fine granularity of each communication operation in the communication operation workflow is the fine granularity of inter-GPU communication.

8. A communication scheduling device for machine learning training tasks, characterized in that, The communication scheduling device for machine learning training tasks includes: A task queue management module, configured to obtain the to-be-trained tasks in the training task queue; A training feature extraction module, configured to extract training features of the to-be-trained task to obtain the training features of the to-be-trained task; A communication operation planning module, configured to obtain the communication operation workflow of the to-be-trained task and the priorities of each communication operation in the communication operation workflow according to the training features; A communication scheduling execution module, configured to obtain a communication scheduling scheme for the to-be-trained task according to the priorities of each communication operation, and perform communication scheduling on the to-be-trained task according to the communication scheduling scheme.

9. A machine learning training task communication scheduling device, characterized in that, The communication scheduling device for machine learning training tasks includes: a memory, a processor, and a machine learning training task communication scheduling program stored on the memory and executable on the processor, where the machine learning training task communication scheduling program is configured to implement the steps of the machine learning training task communication scheduling method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, A machine learning training task communication scheduling program is stored on the storage medium, and when the machine learning training task communication scheduling program is executed by a processor, it implements the steps of the machine learning training task communication scheduling method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Large model training method and system based on green distributed computing power center

    CN120450088A

  • Information processing method and device, equipment and storage medium

    CN120508396A

  • Information processing method, device, equipment and storage medium

    CN120508396B