A Flexible and Efficient Distributed Machine Learning Method in Dynamic Scenarios
By constructing a dynamic network communication model and adaptive scheduling and allocation algorithm for node tasks, the problems of communication bottlenecks and uneven computing power distribution in distributed machine learning are solved, efficient training in dynamic scenarios is achieved, and the communication efficiency and training efficiency of the cluster are improved.
Patent Information
- Application Number
- CN202211663723.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-12-23
AI Technical Summary
The existing distributed machine learning methods face communication bottlenecks, uneven computing power distribution and behind-the-straight nodes in dynamic scenarios, resulting in low training efficiency, especially in the case of complex large-scale clusters or network conditions.
Using dynamic and efficient distributed training framework and greedy ideas, a dynamic network communication model is built, and a heuristic communication model is constructed through node task adaptive scheduling and allocation algorithm and Huffman to realize dynamic orchestration of training nodes and flexible scheduling of computing nodes, and optimize the network communication efficiency in the cluster.
In the case of complex large-scale clusters or network conditions, the communication bottlenecks are effectively alleviated, the communication efficiency and training efficiency of the cluster are improved, the flexibility and adaptability of the cluster are improved, and the efficient computing node scheduling can be maintained under dynamically changing training nodes.
Smart Images

Figure CN116028175B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed machine learning, and in particular to a flexible and efficient distributed machine learning method for dynamically and efficiently organizing and scheduling training nodes in a distributed cluster in a dynamic scenario. Background Art
[0002] In recent years, thanks to the large training datasets and the models trained by complex deep neural networks that can effectively approximate the decision boundaries of difficult problems, artificial intelligence technology has achieved extensive success in many fields including speech recognition, image processing, natural language, autonomous driving, and medical care. However, the training of models, especially deep neural networks, is a computationally intensive task, and the increase in the scale of training datasets and deep neural networks will make model training take longer. Although the computing power of hardware has also been improving in recent years, it is still far from meeting the computing power requirements for complex models. For example, training the BERT model on a single TPU takes more than 1.5 months. Therefore, using a distributed cluster to train complex models is an effective method.
[0003] In distributed machine learning, the training nodes use the stochastic gradient descent algorithm and part of the training data stored locally to train the model. After obtaining the trained local model, the training nodes update the local model to the global model through a series of gradient aggregation operations, and send the global model to all training nodes for further iterative training until the global model converges.
[0004] Using a large-scale training cluster, distributed machine learning can expand the system computing power and effectively reduce the time spent in the model training process. On the other hand, however, the gradient aggregation and gradient distribution operations also introduce high communication overheads, which become an important bottleneck affecting the performance expansion of distributed machine learning. And as the cluster scale increases, the impact of this bottleneck becomes more prominent: some work uses 128 Nvdia V100 GPUs to train ResNet-50 on a high-performance platform. The experimental results show that compared with a single Nvdia V100, the distributed system only achieves a 40-fold speedup effect, and the scaling effect is very low. In fact, the challenges faced by distributed systems are multi-faceted:
[0005] First, the distribution of communication resources within a distributed system is not uniform. Inside a node, benefiting from the advantages brought by the communication bus and internal communication protocols, GPUs can achieve a high communication bandwidth of 20 - 25 GB / s. In particular, the latest NVLink-C2C technology recently proposed in the industry can achieve a communication bandwidth of 900 GB / s. However, the bandwidth of inter-node communication is much lower than the internal communication bandwidth. At the same time, due to the existence of communication interference, data packets transmitted in the network often experience packet loss and latency phenomena.
[0006] Second, the computing performance of the training nodes participating in the training is not the same. The collaborative training of models in a cluster is a relatively synchronous task. Only when the training nodes complete their respective tasks can the next operation be carried out. However, due to the influence of the hardware parameters of the training nodes and the load conditions of the tasks, the computing power they can provide is not the same. Such heterogeneous training nodes will cause the cluster to be unable to complete the training task within a similar time period, resulting in waiting latency.
[0007] Furthermore, there are straggler training nodes in the distributed training system. During the actual training process, due to factors such as communication latency and node failures, training nodes become straggler nodes. Other nodes in the cluster will wait for the straggler training nodes to complete the training task, which will affect the training process of the entire cluster. Although some algorithms alleviate the impact on the entire training process by selectively and directly discarding straggler nodes, they still cannot well handle the problems in dynamic cluster scenarios.
[0008] Currently, existing work has adopted various methods to alleviate the above problems. On the one hand, by reasonably invoking and allocating resources, the utilization rate of resources within the cluster is improved. For example, the training process of a deep learning model can generally be regarded as a directed graph. Based on this characteristic, some work adopts a pipeline-like method to cover the training and communication costs. Usually, this pipeline method reduces the time cost of each iteration by specifying the aggregation order of the gradient model or using historical model weights. However, the acceleration effect of this method has certain limitations. When the bandwidth resources are severely restricted, its acceleration effect will be greatly limited. In addition, some other work starts from communication optimization. According to the characteristics of communication data during the cluster training process, the communication protocol and hardware devices are adjusted to reduce the packet loss rate or communication latency of communication data during the communication process, improve the communication performance of the overall system, and accelerate the cluster training speed.
[0009] On the other hand, reducing the network bandwidth load to alleviate the bandwidth bottleneck faced by the cluster is also a commonly used method. Among them, a more common approach is to compress the transmitted gradient information, that is, using fewer transmitted bits to compress the gradient data or using sparsification methods to select important gradient data, which can reduce the communication volume of the cluster. In addition, having the computing nodes perform multiple iterative trainings locally to reduce the communication frequency is also an effective method to alleviate the network bottleneck. However, although such methods improve the training efficiency of the model, to a certain extent, they will affect the accuracy of the model. Although these works have optimized a series of communication data using methods in related fields such as statistics, the effectiveness of the optimization method in a specific scenario and the additional overhead brought by this method still need to be considered. Summary of the Invention
[0010] The object of the present invention is to provide a flexible and efficient distributed machine learning method in a dynamic scenario for the deficiencies of the prior art. It adopts a method of constructing a network communication model with a dynamically changing network cluster, and uses a dynamic and efficient distributed training framework and a greedy idea to realize the dynamic arrangement of training nodes in the cluster and the flexible scheduling of computing nodes, so as to flexibly cope with various challenges generated by the distributed system in a dynamic scenario, and has achieved a good acceleration effect on a large-scale distributed cluster, greatly improving the communication efficiency of the cluster. This method proposes a node task adaptive scheduling and allocation algorithm, which relies on the above-mentioned constructed dynamic communication network to improve the network communication efficiency between nodes, can efficiently and flexibly schedule the computing nodes under a distributed training cluster, can effectively alleviate the communication bottleneck in the case of a large-scale cluster or complex network conditions, and preferably solves the uneven distribution of computing power commonly existing in the actual network, and can effectively alleviate the communication bottleneck problem in the case of a large-scale cluster or complex network conditions.
[0011] The specific technical solution to achieve the object of the present invention is: a flexible and efficient distributed machine learning method in a dynamic scenario, including the following specific steps:
[0012] S1. Complete the construction of the cluster model
[0013] The cluster nodes participating in the distributed training are divided into two different roles: training nodes participating in model calculation and supervisor nodes for task scheduling. Among them, the training nodes are mainly responsible for the local training of the model and the fusion and transmission of model gradients, and the supervisor nodes are responsible for recording the states of the training nodes and realizing task scheduling according to the relevant information of the nodes.
[0014] In practical situations, although network topologies vary widely, connectivity can be ensured between any two nodes in the same cluster. Therefore, this method abstractly constructs the connections between all nodes as a complete graph, thus eliminating the impact of differences in physical network topologies on the cluster scheduling strategy.
[0015] S2. Construct a dynamic cluster communication network to complete the iterative training of the model
[0016] Based on the cluster node model described in step S1, the present invention analyzes the efficiency of the training process of distributed training. Under a dynamically changing network cluster, the time cost T of distributed machine learning ps is defined by the following equation (e):
[0017]
[0018] where E is the number of training iterations; is the local training time; is the network communication time.
[0019] The present invention proposes a heuristic dynamic network communication model Greed-Comm based on Huffman construction to improve the network communication efficiency between nodes. Its time cost T gw can be expressed by the following equation (a):
[0020]
[0021] where is the time spent on local training of the training node, is the time consumed by node selection, network communication, and gradient aggregation, is the time spent on gradient distribution. Here, w is the gradient model, b is the network bandwidth, Δs is the node selection time, Δt is the gradient aggregation time, p k is the number of training nodes in each aggregation operation, k i is the height of the network communication tree Greed-Comm. Specifically, it includes the following steps:
[0022] S21. First, before a training node participates in cluster training, it needs to perform initialization operations. Specifically, the training node needs to first report its own status to the supervisor, and then perform node initialization operations, such as training environment preparation, training data preparation, etc. At this time, this training node will not participate in cluster training. In particular, if the cluster is performing the first iterative training of the model at this time, the training node can immediately update its status to the supervisor and perform the first round of training. After the training node completes the initialization preparation, it will report to the supervisor that it is ready and wait to participate in subsequent iterative training.
[0023] S22. After the training node reports its ready state to the supervisor, the supervisor will participate in scheduling the training nodes to complete the aggregation of gradient information. After receiving the ready information of the training nodes, the supervisor will check the training node information in the ready pool and select several ready training node information from the ready pool according to the training node selection strategy to provide to the training node for local model gradient aggregation. If there are no qualified nodes in the ready pool, the supervisor will throw this training node into the ready pool and wait for the supervisor to schedule it. On the other hand, when the training node receives the ready information provided by the supervisor, it will pull gradient data from the corresponding training nodes according to this information and fuse these gradient data. After completing the fusion of the reshuffled gradient data, the training node updates its ready state to the supervisor and repeats the above steps until the gradient information of the entire cluster is obtained. During this period, when the ready training node receives a convergence request from other nodes, it will update its own state to the convergence state while providing gradient data, waiting to obtain the global gradient information and perform the next iteration.
[0024] S23. After completing step S22, there will be a training node at the root position of the Greed-Comm tree that has the global model obtained in this round of iteration. This training node will parallelly transfer the global model to all training nodes according to the communication tree constructed in the above steps. Specifically, after receiving the latest global gradient information, the training node will first update the local model and transfer the global gradient in the opposite direction of gradient aggregation to its associated sub-training nodes, and then start the next round of model iteration. In this way, the global gradient model is completed for distribution. Steps S21, S22, and S23 together constitute a dynamic and efficient communication network to alleviate the communication bottleneck in the network.
[0025] S3. Construct a cluster task allocation model
[0026] Aiming at the problem of uneven computing power distribution caused by node heterogeneity or node load in the actual network, the present invention proposes a node task adaptive scheduling and allocation algorithm. Through the post-sampling technology, the computing power of nodes in the cluster is evaluated, and the training tasks are predicted and scheduled according to the evaluation results, so as to optimize the training efficiency of the entire cluster, which specifically includes the following steps:
[0027] S31. Sample and evaluate the computing power of training nodes in the cluster
[0028] Based on the gradient data exchange in step S2, when the training node enters this round of training, it will record a start timestamp t s , when this round of gradient training ends, according to the current timestamp and the start timestamp t s derive the evaluation value ρ of the computing power of the training node in this round of trainingt,i While recording the status of the training nodes, the supervisor also records the computing power evaluation information of the training nodes, and the evaluation value ρ of the computing power t,i is calculated by the following formula (b):
[0029]
[0030] where d t,i represents the size of the training set used by the i-th training node in the t-th iteration, and c t,i represents the time consumed by the i-th training node to complete the t-th iteration training.
[0031] S32. The supervisor performs adaptive task allocation according to the computing power evaluation information of the training nodes
[0032] After obtaining the computing power evaluation information of all training nodes in the training cluster, the supervisor will adjust the workload of the training nodes according to the change of its computing power evaluation value. The training set allocated to each training node in each iteration is based on d defined by the following formula (c) t+1,i for allocation:[[]]
[0033]
[0034] where d t+1,i is the data set used by the training node i in the t + 1-th round, ρ t,i is the computing power evaluation value of the training node i in the t-th round, is the evaluation of the computing power of the entire cluster in the t-th round, D train represents the entire training set, and m represents the number of data blocks. And during gradient aggregation, the local model is weighted and aggregated according to the size of the training set by the following formula (d):
[0035]
[0036] where w t+1 is the global model in the t + 1-th round, w t+1,i is the local model obtained by the training node i in the t-th round of training, and d t,i is the size of the training set used by the training node i in the t-th round of training.
[0037] The reason why the present invention divides all data into several data blocks is that when the number of data blocks is too large, the amount of training data contained in a single data block is small, and it is easy for weak nodes to have gradient oscillations during gradient training on a training set with insufficient data volume, which affects the final convergence effect of the model; while if the number of data block divisions is too small, the granularity of the data blocks will be too large, which will instead affect the effect of task load balancing during model training. In addition, when a training node first participates in the training process, the supervisor will first provide a benchmark task, and then further adjust its task according to the change state of the computing power estimation during the training process until all training nodes in the cluster can be assigned stable training tasks.
[0038] Compared with the prior art, the present invention improves the overall training efficiency of the model. Especially in the case of large-scale clusters or limited network communication, a dynamic training node scheduling technology is realized. This technology can alleviate the communication bottleneck generated during cluster training, improve the flexibility of the cluster in actual application scenarios, and enable the cluster to adaptively respond to dynamically changing training nodes. In addition, for heterogeneous training nodes, the present invention provides an adaptive training task allocation method, which can adapt to the heterogeneous training nodes in the cluster and flexibly and efficiently allocate the computing power of the cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic diagram of the network architecture of the present invention;
[0040] Figure 2 It is an example of constructing a dynamic gradient aggregation tree according to the present invention;
[0041] Figure 3 It is a schematic diagram of the conversion of the training node state in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. Obviously, the listed examples are only used to explain the present invention and are not used to limit the scope of the present invention.
[0043] Embodiment 1
[0044] A flexible and efficient distributed machine learning method in a dynamic scenario according to the present invention specifically includes the following steps:
[0045] S1. Complete the construction of the cluster model
[0046] First, the cluster nodes participating in distributed training are divided into two different roles: training nodes participating in model calculation and supervisor nodes for task scheduling. Among them, the training nodes are mainly responsible for local training of the model and the fusion and transmission of model gradients, and the supervisor nodes are responsible for recording the states of the training nodes and realizing task scheduling according to the relevant information of the nodes.
[0047] In the actual situation, although the network topologies vary greatly, the connectivity can be guaranteed between any two nodes under the same cluster. Therefore, the present invention abstractly constructs the connections between all nodes as a complete graph, so as to get rid of the influence of the differences in physical network topologies on the cluster scheduling strategy.
[0048] S2. Construct a dynamic cluster communication network to complete the iterative training of the model
[0049] Refer to Figure 1 , which shows the cluster architecture model designed by the present invention. Based on the cluster node model described in step S1, the present invention analyzes the efficiency of the distributed training process. Under the dynamically changing network cluster, the time cost T of distributed machine learning ps is defined by the following formula (e):
[0050]
[0051] where E is the number of training iterations, is the local training time, is the network communication time.
[0052] The present invention constructs a heuristic dynamic network communication model Greed-Comm based on Huffman to improve the network communication efficiency between nodes. Its time cost T gw can be expressed by the following formula (a):
[0053]
[0054] where, is the time spent on local training of the training node, is the time consumed by node selection, network communication and gradient aggregation, is the time spent on gradient distribution, where w is the gradient model, b is the network bandwidth, Δs is the node selection time, Δt is the gradient aggregation time, p k is the number of training nodes in each aggregation operation, k i is the height of the network communication tree Greed-Comm. The construction of the dynamic cluster communication network to complete the iterative training of the model is specifically divided into the following steps:
[0055] S21. Before a training node participates in cluster training, it needs to perform initialization operations: Specifically, the training node needs to first report its own status to the supervisor, and then perform node initialization operations, such as preparing the training environment and training data. At this time, the training node will not participate in cluster training. In particular, if the cluster is performing the first iterative training on the model at this time, the training node can immediately update its status to the supervisor and perform the first round of training. After the training node completes the initialization preparation, it will report to the supervisor that it is ready and wait to participate in subsequent iterative training.
[0056] S22. When the training node reports its ready status to the supervisor, the supervisor will participate in scheduling the training node to complete the aggregation of gradient information: After receiving the ready information of the training node, the supervisor will check the training node information in the ready pool and select several ready training node information from the ready pool according to the training node selection strategy to provide local model gradient aggregation for this training node. If there are no qualified nodes in the ready pool, the supervisor will throw this training node into the ready pool and wait for the supervisor to schedule. On the other hand, when the training node receives the ready information provided by the supervisor, it will pull gradient data from the corresponding training node based on this information and fuse these gradient data. After completing the fusion of the shuffled gradient data, this training node updates its ready status to the supervisor and repeats the above steps until the gradient information of the entire cluster is obtained. During this period, when a ready training node receives a convergence request from other nodes, it will update its own status to the convergence state while providing gradient data and wait to obtain the global gradient information and perform the next iteration.
[0057] See Figure 2 , this figure shows an example of the gradient aggregation process: Training nodes 1-7 sequentially complete local model training, update their ready status to the supervisor node, and the supervisor completes local gradient scheduling according to Group 1 and Group 2 respectively, and finally completes the aggregation of the global gradient model.
[0058] S23. After completing step S2, there will be a training node at the root position of the Greed-Comm tree that has the global model obtained in this round of iteration. This training node will parallelly transmit the global model to all training nodes according to the communication tree constructed in the above steps. Specifically, after receiving the latest global gradient information, the training node will first update the local model and transmit the global gradient in the opposite direction of gradient aggregation to its associated sub-training nodes, and then start the next round of model iteration. In this way, the global gradient model is completed for distribution.
[0059] The above steps S21, S22, and S23 together constitute a dynamic and efficient communication network to alleviate the communication bottleneck in the network.
[0060] S3. Construct a cluster task allocation model
[0061] Refer to Figure 3 , for ease of understanding, this figure shows the transformation model of the training node status in the cluster. In view of the problem of uneven computing power distribution caused by node heterogeneity or node load in the actual network, the present invention proposes a node task adaptive scheduling and allocation algorithm. The computing power of nodes in the cluster is evaluated through the post-sampling technology, and the training tasks are predicted and scheduled according to the evaluation results, so as to optimize the training efficiency of the entire cluster. Constructing the cluster task allocation model specifically includes the following steps:
[0062] S31. Sample and evaluate the computing power of training nodes in the cluster
[0063] Based on the gradient data exchange in S2, before the training node enters this round of training, a start timestamp t will be recorded s ; when this round of gradient training ends, according to the current timestamp and the start timestamp t s Derive the evaluation value ρ of the computing power of the training node in this round of training t,i , while the supervisor records the status of the training node, it will also record the computing power evaluation information of the training node. The evaluation value ρ of the computing power t,i Is calculated by the following formula (b):
[0064]
[0065] Where d t,i Represents the size of the training set used by the i-th training node at the t-th iteration, and c t,i Represents the time consumed by the i-th training node to complete the t-th iteration training.
[0066] S32. The supervisor performs adaptive task allocation according to the computing power evaluation information of the training node
[0067] After the supervisor obtains the computing power evaluation information of all training nodes in the training cluster, it will adjust the workload of the training nodes according to the change of its computing power estimation. The training set allocated to each training node in each iteration is defined by d in the following formula (c) t+1,i For allocation:
[0068]
[0069] Where d t+1,i Is the data set used by training node i in the t+1-th round, ρ t,i Is the computing power evaluation value of training node i in the t-th round, Is the evaluation of the computing power of the entire cluster at the t-th round, and D trainLet \(\mathcal{D}\) denote the entire training set, and \(m\) denote the number of data chunks. When aggregating gradients, the local models are weighted and aggregated according to the size of the training set by the following equation (d):
[0070]
[0071] where \(w^{(t + 1)}\) t+1 is the global model at the \((t + 1)\)-th round, \(w_i^{(t)}\) t+1,i is the local model obtained by training node \(i\) at the \(t\)-th round, and \(d_i^{(t)}\) t,i is the size of the training set used by training node \(i\) during the \(t\)-th round of training.
[0072] It should be noted that the reason for dividing all the data into several data chunks in the present invention is that when the number of data chunks is too large, the amount of training data contained in a single data chunk is small, and it is easy for weak nodes to experience gradient oscillations during gradient training on a training set with insufficient data volume, which affects the final convergence effect of the model; while if the number of data chunks is too small, the granularity of the data chunks will be too large, which will instead affect the effect of task load balancing during model training. In addition, when a training node first participates in the training process, the supervisor will first provide a benchmark task, and then further adjust its task according to the changing state of the computing power estimation during the training process until all training nodes in the cluster can be assigned stable training tasks. The present invention proposes a dynamic and efficient distributed training framework to alleviate the limitation of the communication bottleneck and reasonably allocate system resources, realizing dynamic training node scheduling, better alleviating the communication bottleneck generated during cluster training, improving the overall training efficiency of the model, and achieving good acceleration effects especially in large-scale clusters or when network communication is restricted.
[0073] The above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A flexible and efficient distributed machine learning method in a dynamic scenario, characterized in that The method includes the following specific steps: S1. Construction of the cluster node model The cluster nodes participating in distributed training are divided into training nodes for model calculation and supervisor nodes for task scheduling according to two different roles. The connections between all nodes are abstractly constructed as a complete graph to complete the construction of the cluster model. The training nodes are responsible for local training of the model and fusion and transmission of model gradients; the supervisor nodes are responsible for recording the states of the training nodes and implementing task scheduling according to the relevant information of the nodes; S2. Construction of the cluster communication network Under the dynamically changing network cluster, a heuristic dynamic network communication model Greed-Comm is constructed based on the Huffman algorithm to complete the iterative training of the model; The construction of the dynamic network communication model Greed-Comm specifically includes the following steps: S21. Initialization operation Before participating in cluster training, the training node needs to report its own state to the supervisor, and then perform operations including: preparation of the training environment and preparation of training data. After completing the initialization operation of the training node, it reports to the supervisor that the training node is ready and waits to participate in subsequent iterative training; S22. Aggregation of gradient information After receiving the ready information of the training node, the supervisor checks the information of the training nodes in the ready pool and selects several ready training node information from the ready pool according to the training node selection strategy to provide local model gradient aggregation for this training node; if there are no qualified nodes in the ready pool, the supervisor will throw this training node into the ready pool and wait for the supervisor to schedule; the training node pulls gradient data from the corresponding training nodes according to the information in the ready pool and fuses these gradient data. After completing the fusion of this group of gradient data, this training node updates its ready state to the supervisor and repeats the above steps until the gradient information of the entire cluster is obtained; during this period, when a ready training node receives a convergence request from other nodes, it will update its own state to the convergence state while providing gradient data and wait to obtain the global gradient information and perform the next iteration; After step S22 is completed, there will be a training node at the root position of the Greed-Comm tree that has the global model obtained in this round of iteration. This training node will transmit the global model in parallel to all training nodes according to the communication tree constructed in the above steps; After receiving the latest global gradient information, the training node will first update the local model, and transmit the global gradient to its associated sub-training nodes in the opposite direction of gradient aggregation, and then start the next round of model iteration to complete the construction of the dynamic network communication model Greed-Comm on the cluster node model; S3. Construction of the cluster task allocation model The node task adaptive scheduling and allocation algorithm is adopted to construct a cluster task allocation model through the post-sampling technology, evaluate the computing power of the nodes in the cluster, and perform predictive scheduling on the training tasks according to the evaluation results, so as to optimize the training efficiency of the entire cluster.
2. The flexible and efficient distributed machine learning method in a dynamic scenario according to claim 1, characterized in that The cluster communication network in step S2 is a heuristic dynamic network communication model Greed-Comm constructed based on the Huffman algorithm, and the time cost of its distributed machine learning is defined by the following formula (a): (a); Among them, is the time spent on local training of the training node; is the time consumed for node selection, network communication, and gradient aggregation; is the time spent on gradient distribution; is the gradient norm; is the network bandwidth; is the node selection time; is the time for gradient aggregation; is the number of training nodes in each group of aggregation operations; is the height of the network communication tree Greed-Comm; 3. The flexible and efficient distributed machine learning method in a dynamic scenario according to claim 1, characterized in that After the cluster task allocation model in step S3 is used, the post-sampling technology is adopted to jointly evaluate the computing power of nodes, and the training tasks are allocated according to the evaluation results, which specifically include the following steps: S31. Sample and evaluate the computing power of training nodes in the cluster Before entering the current round of training at the training node, a start timestamp will be recorded ; after the current round of gradient training is completed, based on the current timestamp and the start timestamp derive the evaluation value of the computing power of the training node in the current round of training , while recording the status of the training node, the supervisor will also record the computing power evaluation information of the training node; the evaluation value of the computing power is calculated by the following formula (b): (b); Among them, represents the size of the training set used by the th training node at the th iteration, and represents the time consumed by the th training node to complete the th iteration training; S32. The supervisor performs adaptive task allocation according to the computing power evaluation information of the training nodes, which specifically includes: S321. After obtaining the evaluation information of the computing power of all training nodes in the training cluster, the supervisor adjusts the workload of the training nodes according to the change of its computing power evaluation value and distributes the training set allocated to the training nodes in each iteration according to the following formula (c): for distribution: (c); Among them, is the data set used by the training node in the round; is the computing power evaluation value of the training node in the round; is the evaluation of the computing power of the entire cluster at the round; represents the entire training set; represents the number of data blocks; S322. During gradient aggregation, the local model performs weighted aggregation according to the size of the training set by the following formula (d): Aggregation: (d); Among them, is the global model of the th round; is the local model obtained by the training node in the th round of training; is the size of the training set used by the training node in the th round of training. S323. When a training node first participates in the training process, the supervisor will first provide a benchmark task, and then further adjust its task according to the change state of the computing power valuation during the training process until all training nodes in the cluster can be allocated stable training tasks.
Citation Information
Patent Citations
Load balancing method based on automatic workload tuning
CN110888744A
Task configuration method for heterogeneous distributed machine learning cluster
CN113590321A