Distributed machine learning training system and method based on network flow rate control
By deploying traffic information collectors and job schedulers on switches, the priority and sending rate of DNN training jobs are dynamically adjusted, and the communication contention problem in distributed machine learning training is solved, network link utilization is improved and average job completion time is reduced.
Patent Information
- Application Number
- CN202411815653.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art has failed to completely solve the communication contention problem of different DNN training jobs on the same bottleneck link in distributed machine learning training, resulting in low network link utilization and long average job completion time.
By deploying the traffic information collector module and job scheduler module on the switch, the communication mode information of the job flow is collected and the remaining iteration completion time is calculated, and the job priority and sending rate are dynamically adjusted to prioritize the job flow with the shortest remaining iteration completion time.
It significantly improves the utilization rate of network links, reduces network congestion and idle time, and effectively reduces the average job completion time.
Smart Images

Figure CN119918618A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high performance computing and cloud computing, and specifically to a network flow rate control technology in distributed machine learning training, which can be used to improve the communication efficiency and overall training speed of deep neural network (DNN) training operations in data center networks. Background Art
[0002] In the field of distributed machine learning training, especially deep neural network (DNN) training, existing training acceleration technologies are mainly concentrated in the following aspects: machine learning scheduling technology, changing congestion control algorithms, network-aware DNN schedulers, and multi-resource scheduling technology.
[0003] Machine learning scheduling technologies include Gandiva, Optimus, Tiresias, Themis, Pollux, and Cassini.
[0004] Gandiva: Proposed by a team from Beihang University and others, it considers the characteristics of deep learning workflows, makes primitives more efficient through time slicing and migration, and proposes an introspective scheduling framework to improve early feedback time and improve training group efficiency.
[0005] Optimus: Proposed by a team from the University of Hong Kong, it builds a performance model to track training progress, uses online fitting to predict the number of model convergence steps, and proposes a task placement scheme to reduce communication overhead.
[0006] Tiresias: proposed by Juncheng Gu et al., proposed a discrete two-dimensional indicator scheduling algorithm to minimize the average job completion time, and proposed a placement algorithm to relax the integrated placement constraints.
[0007] Themis: Proposed by a team from the University of Wisconsin, it uses an auction algorithm to balance the fairness and efficiency of completion time and uses a two-level scheduling architecture.
[0008] Pollux: Proposed by a team from Carnegie Mellon University, the adaptive scheduler jointly optimizes system throughput and statistical efficiency to configure the right combination of resource allocation and training parameters for DL jobs.
[0009] Cassini: Proposed by a team from MIT, it uses a centralized scheduler to adjust the job communication start time by determining the time shift value through an affinity graph, so that the job communication pattern is staggered.
[0010] Changed congestion control algorithms include DCTCP and DCQCN.
[0011] DCTCP: An ECN-based congestion control algorithm that adjusts the congestion window through fine-grained congestion feedback to reduce queue length and delay.
[0012] DCQCN: Combining the advantages of ECN and QCN, DCQCN adjusts the sending rate through explicit congestion notification and quantitative feedback mechanism to maintain low queue length and low latency.
[0013] Network-aware DNN schedulers include TACCL, BytePS, TicTac, ByteScheduler, and SYNDICATE.
[0014] TACCL: Proposed by a team from the University of Texas at Austin, it uses communication sketches to guide collective algorithm synthesis, adaptively select and optimize collective communication algorithms.
[0015] BytePS: Proposed by a team from Tsinghua University, it provides a unified framework to optimize computing and communication paths and fully utilize GPU and CPU resources in heterogeneous clusters.
[0016] TicTac: Identify and dynamically adjust communication and computation order, using priorities to improve iteration time.
[0017] ByteScheduler: proposes a general communication scheduler that uses a priority-based scheduling strategy to dynamically adjust the order of communication operations.
[0018] SYNDICATE: A joint optimization framework is used to combine scheduling strategies and execution plan optimization to achieve efficient distributed training.
[0019] Multi-resource scheduling technologies include Muri and Synergy.
[0020] Muri: Proposed by a team from Peking University, this scheduling technique staggers the use of key resources to achieve high resource utilization and reduce job completion time.
[0021] Synergy: Proposed by a team from the University of Texas at Austin, it proposes a multi-resource staggered scheduling method and uses optimistic analysis to infer the sensitivity of DNN jobs to resources.
[0022] The above existing technologies have different focuses in the field of distributed machine learning training, but none of them can completely solve the communication contention problem of different DNN training jobs on the same bottleneck link. The present invention aims to provide a new distributed machine learning training acceleration solution based on a network flow rate control method to improve network link utilization and reduce average job completion time. Summary of the invention
[0023] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a distributed machine learning training system and method based on network flow rate control, which aims to improve network link utilization and reduce average job completion time through a method based on network flow rate control.
[0024] In order to solve the above technical problems, the technical solutions proposed in the present invention are: 1. A distributed machine learning training system based on network flow rate control, characterized in that: it includes a traffic information collector module and a job scheduler module; the traffic information collector module collects the number of bytes in the job flow communication mode and records the timestamp of the arrival of the data packet on the switch, and calculates the remaining iteration completion time of the job flow; the job scheduler module determines the priority of the job according to the remaining iteration completion time of the job flow, so that the switch gives priority to sending the job flow with the shortest remaining iteration completion time.
[0025] In the above-mentioned distributed machine learning training system based on network flow rate control, preferably, the traffic information collector module deploys two Count-Min Sketches on the switch, one for collecting the number of bytes in the job flow communication mode, and the other for recording the timestamp of the arrival of the data packet.
[0026] In the above-mentioned distributed machine learning training system based on network flow rate control, preferably, the flow information collector module collects and stores the flow ID of each job flow and the total number of bytes of the previous iteration of each job flow by deploying KV-Store on the switch; at the same time, the remaining iteration completion time of the job flow is calculated.
[0027] In the above-mentioned distributed machine learning training system based on network flow rate control, preferably, the method for calculating the remaining iteration completion time of the job flow includes the following steps: ① calculating the ratio TimeGap of the time difference between the arrival of two adjacent data packets of the job flow, and determining the total number of bytes transmitted in each iteration of the job flow; ; ② Determine whether the value of TimeGap is less than the preset threshold; if the value of TimeGap is less than or equal to the preset threshold, the number stored in Count-Min Sketch is the total number of bytes transmitted in one round of communication phase; if the value of TimeGap is greater than the preset threshold, it means that this round of communication phase has not ended, and the size of the current data packet continues to be added to Count-Min Sketch to continuously calculate the statistics. ③ After one round of iterative training transmission, the total number of bytes of iterative training in the communication phase obtained by the switch is used in the next round In rounds of iterative training; ④ Use the total number of bytes of iterative training in the communication phase obtained by the switch in step 3) to calculate the remaining iteration time T ric ; .
[0028] In the above-mentioned distributed machine learning training system based on network flow rate control, preferably, the job scheduler module uses the congestion control and rate adjustment module to mark the job flow whose remaining iteration completion time exceeds the preset time with an ECN mark or send a CNP packet to reduce its sending rate, thereby giving priority to sending the job flow with the shortest remaining iteration completion time.
[0029] In the above-mentioned distributed machine learning training system based on network flow rate control, preferably, if the flow ID of the job flow stored on the KV-Store and the total number of bytes of the previous iteration of each job flow do not change within a preset time, they are directly deleted or overwritten with a new job flow.
[0030] A distributed machine learning training method based on network flow rate control includes the following steps: 1) Deploy a traffic information collector module, a job scheduler module, and a congestion control and rate adjustment module on the switch; the traffic information collector module includes a KV-Store and two Count-Min Sketches; the two Count-Min Sketches, one for collecting the number of bytes in the job flow communication mode, and the other for recording the timestamp of the arrival of the data packet; 2) The KV-Store on the flow information collector module collects and stores the flow ID and the total number of bytes of the previous iteration of each job flow; at the same time, the remaining iteration completion time of the job flow is calculated; 3) The job scheduler module determines the priority of the job according to the remaining iteration completion time of the job flow, so that the switch preferentially sends the job flow with the shortest remaining iteration completion time; 4) The congestion control and rate adjustment module deployed on the switch will mark the job flow whose remaining iteration completion time exceeds the preset time with an ECN mark or send a CNP packet to reduce its sending rate, thereby dynamically adjusting the job sending rate; 5) Repeat steps 2)-4) until all job flows have completed training.
[0031] The above-mentioned distributed machine learning training method based on network flow rate control, preferably, the specific method of step 2) includes the following steps: ① Calculate the ratio of the time difference between the arrival of two adjacent data packets in the job flow, TimeGap, and determine the total number of bytes transmitted in each iteration of the job flow; ; ② Determine whether the value of TimeGap is less than the preset threshold; if the value of TimeGap is less than or equal to the preset threshold, the number stored in Count-Min Sketch is the total number of bytes transmitted in one round of communication phase; if the value of TimeGap is greater than the preset threshold, it means that this round of communication phase has not ended, and the size of the current data packet continues to be added to Count-Min Sketch to continuously calculate the statistics. ③ After one round of iterative training transmission, the total number of bytes of iterative training in the communication phase obtained by the switch is used in the next round In rounds of iterative training; ④ Use the total number of bytes of iterative training in the communication phase obtained by the switch in step 3) to calculate the remaining iteration time T ric ; .
[0032] In the above-mentioned distributed machine learning training method based on network flow rate control, preferably, in the step 1), if the flow ID of the job flow stored on the KV-Store and the total number of bytes of the previous iteration of each job flow do not change within a preset time, they are directly deleted or overwritten with a new job flow.
[0033] Compared with the prior art, the advantages of the present invention are: the distributed machine learning training system and method based on network flow rate control of the present invention can significantly improve the utilization rate of network links; by optimizing the communication phase of the job flow, training tasks, such as DNN training tasks, can share network resources more efficiently, thereby reducing network congestion and idle time. By giving priority to the transmission of job flows with shorter remaining iteration completion time, the average completion time of the job is effectively reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is the overall structural diagram of the distributed machine learning training system based on network flow rate control in Example 1.
[0035] Figure 2 This is a schematic diagram of the traffic information collector module in Example 1 using the Count-Min Sketch algorithm.
[0036] Figure 3 This is a schematic diagram of the traffic information collector module in Example 1 designed using KV-Store.
[0037] Figure 4 Comparison of network link utilization when the VGG model, ResNet model, and BERT model compete on the network path.
[0038] Figure 5The figure shows the overall end-to-end performance comparison of the VGG model, ResNet model, and BERT model when competing on the network path.
[0039] Figure 6 The figure is a comparison of network link utilization when MLTCP and MLSC compete on the network path.
[0040] Figure 7 The figure is a comparison of the overall end-to-end performance of MLTCP and MLSC when there is contention on the network path.
[0041] Figure 8 The figure is a comparison of network link utilization when CRUX and MLSC compete on the network path.
[0042] Fig. 9 The figure is a comparison of the overall end-to-end performance between CRUX and MLSC when there is contention on the network path. DETAILED DESCRIPTION
[0043] In order to facilitate the understanding of the present invention, the present invention will be described more comprehensively and carefully in combination with preferred embodiments below, but the protection scope of the present invention is not limited to the following specific embodiments.
[0044] It should be noted that when an element is described as being "fixed, fixed, connected or connected to" another element, it can be directly fixed, fixed, connected or connected to the other element, or it can be indirectly fixed, fixed, connected or connected to the other element through other intermediate connectors.
[0045] Unless otherwise defined, all the professional terms used below have the same meanings as those generally understood by those skilled in the art. The professional terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the scope of protection of the present invention. Example
[0046] A distributed machine learning training system based on network flow rate control, such as Figure 1 As shown, it includes a flow information collector module and a job scheduler module; the flow information collector module collects the number of bytes in the job flow communication mode and records the timestamp of the arrival of the data packet on the switch, and calculates the remaining iteration completion time of the job flow; the job scheduler module determines the priority of the job according to the remaining iteration completion time of the job flow, so that the switch gives priority to sending the job flow with the shortest remaining iteration completion time; the congestion control and rate adjustment module will mark the job flow with the remaining iteration completion time exceeding the preset time with an ECN mark or send a CNP packet to reduce its sending rate.
[0047] In this embodiment, the traffic information collector module uses the Count-Min Sketch algorithm to approximately count the total number of bytes transmitted by different job flows in each iteration training. Figure 2 As shown in the figure, two Count-Min Sketches are deployed, one for recording the number of bytes (size) and the other for recording the timestamp (time) of the arrival of the data packet; this information is used for subsequent job scheduling decisions.
[0048] At the same time, in this embodiment, the flow information collector module also adopts KV-Store design; Figure 3 As shown in the figure, KV-Store is designed to store and store the stream ID of the job stream and the total number of bytes of the previous iteration; at the same time, it calculates the remaining iteration completion time of the job stream. KV-Store also designs an aging mechanism to automatically adjust entries that have not been updated for a long time to save memory resources. The aging mechanism designed on KV-Store is as follows: if the stream ID of the job stream stored on KV-Store and the total number of bytes of the previous iteration of each job stream do not change within the preset time, they will be directly deleted or overwritten with a new job stream.
[0049] In this embodiment, the method for the flow information collector module to calculate the remaining iteration completion time of the job flow includes the following steps: ① calculating the ratio TimeGap of the time difference between the arrival of two adjacent data packets of the job flow, and determining the total number of bytes transmitted in each iteration of the job flow; ; ② Determine whether the value of TimeGap is less than the preset threshold; if the value of TimeGap is less than or equal to the preset threshold, the number stored in Count-Min Sketch is the total number of bytes transmitted in one round of communication phase; if the value of TimeGap is greater than the preset threshold, it means that this round of communication phase has not ended, and the size of the current data packet continues to be added to Count-Min Sketch to continuously calculate the statistics. ③ After one round of iterative training transmission, the total number of bytes of iterative training in the communication phase obtained by the switch is used in the next round In rounds of iterative training; ④ Use the total number of bytes of iterative training in the communication phase obtained by the switch in step 3) to calculate the remaining iteration time T ric ; .
[0050] In this embodiment, the job scheduler module adopts a job preference strategy to determine the priority transmission of the switch, and dynamically adjusts the job sending rate. The job preference strategy is as follows: the job scheduler module determines the priority of the job according to the remaining iteration completion time of the job flow. Priority is given to those jobs with the shortest remaining iteration completion time to reduce communication contention between jobs. The method for dynamically adjusting the job sending rate is as follows: according to the remaining iteration completion time ratio of the job flow, the RDMA-based DCQCN congestion control algorithm is used to dynamically adjust the sending rate of the job flow. For job flows with longer remaining iteration completion times, their sending rates are reduced by marking them with ECN marks or sending CNP packets.
[0051] In this embodiment, the job scheduler module adopts the RDMA-based DCQCN congestion control algorithm through the congestion control and rate adjustment module to control congestion by marking the data packets with ECN marks; that is, the RDMA-based DCQCN congestion control algorithm is used to dynamically adjust the sending rate of the job flow. For job flows with a long remaining iteration completion time, the sending rate is reduced by marking them with ECN marks or sending CNP packets, so as to give priority to sending the job flow with the shortest remaining iteration completion time. This method allows the system to adaptively adjust the sending rate of the job flow without knowing the job information in advance.
[0052] In this embodiment, the job scheduler module dynamically adjusts the sending rate of the job flow according to the remaining iteration completion time of the job flow and the current sending rate to achieve interleaved transmission between jobs.
[0053] This embodiment also provides a distributed machine learning training method based on network flow rate control, comprising the following steps: 1) Deploy a traffic information collector module, a job scheduler module and a congestion control and rate adjustment module on the switch; the traffic information collector module includes a KV-Store and two Count-Min Sketches; the two Count-Min Sketches, one for collecting the number of bytes in the job flow communication mode, and the other for recording the timestamp of the arrival of the data packet; if the flow ID of the job flow stored on the KV-Store and the total number of bytes of the previous iteration of each job flow do not change within a preset time, they will be directly deleted or overwritten with a new job flow.
[0054] 2) The KV-Store on the flow information collector module collects and stores the flow ID and the total number of bytes of the previous iteration of each job flow; at the same time, the remaining iteration completion time of the job flow is calculated. The specific method includes the following steps: ① Calculate the ratio of the time difference between the arrival of two adjacent data packets in the job flow, TimeGap, and determine the total number of bytes transmitted in each iteration of the job flow; ; ② Determine whether the value of TimeGap is less than the preset threshold; if the value of TimeGap is less than or equal to the preset threshold, the number stored in Count-Min Sketch is the total number of bytes transmitted in one round of communication phase; if the value of TimeGap is greater than the preset threshold, it means that this round of communication phase has not ended, and the size of the current data packet continues to be added to Count-Min Sketch to continuously calculate the statistics. ③ After one round of iterative training transmission, the total number of bytes of iterative training in the communication phase obtained by the switch is used in the next round In rounds of iterative training; ④ Use the total number of bytes of iterative training in the communication phase obtained by the switch in step 3) to calculate the remaining iteration time T ric ; .
[0055] 3) The job scheduler module determines the priority of the job according to the remaining iteration completion time of the job flow, so that the switch preferentially sends the job flow with the shortest remaining iteration completion time.
[0056] 4) The congestion control and rate adjustment module deployed on the switch will mark the job flow whose remaining iteration completion time exceeds the preset time with ECN mark or send CNP packets to reduce its sending rate, thereby dynamically adjusting the job sending rate.
[0057] 5) Repeat steps 2)-4) until all job flows have completed training.
[0058] The present invention can effectively solve the communication contention problem in distributed machine learning training, improve the bandwidth utilization of network links, reduce the average job completion time, and accelerate the execution of training jobs, such as DNN training jobs. The distributed machine learning training system and method based on network flow rate control of this embodiment is highly adaptable, flexible and hardware-friendly, and is suitable for DNN training jobs of various scales and characteristics.
[0059] In order to verify the excellent performance of the distributed machine learning training system based on network flow rate control of Example 1, the present invention provides the following experimental evaluation.
[0060] Evaluation indicators: The experimental evaluation of the present invention mainly uses network link utilization and average job completion time as evaluation indicators. Network link utilization reflects the utilization efficiency of network resources and reflects the degree of optimization of network bandwidth by the scheduling scheme; the average job completion time measures the execution efficiency of the overall job and is directly related to the training efficiency of the cluster. These two indicators can be used to comprehensively evaluate the performance of the distributed machine learning training acceleration scheduling method (MLSC) based on network flow rate control in Example 1.
[0061] Experimental environment: The experiment of the present invention is based on the NS3 network simulation platform, simulating a real GPU cluster environment. The specific experimental environment is as follows: 1) Topology: The experiment of the present invention adopts a single bottleneck dumbbell structure, including 5 groups of hosts (G1 and G6, G2 and G7, G3 and G8, G4 and G9, G5 and G10), each group of hosts is connected to the switch through a 100Gb / s link, and the link between the two switches is a bottleneck link. At the same time, the experiment adopts a hierarchical topology structure, including 6 switches and multiple bottleneck links, to evaluate the impact of multiple network bottlenecks on different DNN training tasks.
[0062] 2) DNN training task: In the experimental setting of the present invention, the test tasks include VGG, ResNet, BERT and GPT models, representing small-scale, medium-scale and large-scale DNN training tasks, respectively. The iteration time for each round is set to: 0.05ms for VGG and ResNet, 5ms for BERT, and 0.5s for GPT. Among them, the VGG (Visual Geometry Group) model is a deep convolutional neural network model proposed by the Visual Geometry Group of the University of Oxford. The VGG model is known for its simple and effective architecture. Its main feature is the use of multiple small-sized convolution kernels (usually 3x3) to build a deep network. The ResNet (Residual Network) model is a deep residual network proposed by Microsoft Research. ResNet solves the gradient vanishing problem in deep neural networks by introducing residual blocks, making the network deeper. The BERT (Bidirectional Encoder Representations from Transformers) model is a pre-trained language model proposed by Google. BERT uses a bidirectional Transformer encoder to capture contextual information, significantly improving the performance of natural language processing tasks. GPT (Generative Pre-trained Transformer) is a generative pre-trained language model proposed by OpenAI. GPT generates text through a unidirectional Transformer decoder and is suitable for generative tasks. The network link utilization and overall end-to-end performance of the VGG model, ResNet model, and BERT model when there is contention on the network path are shown in Figure 2. Figure 4 and Figure 5 shown.
[0063] Performance: The distributed machine learning training acceleration scheduling method (MLSC) based on network flow rate control in Example 1 performs well in two key indicators: network link utilization and average job completion time. Compared with the fairness network, the network link utilization of ResNet, BERT and GPT models is improved by 42%, and the average job completion time is reduced by 14% (see Figure 4). In addition, compared with MLTCP (Rajasekaran, S., Narang, S., Zabreyko, AA, et al. (2024). MLTCP: Congestion Control for DNN Training.), MLSC improves network link utilization by 12.5%, as shown in Figure 4. Figure 6; The average task completion time was improved by 45.8%, Figure 7 As shown in Figure 2, compared with CRUX (Cao, J., Guan, Y., Qian, K., et al. (2024). CRUX: GPU-EfficientCommunication Scheduling for Deep Learning Training.), MLSC improves network link utilization by 2.9%, as shown in Figure 2. Figure 8 As shown in the figure, the average job completion time was improved by 98.4%. Fig. 9 These experimental results show that the MLSC scheduling scheme significantly outperforms existing state-of-the-art work in improving network link utilization and reducing average job completion time.
Claims
1. A distributed machine learning training system based on network flow rate control, characterized in that: It includes a flow information collector module and a job scheduler module; The flow information collector module collects the number of bytes in the job flow communication mode and records the timestamp of the arrival of the data packet on the switch, and calculates the remaining iteration completion time of the job flow; The job scheduler module determines the priority of the job according to the remaining iteration completion time of the job flow, so that the switch preferentially sends the job flow with the shortest remaining iteration completion time.
2. The distributed machine learning training system based on network flow rate control according to claim 1, characterized in that: The flow information collector module deploys two Count-Min Sketches on the switch, one for collecting the number of bytes in the job flow communication mode, and the other for recording the timestamp of the arrival of the data packet.
3. The distributed machine learning training system based on network flow rate control according to claim 2, characterized in that: The flow information collector module collects and stores the flow ID of each job flow and the total number of bytes of the previous iteration of each job flow by deploying KV-Store on the switch; at the same time, it calculates the remaining iteration completion time of the job flow.
4. The distributed machine learning training system based on network flow rate control according to claim 3 is characterized in that: The method for calculating the remaining iteration completion time of the job flow comprises the following steps: ① calculating the ratio TimeGap of the time difference between the arrival of two adjacent data packets of the job flow, and determining the total number of bytes transmitted in each iteration of the job flow; ; ② Determine whether the value of TimeGap is less than the preset threshold; if the value of TimeGap is less than or equal to the preset threshold, the number stored in Count-Min Sketch is the total number of bytes transmitted in one round of communication; if the value of TimeGap is greater than the preset threshold, it means that this round of communication has not ended, and the size of the current data packet continues to be added to Count-Min Sketch to continuously collect statistics; ③After one iteration of training transmission is completed, the total number of bytes of iterative training in the communication phase obtained by the switch will be used in the next round In rounds of iterative training; ④ Use the total number of bytes of iterative training in the communication phase obtained by the switch in step 3) to calculate the remaining iteration time T ric ; 。 5. The distributed machine learning training system based on network flow rate control according to claim 3, characterized in that: The job scheduler module uses the congestion control and rate adjustment module to mark the job flow whose remaining iteration completion time exceeds the preset time with an ECN mark or send a CNP packet to reduce its sending rate, thereby giving priority to sending the job flow with the shortest remaining iteration completion time.
6. The distributed machine learning training system based on network flow rate control according to claim 3, characterized in that: If the stream ID of the job stream stored in the KV-Store and the total number of bytes of the last iteration of each job stream do not change within a preset time, they are directly deleted or overwritten with a new job stream.
7. A distributed machine learning training method based on network flow rate control, characterized in that: The following steps are involved: 1) Deploy a traffic information collector module, a job scheduler module, and a congestion control and rate adjustment module on the switch; the traffic information collector module includes a KV-Store and two Count-Min Sketches; the two Count-Min Sketches, one for collecting the number of bytes in the job flow communication mode, and the other for recording the timestamp of the arrival of the data packet; 2) The KV-Store on the flow information collector module collects and stores the flow ID and the total number of bytes of the previous iteration of each job flow; at the same time, the remaining iteration completion time of the job flow is calculated; 3) The job scheduler module determines the priority of the job according to the remaining iteration completion time of the job flow, so that the switch preferentially sends the job flow with the shortest remaining iteration completion time; 4) The congestion control and rate adjustment module deployed on the switch will mark the job flow whose remaining iteration completion time exceeds the preset time with an ECN mark or send a CNP packet to reduce its sending rate, thereby dynamically adjusting the job sending rate; 5) Repeat steps 2)-4) until all job flows have completed training.
8. The distributed machine learning training method based on network flow rate control according to claim 7 is characterized in that: The specific method of step 2) comprises the following steps: ① Calculate the ratio of the time difference between the arrival of two adjacent data packets in the job flow, TimeGap, and determine the total number of bytes transmitted in each iteration of the job flow; ; ② Determine whether the value of TimeGap is less than the preset threshold; if the value of TimeGap is less than or equal to the preset threshold, the number stored in Count-Min Sketch is the total number of bytes transmitted in one round of communication phase; ③After one iteration of training transmission is completed, the total number of bytes of iterative training in the communication phase obtained by the switch will be used in the next round In rounds of iterative training; ④ Use the total number of bytes of iterative training in the communication phase obtained by the switch in step 3) to calculate the remaining iteration time T ric ; 。 9. The distributed machine learning training method based on network flow rate control according to claim 7, characterized in that: In the step 1), if the stream ID of the job stream stored in the KV-Store and the total number of bytes of the last iteration of each job stream do not change within a preset time, they are directly deleted or overwritten with a new job stream.