Network congestion control method for distributed model training and related device

By monitoring and controlling the network bandwidth usage of model training tasks in a distributed system, the task extension problem caused by network congestion in distributed model training is solved, and more efficient network utilization and training time is achieved.

WO2025130623A1PCT designated stage expired Publication Date: 2025-06-26CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

Patent Information

Application Number
PCT/CN2024/136899
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-04
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In distributed model training, network congestion problems lead to an extended execution time of model training tasks, especially when multiple model training tasks are executed simultaneously, network congestion on the shared link increases the overall training time of the task.

Method used

By monitoring the target model training task to be executed in the distributed system, it is determined whether there is a periodic peak network bandwidth occupancy time period. Based on the pre-constructed network traffic model, the corresponding network congestion control parameters are called to control the distributed system network congestion, so that the peak network bandwidth occupancy time period of the target model training task is different from the existing model training task.

Benefits of technology

It reduces the probability of network congestion, improves network utilization, and reduces the overall training time of model training tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136899_26062025_PF_FP_ABST
    Figure CN2024136899_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of artificial intelligence. Provided are a network congestion control method and apparatus for distributed model training, a device, a medium and a computer program product. The method comprises: detecting a target model training task to be executed in a distributed system, the target model training task being a model training task having a periodic peak network bandwidth occupation time period, and the peak network bandwidth occupation time period being a task time period when the network bandwidth occupation is higher than a preset threshold in the execution process of the model training task; and, on the basis of a pre-constructed network traffic model of the distributed system and the target model training task, invoking a corresponding network congestion control parameter to perform network congestion control on the distributed system, such that the peak network bandwidth occupation time period of the target model training task is different from a peak network bandwidth occupation time period of an existing model training task in the distributed system. The present disclosure can reduce the occurrence probability of network congestion and improve the network utilization rate.
Need to check novelty before this filing date? Find Prior Art

Description

Network congestion control method and related equipment for distributed model training

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This disclosure claims priority to Chinese patent application number 202311789769.1, filed on December 22, 2023, entitled “Network congestion control method, apparatus, device and medium for distributed model training”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of artificial intelligence technology, and in particular to a network congestion control method, apparatus, device, medium, and computer program product for distributed model training. Background Art

[0004] Because machine learning applications for very large models often require extensive data computation, distributed model training is often employed as an engineering solution for implementing these applications. Distributed machine learning distributes the model training process across multiple computing nodes, with each node processing a portion of the data. This accelerates the model training process, enables the processing of large datasets, and enables the training of more complex models.

[0005] In distributed systems based on Remote Direct Memory Access (RDMA), a fairness strategy is often used for network congestion control. While RDMA technology improves the network communication performance of distributed systems, its general adaptability is not optimized for machine learning applications. For example, when the data traffic of two model training tasks simultaneously puts pressure on the network, the RDMA congestion control strategy will fairly allocate network resources to the different training tasks. However, this fair allocation also causes the two model training tasks to periodically compete for and release network resources during execution. This situation is particularly pronounced when the two model training tasks are of the same type, as network congestion in the system will cause the execution of both tasks to be relatively delayed.

[0006] In the related art, in the actual machine learning model training process, model training tasks do not require high network traffic at all times. These tasks may generate periodic high-bandwidth pressure on the system network. When multiple model training tasks appear in the system, the data synchronization phase of the same model training task may overlap periodically, resulting in periodic network congestion on the shared link, thereby extending the overall training time of each model training task. Network congestion refers to the situation where the number of packets transmitted in a packet switching network is too large, and the network transmission performance is reduced due to the limited resources of the storage and forwarding nodes, resulting in a continuously overloaded network state, which may cause additional network overhead and reduce the efficiency of model training.

[0007] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0008] The present disclosure provides a network congestion control method, apparatus, device, medium and computer program product for distributed model training, which at least to a certain extent overcome the problem of network congestion occurring during model training in related technologies.

[0009] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0010] According to one aspect of the present disclosure, a network congestion control method for distributed model training is provided, comprising: monitoring a target model training task to be executed in a distributed system, wherein the target model training task is a model training task having a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during which the network bandwidth occupancy is higher than a preset threshold during the execution of the model training task; based on a pre-constructed network traffic model of the distributed system and the target model training task, calling corresponding network congestion control parameters to perform network congestion control on the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of existing model training tasks in the distributed system.

[0011] In some embodiments, the monitoring of the target model training task to be executed in the distributed system includes: monitoring whether the distributed system receives the model training task to be executed; if the monitoring distributed system receives the model training task to be executed, determining whether the model training task to be executed has a periodic peak network bandwidth occupancy time period; if the model training task to be executed has a periodic peak network bandwidth occupancy time period, determining the model training task to be executed as the target model training task.

[0012] In some embodiments, determining whether the model training task to be executed has a periodic peak network bandwidth occupancy time period includes: determining the peak network bandwidth occupancy time period of the model training task to be executed during the execution of the model training task according to a preset network congestion control algorithm; and determining whether the model training task to be executed has a periodic peak network bandwidth occupancy time period according to the peak network bandwidth occupancy time period of the model training task to be executed.

[0013] In some embodiments, the preset network congestion control algorithm is a Data Center Quantized Congestion Notification (DCQCN) algorithm.

[0014] In some embodiments, based on the pre-built network traffic model of the distributed system and the target model training task, corresponding network congestion control parameters are called to perform network congestion control on the distributed system, including: obtaining a first execution priority and a second execution priority, wherein the first execution priority is the task execution priority of the target model training task, and the second execution priority is the task execution priority of the existing model training task; according to the first execution priority and the second execution priority, the execution order of the target model training and the existing model training task is determined, and the corresponding network congestion control parameters are called to perform network congestion control on the distributed system, so that the model training task with a higher task execution priority is executed first, and the model training task with a higher task execution priority is postponed.

[0015] In some embodiments, the pre-built network traffic model of the distributed system is a network traffic model constructed based on one or more existing model training tasks of periodic peak network bandwidth occupancy time periods in the distributed system.

[0016] In some embodiments, the calling of corresponding network congestion control parameters to perform network congestion control on the distributed system so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system includes: the calling of corresponding network congestion control parameters to perform network congestion control on the distributed system so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of any existing model training task in the distributed system.

[0017] According to another aspect of the present disclosure, a network congestion control device for distributed model training is also provided, including: a training task monitoring module, configured to monitor a target model training task to be executed in a distributed system, wherein the target model training task is a model training task with a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during the execution of the model training task during which the network bandwidth occupancy is higher than a preset threshold; a network congestion control module, configured to call corresponding network congestion control parameters to perform network congestion control on the distributed system based on a pre-constructed network traffic model of the distributed system and the target model training task, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.

[0018] According to another aspect of the present disclosure, an electronic device is also provided, which includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the network congestion control method for distributed model training described in any one of the above-mentioned instructions by executing the executable instructions.

[0019] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the network congestion control method for distributed model training described in any one of the above is implemented.

[0020] According to another aspect of the present disclosure, a computer program product is also provided, including a computer program, which, when executed by a processor, implements any of the above-mentioned network congestion control methods for distributed model training.

[0021] The network congestion control method, apparatus, device, medium, and computer program product for distributed model training provided in the embodiments of the present disclosure, after detecting the existence of a target model training task to be executed in a distributed system with a periodic peak network bandwidth occupancy period, combines a pre-built network traffic model and calls corresponding network congestion control parameters to control network congestion in the distributed system, so that the peak network bandwidth occupancy period of the target model training task is different from that of the existing model training tasks in the distributed system. The embodiments of the present disclosure can reduce the probability of network congestion and improve network utilization by automatically adjusting the network congestion parameters of the model training task according to user needs.

[0022] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0024] FIG1 shows a schematic diagram of an exemplary application system architecture of a network congestion control method for distributed model training according to an embodiment of the present disclosure;

[0025] FIG2 shows a schematic diagram of a network perception system according to an embodiment of the present disclosure;

[0026] FIG3 shows a flow chart of a network congestion control method for distributed model training according to an embodiment of the present disclosure;

[0027] FIG4 shows a flow chart of another network congestion control method for distributed model training according to an embodiment of the present disclosure;

[0028] FIG5 shows a schematic diagram of a network congestion control device for distributed model training according to an embodiment of the present disclosure;

[0029] FIG6 shows a block diagram of an electronic device according to an embodiment of the present disclosure;

[0030] FIG7 shows a schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0032] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0033] For ease of understanding, before introducing the embodiments of the present disclosure, several terms involved in the embodiments of the present disclosure are first explained as follows:

[0034] RDMA: Remote Direct Memory Access, remote direct memory access;

[0035] PFC: Priority Flow Control, is a flow control technology used to manage the priority of data flows in Ethernet networks;

[0036] DCQCN: Data Center Quality of Service Congestion Notification, a technology used for congestion notification and quality of service in data center networks.

[0037] The specific implementation of the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.

[0038] Figure 1 shows a schematic diagram of an exemplary application system architecture of a network congestion control method for distributed model training in an embodiment of the present disclosure. As shown in Figure 1, in the related art, the system architecture may include a spine device 101, a leaf device 102, a server 103, and a client 104. Among them, the spine device 101 can be connected to the first leaf device 1021 and the second leaf device 1022 respectively, and the first leaf device can be connected to the first server 1031, the second server 1032, the third server 1033, and the fourth server 1034 respectively. Taking the first server 1031 as an example here, the first server 1031 is connected to the client 104.

[0039] In the related technology, the current cluster-based distributed model training system performs end-to-end model training by supporting different frameworks (such as TensorFlow, etc.). At this time, the model training task is carried on multiple nodes of the cluster. The distributed system divides the model training task into smaller batches, performs calculations on each node, and then updates the unified model weights through a synchronization mechanism. In the scenario of face recognition, a combination of multiple graphics processors (GPUs) on multiple servers can be selected for model training tasks. During the training of commonly used distributed frameworks, the model weights are synchronized between multiple servers through the network every time the model undergoes an iteration or a fixed period. Multiple servers usually encounter situations where weights are updated using a shared network link.

[0040] In RDMA-based distributed systems, a fair strategy is used for network congestion control. However, in special scenarios (such as large model training scenarios), both the PFC-based port congestion control and the flow-based DCQCN congestion control mechanism have adjustment mechanisms that can set the network bandwidth for specific tasks (such as large model training tasks).

[0041] Therefore, in one embodiment of the present disclosure, a network perception system 105 is added to the network architecture of related technologies. Specifically, Figure 2 shows a schematic diagram of a network perception system according to an embodiment of the present disclosure. System 105 includes a traffic model construction module 1051, a network congestion perception module 1052, and a parameter memory module 1053.

[0042] In one embodiment of the present disclosure, the traffic model construction module 1051 has two user interfaces. One interface is oriented towards the current distributed system, and is mainly used to perform system topology perception and network perception of the model training task, and record the bandwidth sensitivity of each node of the model training task. The other interface is oriented towards the model training task, and searches and adapts a portion of the hyperparameters of the model training task within a preset area. For example, in a data parallel architecture, the batch size of the model training task, the working node, and other hyperparameters of the image data flow are configured with discrete values. After adapting different hyperparameter combinations, a traffic model table for the initial model training task in the system is constructed. The user can refer to the traffic model table to intervene in the actual model training plan and select the most appropriate parameter configuration under the current system topology to complete the training of the initial traffic model.

[0043] Specifically, the traffic model table can show the traffic situation of the model training task under different parameters, as shown in Table 1:

[0044] Table 1

[0045] W1 can be a model parameter group consisting of multiple default hyperparameters in the system, and TH1 and LT1 can be the default network throughput and default task latency calculated based on the current default model parameter group. W2 can be a model parameter group consisting of multiple hyperparameters for the initial model training task, and TH2 and LT2 can be the network throughput and task latency calculated based on the model parameter group for the current initial model training task. Compare the results of the two model parameter groups and select the more optimal data to complete the initial traffic model training.

[0046] It should be noted that the above-mentioned preset area is determined by the customized search range selected based on the hyperparameters of training the initial traffic model, and can be adjusted according to actual conditions. The embodiments of the present disclosure do not make specific limitations on this.

[0047] In one embodiment of the present disclosure, the network congestion perception module 1052 is used to perceive network communication traffic during the model training task execution phase, including an analysis perception phase of constructing a traffic model (i.e., an initial traffic model) for existing simulation training tasks in the system at startup, a sampling perception phase of newly accessed network traffic (i.e., a new model training task) during the execution of the model training task, and an inherent perception phase of adjusting the network congestion control parameters of the model training task.

[0048] In one embodiment of the present disclosure, the analysis and perception stage refers to the scenario in which there is a model training task in the system, the network usage rate of the model training task is perceived, and a network traffic model of the current model training task is constructed as the initial traffic model of the current system. Specifically, a preset network congestion control algorithm can be used to perceive the periodic data flow and ignore the bursty data flow. The above-mentioned preset network congestion control algorithm can refer to the DCQCN algorithm, that is, the traffic in the network link is evaluated by analyzing the number of explicit congestion notification ECN tags in the algorithm. It should be noted that any network congestion control algorithm that can perceive the periodic data flow of the model training task can be used, and the present disclosure does not make specific limitations on this.

[0049] In one embodiment of the present disclosure, after obtaining the initial traffic model, the sampling perception phase begins. When a new model training task is running in the system, a preset network congestion control algorithm is used to perceive abnormal congestion in the network traffic to determine whether the current model training task has periodic data traffic (i.e., periodic peak network bandwidth occupancy time period). If so, the system returns to the inherent perception phase and re-perceives whether there is any new model training task accessing the system.

[0050] In one embodiment of the present disclosure, the inherent perception stage is used to start the congestion parameter search function. For example, there are two model training tasks in the current system. By adjusting the RDMA network congestion control parameters of a certain node, the network occupancy rate is adjusted to the priority first training task according to a certain ratio, so that the system's network resources are more inclined to complete the first training task when congestion occurs. After adjustment, after several rounds of training iterations, the network congestion perception module determines whether the system actively postpones the time point when the second training task occupies high network bandwidth (that is, whether a concurrent periodicity suitable for the two model training tasks is found). When the network congestion perception stage recognizes that the system's congestion control reaches the optimal state for the current situation, the system reaches stability, and the completion time of the model training task is expected to be the shortest. The network occupancy rate of the two tasks on the shared link will be improved in the high occupancy demand stage, similar to shifting the phase of two periodic sine waves to avoid the area where the peaks overlap, reducing the pressure on the overall network bandwidth, and increasing the network bandwidth occupied by each.

[0051] It should be noted that the priority of the model training task can be the task execution priority pre-set by the user. It should be noted that the embodiment of the present disclosure does not make specific limitations on this.

[0052] In one embodiment of the present disclosure, the parameter memory module 1053 has two interfaces: one interface for model training tasks and the other interface for network congestion control parameters during the execution of the model training tasks. This module is used to register and store the hyperparameters of the model training tasks, the network congestion control parameters, and the network traffic model in the current system. It is also used to provide and configure the initialization values ​​of the hyperparameters of the model training tasks and the network congestion control parameters. This module can also be opened to users, making it convenient for users to set parameter information in a personalized system, thereby improving the flexibility of the system.

[0053] It should be noted that, in actual applications, the above initialization values ​​are usually provided by the system as default values, and can also be configured independently by the user, and the embodiments of the present disclosure do not specifically limit this.

[0054] In one embodiment of the present disclosure, the system can provide users with the ability to analyze application services and system network traffic, identify network traffic information for model training tasks, and help users optimize network configuration in different systems without specifically specifying the system topology of the model training job. It also combines business operations and network performance parameters, comprehensively considering overall training accuracy and training efficiency, and provides users with recommended parameters based on both training operations and network configuration when using the system, providing a two-dimensional training operation parameter model.

[0055] Those skilled in the art will appreciate that the number of spine devices, leaf devices, servers, and clients in FIG1 is merely illustrative, and any number of spine devices, leaf devices, servers, and clients may be used as needed, and the present disclosure does not limit this.

[0056] Under the above system architecture, an embodiment of the present disclosure provides a network congestion control method for distributed model training, which can be executed by any electronic device with computing and processing capabilities.

[0057] FIG3 shows a flow chart of a network congestion control method for distributed model training according to an embodiment of the present disclosure. As shown in FIG3 , the method includes the following steps:

[0058] S302, monitoring a target model training task to be executed in a distributed system, wherein the target model training task is a model training task with a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during which the network bandwidth occupancy is higher than a preset threshold during the execution of the model training task.

[0059] In one embodiment of the present disclosure, among one or more model training tasks newly connected to a distributed system, a model training task that has a periodic peak network bandwidth occupancy period is determined as a target model training task. The peak network bandwidth occupancy period may refer to a task period during which the network bandwidth occupancy of the model training task exceeds a preset threshold during task execution. It should be noted that the preset threshold can be set by the user based on actual circumstances and is not specifically limited in this embodiment of the present disclosure.

[0060] S304, based on the pre-built network traffic model of the distributed system and the target model training task, call the corresponding network congestion control parameters to control the network congestion of the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.

[0061] In one embodiment of the present disclosure, a pre-built network traffic model of a distributed system may be a network traffic model constructed based on one or more existing model training tasks of periodic peak network bandwidth occupancy time periods in the distributed system; the peak network bandwidth occupancy time period of the model training task is adjusted by calling the corresponding network congestion control parameters in the target model training task to distinguish it from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.

[0062] As can be seen from the above, after detecting the existence of a target model training task to be executed in a distributed system with periodic peak network bandwidth occupancy time periods, the embodiment of the present disclosure, in combination with a pre-built network traffic model, calls the corresponding network congestion control parameters to control network congestion in the distributed system, so that the peak network bandwidth occupancy time periods of the target model training task and the existing model training tasks in the distributed system are different. The embodiment of the present disclosure can reduce the probability of network congestion and improve network utilization by automatically adjusting the network congestion parameters of the model training task according to user needs.

[0063] In one embodiment of the present disclosure, the above S302 includes monitoring whether the distributed system receives a model training task to be executed; if the distributed system receives a model training task to be executed, determining whether the model training task to be executed has a periodic peak network bandwidth occupancy time period; if the model training task to be executed has a periodic peak network bandwidth occupancy time period, determining the model training task to be executed as the target model training task.

[0064] In one embodiment of the present disclosure, the DCQCN algorithm can be used to determine whether there is a periodic peak network bandwidth occupancy time period in the model training task to be executed. It should be noted that the embodiment of the present disclosure can also adopt any other network congestion control algorithm that can determine whether the model training task has a periodic peak network bandwidth occupancy time period, and the embodiment of the present disclosure does not make specific limitations on this.

[0065] In one embodiment of the present disclosure, it is determined whether a model training task to be executed has a periodic peak network bandwidth occupancy time period, including: determining, according to a preset network congestion control algorithm, the peak network bandwidth occupancy time period of the model training task to be executed during the execution of the model training task; and determining, according to the peak network bandwidth occupancy time period of the model training task to be executed, whether the model training task to be executed has a periodic peak network bandwidth occupancy time period.

[0066] In one embodiment of the present disclosure, the preset network congestion control algorithm is a Data Center Quantized Congestion Notification (DCQCN) algorithm.

[0067] In one embodiment of the present disclosure, the above S304 includes obtaining a first execution priority and a second execution priority, wherein the first execution priority is the task execution priority of the target model training task, and the second execution priority is the task execution priority of the existing model training task; according to the first execution priority and the second execution priority, the execution order of the target model training and the existing model training tasks is determined, and the corresponding network congestion control parameters are called to perform network congestion control on the distributed system, so that the model training tasks with higher task execution priorities are executed first, and the model training tasks with higher task execution priorities are postponed.

[0068] It should be noted that both the first execution priority and the second execution priority can be set in advance by the user, and the embodiment of the present disclosure does not specifically limit this.

[0069] In one embodiment of the present disclosure, the pre-built network traffic model of the distributed system is a network traffic model constructed based on model training tasks of one or more existing periodic peak network bandwidth occupancy time periods in the distributed system.

[0070] In one embodiment of the present disclosure, DCQCN can be used to determine whether an existing model training task has a periodic peak network bandwidth occupancy time period. It should be noted that the embodiment of the present disclosure can also adopt any other network congestion control algorithm that can determine whether a model training task has a periodic peak network bandwidth occupancy time period, and the embodiment of the present disclosure does not make specific limitations on this.

[0071] In one embodiment of the present disclosure, the above S304 includes: calling corresponding network congestion control parameters to perform network congestion control on the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of any existing model training task in the distributed system.

[0072] In one embodiment of the present disclosure, in a multi-user cloud computing business scenario, based on different businesses, the business that can maximize network performance can be virtualized to similar hardware implementations, providing users with effective network configuration and system topology recommendations.

[0073] In one embodiment of the present disclosure, under the condition that the cloud service provider provides the same resources, users can be provided with more efficient services, while also helping and guiding users to have more control over resource utilization, thereby improving their cognitive experience. This can also help cloud service providers improve the efficiency of user resource allocation. Cloud service providers can provide complementary resources to multiple users based on user needs, thus saving resource costs.

[0074] FIG4 shows a flow chart of another network congestion control method for distributed model training according to an embodiment of the present disclosure. As shown in FIG4 , the method includes the following steps:

[0075] S401, when the system starts, registers a training task in the traffic model construction module to build an initial network traffic model.

[0076] In one embodiment of the present disclosure, an initial model training task received in a distributed system is judged, and an initial network traffic model is trained based on the model training task with periodic peak network bandwidth occupancy time periods. That is, the model training task with periodic peak network bandwidth occupancy time periods is registered in a user-facing interface, while the model training task with only sudden peak network bandwidth occupancy is not registered.

[0077] In one embodiment of the present disclosure, the DCQCN algorithm can be used to determine whether an existing model training task has a periodic peak network bandwidth occupancy time period. It should be noted that the embodiment of the present disclosure can also adopt any other network congestion control algorithm that can determine whether a model training task has a periodic peak network bandwidth occupancy time period, and the embodiment of the present disclosure does not make specific limitations on this.

[0078] S402: After the initial network traffic model is built, the sampling perception phase begins.

[0079] S403: Determine whether a new model training task is connected to the current distributed system. If yes, execute S404; if not, execute S410.

[0080] In one embodiment of the present disclosure, DCQCN can be used to determine whether the network traffic execution cycle in the current distributed system has changed, so as to determine whether a new model training task is connected to the system. It should be noted that the embodiment of the present disclosure can also adopt any other network congestion control algorithm that can determine whether the model training task has a periodic peak network bandwidth occupancy time period, and the embodiment of the present disclosure does not make specific limitations on this.

[0081] S404: Sense network congestion and start a congestion sensing function.

[0082] S405: Determine whether the received new model training task has a periodic peak network bandwidth occupancy period. If so, execute S406; if not, mark the task and do not process it further, and execute S407.

[0083] In one embodiment of the present disclosure, DCQCN can be used to determine whether a model training task to be executed has a periodic peak network bandwidth occupancy time period. It should be noted that the embodiment of the present disclosure can also adopt any other network congestion control algorithm that can determine whether a model training task has a periodic peak network bandwidth occupancy time period, and the embodiment of the present disclosure does not make specific limitations on this.

[0084] S406: Search and configure parameters, and optimize the distributed system.

[0085] In one embodiment of the present disclosure, the network congestion control parameters for a model training task are adjusted based on the task execution priority preset by the user, so that the peak network bandwidth usage period of the model training task differs from the peak network bandwidth usage period of existing model training tasks in the distributed system. While maintaining the current network congestion control parameters, the hyperparameters for the model training task are searched and configured to optimize the distributed system.

[0086] S407: Store the current parameters.

[0087] S408: Determine whether the current system is optimal. If so, execute S409; if not, execute S406.

[0088] In one embodiment of the present disclosure, when multiple model training tasks are executed simultaneously in the system, the system enters a stable periodic operation. When the estimated completion time of the model training task is the shortest, it proves that the current distributed system has achieved optimal performance.

[0089] S409: Maintain the current parameter configuration.

[0090] S410: Determine whether the current system satisfies both the system termination and the number of new tasks registered in the current system is less than 1. If so, execute S411; if not, execute S402.

[0091] S411: Output the model training report to the user.

[0092] When the system ends and the number of model training tasks registered in the system is less than 1, the module completes the congestion perception task of this stage. During the model training, it organizes and analyzes the search for hyperparameters, the adjustment of network congestion parameters, and the network traffic model obtained by the final training. The above data is sent to the user as a model training report and reasonable suggestions are made.

[0093] From the above, it can be seen that after the embodiment of the present disclosure detects that there is a target model training task to be executed in a distributed system with a periodic peak network bandwidth occupancy time period, it combines the pre-built network traffic model and calls the corresponding network congestion control parameters to control the network congestion of the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from that of the existing model training tasks in the distributed system. The model training tasks can be marked and the network monitored. When executing multiple training tasks at the same time, the network priority can be adjusted with emphasis. Without affecting the accuracy of the training business, the network traffic model of the concurrent training tasks can be identified. Through customized congestion control, that is, according to the system's recommendation, the system is implicitly put into a high utilization mode to maximize the network utilization.

[0094] In one embodiment of the present disclosure, when a new model training task is connected to the current distributed system, while starting to perceive the network congestion situation, the monitoring task of the new model training task in the system can be triggered at the same time. At this time, the system continues to monitor whether there is a new model training task.

[0095] In one embodiment of the present disclosure, network congestion perception can be performed using the following formula: t =R o ×(1+g) (1)

[0096] Among them, R t Indicates the network congestion parameter group when congestion is perceived, R o represents the congestion parameter group before congestion perception, g represents the change rate of congestion control bandwidth, g∈[-1,1].

[0097] Among them, the further away the g value is from 0, the higher the rate of change, and the faster the speed of returning to high-speed bandwidth. It should be noted that the g value should also take into account the robustness of changes in the entire system; the above-mentioned network congestion parameter group is a general term for a series of network congestion parameters used to adjust the peak network bandwidth occupancy time period of the model training task, such as network bandwidth. All network congestion parameters involved in actual model training can be added to the network congestion parameter group. The embodiment of the present disclosure does not specifically limit the parameter types included in the network congestion parameter group.

[0098] In one embodiment of the present disclosure, the time difference between tasks before and after parameter search can be expressed by the following formula:

[0099] Where ΔT represents the time difference between the completion time of the model training task after the parameter search configuration and the completion time of the original model training task, T×(·) represents the completion time of the model training task under the current parameter configuration, Represents the comprehensive value of the hyperparameters of the model training task after parameter search, represents the comprehensive value of the network congestion parameter of the model training task after parameter search, P train_n_nonsense represents the comprehensive value of hyperparameters of model training before parameter search, P net_n_nonsense Represents the comprehensive value of the network congestion parameters of the model training before parameter search.

[0100] It should be noted that the above ΔT, T, P train_n_nonsense and P train_n_nonsense are all constants; the above network congestion parameters can be obtained by R at the current data collection time. t 、R oThe comprehensive value of the network congestion parameters composed of and g is a constant; the hyperparameters are used to optimize the output results after the current model training, such as batch size, working nodes, etc. The above hyperparameters can be the comprehensive value of a hyperparameter group composed of one or more hyperparameters obtained by adjusting the current network congestion parameters at the current data collection moment, which is a constant.

[0101] Based on the same inventive concept, the present disclosure also provides a network congestion control device for distributed model training, as described in the following embodiments. Since the principles of the device embodiment to solve the problem are similar to those of the above-mentioned method embodiment, the implementation of the device embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated.

[0102] FIG5 shows a schematic diagram of a network congestion control device for distributed model training in an embodiment of the present disclosure. As shown in FIG5 , the device includes: a training task monitoring module 501 and a network congestion control module 502 .

[0103] Among them, the training task monitoring module 501 is configured to monitor the target model training task to be executed in the distributed system, wherein the target model training task is a model training task with a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during the execution of the model training task when the network bandwidth occupancy is higher than a preset threshold; the network congestion control module 502 is configured to call the corresponding network congestion control parameters to perform network congestion control on the distributed system based on a pre-built network traffic model and target model training task of the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.

[0104] As can be seen from the above, after detecting the existence of a target model training task to be executed in a distributed system with periodic peak network bandwidth occupancy time periods, the embodiment of the present disclosure, in combination with a pre-built network traffic model, calls the corresponding network congestion control parameters to control network congestion in the distributed system, so that the peak network bandwidth occupancy time periods of the target model training task and the existing model training tasks in the distributed system are different. The embodiment of the present disclosure can reduce the probability of network congestion and improve network utilization by automatically adjusting the network congestion parameters of the model training task according to user needs.

[0105] In one embodiment of the present disclosure, the above-mentioned training task monitoring module 501 is also configured to monitor whether the distributed system receives a model training task to be executed; if the monitoring distributed system receives a model training task to be executed, it is determined whether the model training task to be executed has a periodic peak network bandwidth occupancy time period; if the model training task to be executed has a periodic peak network bandwidth occupancy time period, the model training task to be executed is determined as the target model training task.

[0106] In one embodiment of the present disclosure, the above-mentioned training task monitoring module 501 is also configured to determine the peak network bandwidth occupancy time period of the model training task to be executed during the execution of the model training task according to a preset network congestion control algorithm; and based on the peak network bandwidth occupancy time period of the model training task to be executed, determine whether there is a periodic peak network bandwidth occupancy time period for the model training task to be executed.

[0107] In one embodiment of the present disclosure, the preset network congestion control algorithm is a Data Center Quantized Congestion Notification (DCQCN) algorithm.

[0108] In one embodiment of the present disclosure, the above-mentioned network congestion control module 502 is also configured to obtain a first execution priority and a second execution priority, wherein the first execution priority is the task execution priority of the target model training task, and the second execution priority is the task execution priority of the existing model training task; according to the first execution priority and the second execution priority, the execution order of the target model training and the existing model training tasks is determined, and the corresponding network congestion control parameters are called to perform network congestion control on the distributed system, so that the model training tasks with higher task execution priorities are executed first, and the model training tasks with higher task execution priorities are postponed.

[0109] In one embodiment of the present disclosure, the pre-built network traffic model of the distributed system is a network traffic model constructed based on model training tasks of one or more existing periodic peak network bandwidth occupancy time periods in the distributed system.

[0110] In one embodiment of the present disclosure, the above-mentioned network congestion control module 502 is also configured to call corresponding network congestion control parameters to perform network congestion control on the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of any existing model training task in the distributed system.

[0111] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods, or program products. Therefore, various aspects of the present disclosure may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."

[0112] The electronic device 600 according to this embodiment of the present disclosure is described below with reference to Figure 6. The electronic device 600 shown in Figure 6 is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0113] FIG6 shows a block diagram of an electronic device according to an embodiment of the present disclosure. An electronic device 600 according to this embodiment of the present disclosure is described below with reference to FIG6 . The electronic device 600 shown in FIG6 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.

[0114] As shown in Figure 6, electronic device 600 is implemented as a general-purpose computing device. Components of electronic device 600 may include, but are not limited to, the aforementioned at least one processing unit 610, the aforementioned at least one storage unit 620, and a bus 630 connecting various system components (including storage unit 620 and processing unit 610).

[0115] Wherein, the storage unit stores a program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 performs the steps described in the above "Exemplary Method" section of this specification according to various exemplary embodiments of the present disclosure. For example, the processing unit 610 can perform the following steps of the above method embodiment: monitoring a target model training task to be executed in a distributed system, wherein the target model training task is a model training task with a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during the execution of the model training task when the network bandwidth occupancy is higher than a preset threshold; based on a pre-constructed network traffic model of the distributed system and the target model training task, calling the corresponding network congestion control parameters to perform network congestion control on the distributed system, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.

[0116] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .

[0117] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0118] Bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0119] The electronic device 600 can also communicate with one or more external devices 640 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can occur via an input / output (I / O) interface 650. Furthermore, the electronic device 600 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 660. As shown, the network adapter 660 communicates with other modules of the electronic device 600 via a bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 600, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0120] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0121] In particular, according to an embodiment of the present disclosure, the process described in the above reference flowchart can be implemented as a computer program product, which includes: a computer program, which implements the above-mentioned network congestion control method of distributed model training when executed by a processor.

[0122] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, which may be a readable signal medium or a readable storage medium. FIG7 shows a schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure. As shown in FIG7 , a program product 700 capable of implementing the above-mentioned method of the present disclosure is stored on the computer-readable storage medium. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above.

[0123] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] In the present disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0125] In some embodiments, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0126] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0127] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0128] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0129] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0130] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

Claims

1. A network congestion control method for distributed model training, comprising: A target model training task to be executed in the distributed system is monitored, wherein the target model training task is a model training task with a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during which the network bandwidth occupancy is higher than a preset threshold during the execution of the model training task; Based on the pre-built network traffic model of the distributed system and the target model training task, the corresponding network congestion control parameters are called to perform network congestion control on the distributed system so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.

2. The network congestion control method for distributed model training according to claim 1, wherein: The target model training task to be executed in the distributed system is monitored, including: Monitor whether the distributed system has received model training tasks to be executed; If the monitoring distributed system receives the model training task to be executed, it is determined whether there is a periodic peak network bandwidth occupancy time period for the model training task to be executed; If the model training task to be executed has a periodic peak network bandwidth occupancy time period, the model training task to be executed is determined as the target model training task.

3. The network congestion control method for distributed model training according to claim 2, wherein: Determining whether the model training task to be executed has a periodic peak network bandwidth occupancy time period includes: According to a preset network congestion control algorithm, determine the peak network bandwidth occupancy time period of the model training task to be executed during the execution of the model training task; According to the peak network bandwidth occupancy time period of the model training task to be executed, determine whether the model training task to be executed has a periodic peak network bandwidth occupancy time period.

4. The network congestion control method for distributed model training according to claim 3, wherein: The preset network congestion control algorithm is the Data Center Quantized Congestion Notification DCQCN algorithm.

5. The network congestion control method for distributed model training according to claim 1, wherein: Based on the pre-built network traffic model of the distributed system and the target model training task, calling corresponding network congestion control parameters to perform network congestion control on the distributed system includes: Obtain a first execution priority and a second execution priority, wherein the first execution priority is the task execution priority of the target model training task, and the second execution priority is the task execution priority of the existing model training task; According to the first execution priority and the second execution priority, the execution order of the target model training and the existing model training tasks is determined, and the corresponding network congestion control parameters are called to perform network congestion control on the distributed system, so that the model training tasks with higher task execution priorities are executed first, and the model training tasks with higher task execution priorities are postponed.

6. The network congestion control method for distributed model training according to claim 1, wherein: The pre-built network traffic model of the distributed system is a network traffic model built based on one or more existing model training tasks of periodic peak network bandwidth occupancy time periods in the distributed system.

7. The network congestion control method for distributed model training according to claim 6, wherein: The calling of corresponding network congestion control parameters to perform network congestion control on the distributed system so that the peak network bandwidth occupation time period of the target model training task is different from the peak network bandwidth occupation time period of the existing model training tasks in the distributed system includes: The corresponding network congestion control parameters are called to perform network congestion control on the distributed system so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of any existing model training task in the distributed system.

8. A network congestion control device for distributed model training, comprising: A training task monitoring module is configured to monitor a target model training task to be executed in a distributed system, wherein the target model training task is a model training task with a periodic peak network bandwidth occupancy time period, and the peak network bandwidth occupancy time period is a task time period during which the network bandwidth occupancy is higher than a preset threshold during the execution of the model training task; The network congestion control module is configured to call corresponding network congestion control parameters to perform network congestion control on the distributed system based on the pre-built network traffic model of the distributed system and the target model training task, so that the peak network bandwidth occupancy time period of the target model training task is different from the peak network bandwidth occupancy time period of the existing model training tasks in the distributed system.

9. An electronic device, comprising: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the network congestion control method for distributed model training as described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the network congestion control method for distributed model training as described in any one of claims 1 to 7.

11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the network congestion control method for distributed model training as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Distributed deep learning flow scheduling method, system and device

    CN111131080A

  • Model training method and device and storage medium

    CN114089889A

  • Model training task scheduling method and device and electronic equipment

    CN115220899A

  • Network congestion control method and device for distributed model training, equipment and medium

    CN117544565A

  • Dynamic network bandwidth in distributed deep learning training

    US20220012642A1

Cited By

  • Self-construction network construction method and system

    CN120639639A

  • Network quality prediction method and device for wide area network, equipment and medium

    CN120639644A