Network Traffic Simulation Method, Device, Equipment and Medium for Large Model Training

By obtaining user configuration information to define network topology structure and training parameters, generating communication load matrix, and performing traffic simulation, the problem of inflexible topology structure configuration in large-scale model training is solved, and network traffic simulation of the large-scale model training process is realized, improving applicability and accuracy.

CN119728454BActive Publication Date: 2025-07-04BEIJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510223771.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-04
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the prior art, general network simulation tools are difficult to flexibly adjust the network topology, which limits the applicability in training clusters of different scales, and cannot accurately simulate the specific communication modes and performance indicator analysis of large-scale model training.

Method used

By obtaining user configuration information, defining network topology and training parameters, generating communication load matrix, performing traffic simulation, simulating network traffic during large model training, supporting custom training parameters and parallel strategies, and using congestion control and load balancing algorithms for testing and evaluation.

Benefits of technology

The network traffic simulation of the large model training process is realized, which improves flexibility and applicability, can reflect the network communication needs in actual training, and supports diversified network configuration optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119728454B_ABST
    Figure CN119728454B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer technologies, and provides a method, apparatus, device, and medium for simulating network traffic in large model training. The method includes: obtaining configuration information of a user; defining a network topology structure and training parameters of a large model to be trained according to the configuration information; generating a communication load matrix based on the training parameters and the network topology structure, where the communication load matrix is used to characterize the computing time and data transmission requirements of each network node in the network topology structure; and performing traffic simulation according to the communication load matrix to simulate the network traffic during the model training process of the large model to be trained. By defining the network topology structure and the training parameters of the model through the configuration information of the user, the user can flexibly adjust the network structure to adapt to the simulation of the network traffic of training clusters of different scales and structures, thereby improving the flexibility and applicability of the simulation of the network traffic in model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method, device, equipment and medium for simulating network traffic in large model training. Background Art

[0002] With the rapid development of deep learning and artificial intelligence technologies, training large-scale machine learning models has become a core requirement for promoting technological progress in various industries. Such models usually contain hundreds of millions or even hundreds of billions of parameters, and the training process relies on distributed computing clusters. In the distributed training process, network communication plays a key role, responsible for data synchronization between nodes and the transfer of model parameters. Efficient network communication requires the network topology to have high bandwidth, low latency, and good load balancing capabilities to ensure that data transmission during the training process can be carried out in a timely and stable manner.

[0003] In the prior art, using general network simulation tools to simulate distributed training has achieved certain results in terms of network performance in the distributed training environment, but there are still significant deficiencies in meeting the specific requirements of large-scale model training. This is mainly reflected in the definition and configuration of the network topology structure, which is difficult for users to flexibly adjust, restricting its applicability in training clusters of different scales. Summary of the Invention

[0004] The present invention provides a method, device, equipment and medium for simulating network traffic in large model training, aiming to solve the defect that in the prior art, when using general network simulation tools to simulate distributed training, for the specific requirements of large-scale model training, it is difficult for users to flexibly adjust the definition and configuration of the network topology structure, restricting the applicability of network simulation in training clusters of different scales.

[0005] The present invention provides a method for simulating network traffic in large model training, including the following steps:

[0006] Obtain the configuration information of the user;

[0007] Define the network topology structure and the training parameters of the large model to be trained according to the configuration information;

[0008] Generate a communication load matrix based on the training parameters and the network topology structure; the communication load matrix is used to characterize the computing time and data transmission requirements of each network node in the network topology structure;

[0009] Perform traffic simulation according to the communication load matrix to simulate the network traffic during the model training process of the large model to be trained.

[0010] According to the method for simulating network traffic in large model training provided by the present invention, the defining the network topology structure and the training parameters of the large model to be trained according to the configuration information includes:

[0011] Determine the number of nodes according to the network configuration information in the configuration information, and set network nodes according to the number of nodes; the network configuration information includes node configuration information;

[0012] Configure the key parameters of the network nodes based on the node configuration information to define the network topology; the key parameters include the connection method between the network nodes, link bandwidth, and link delay;

[0013] Define the training parameters of the large model to be trained according to the model configuration information in the configuration information; the training parameters include the scale of model parameters, training parallel strategy, and training batch size.

[0014] According to the network traffic simulation method for large model training provided by the present invention, the training parameters further include the hidden layer size, the number of model layers, the number of attention heads of the attention mechanism, the hidden layer size of the feed-forward network, the vocabulary size, and the sequence length;

[0015] The training parallel strategy includes data parallelism, tensor parallelism, and pipeline parallelism;

[0016] The training batch size includes the global batch size and the micro-batch size.

[0017] According to the network traffic simulation method for large model training provided by the present invention, the generating a communication load matrix based on the training parameters and the network topology includes:

[0018] Based on the training parameters and the network topology, determine the traffic pattern of model parameter synchronization and gradient transmission during the training of the large model to be trained;

[0019] Generate a communication load matrix according to the traffic pattern; the communication load matrix includes the forward calculation duration, the forward communication duration, and the backward calculation duration.

[0020] According to the network traffic simulation method for large model training provided by the present invention, the performing traffic simulation according to the communication load matrix to simulate the network traffic during the model training process of the large model to be trained includes:

[0021] Determine a set of calculation processes and a set of communication processes according to the communication load matrix; the set of calculation processes includes a set of forward calculation processes and a set of backward calculation processes;

[0022] Input the set of calculation processes and the set of communication processes into a global pipeline scheduler for allocation; the global pipeline scheduler determines the pipeline stage according to the number of forward executions and the number of backward executions of the network nodes;

[0023] Execute the target communication process corresponding to the pipeline stage to simulate the network traffic in the model training process of the large model to be trained.

[0024] According to the network traffic simulation method for large model training provided by the present invention, the execution of the target communication process corresponding to the pipeline stage includes:

[0025] Obtain a preset algorithm set, and select a target algorithm combination from the algorithm set; the algorithm set includes multiple algorithm combinations, and each algorithm combination is obtained by combining a congestion control algorithm and a load balancing algorithm; the target algorithm combination is any one of the multiple algorithm combinations;

[0026] Based on the target algorithm combination, execute the target communication process corresponding to the pipeline stage;

[0027] Return to and execute the step of selecting the target algorithm combination from the algorithm set until the target algorithm combination is the last one in the algorithm set.

[0028] According to the network traffic simulation method for large model training provided by the present invention, after the execution of the target communication process corresponding to the pipeline stage, it further includes:

[0029] Collect performance metrics in the target communication process; the performance metrics include network performance metrics and training performance metrics; the network performance metrics include link utilization, end-to-end delay, and network throughput, and the training performance metrics include single training iteration time, training efficiency, and communication overhead ratio;

[0030] Evaluate the network topology structure and the training process of the large model to be trained according to the performance metrics to obtain performance evaluation information.

[0031] The present invention also provides a network traffic simulation device for large model training, including the following modules:

[0032] An information acquisition module, configured to acquire configuration information of a user;

[0033] A simulation configuration module, configured to define a network topology structure and training parameters of a large model to be trained according to the configuration information;

[0034] A load generation module, configured to generate a communication load matrix based on the training parameters and the network topology structure; the communication load matrix is used to characterize the computing time and data transmission requirements of each network node in the network topology structure;

[0035] A simulation and emulation module, configured to perform traffic emulation according to the communication load matrix to simulate the network traffic in the model training process of the large model to be trained.

[0036] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the network traffic simulation method for large model training as described in any one of the above is implemented.

[0037] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the network traffic simulation method for large model training as described in any one of the above is implemented.

[0038] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the network traffic simulation method for large model training as described in any one of the above is implemented.

[0039] The network traffic simulation method, device, equipment, and medium for large model training provided by the present invention define the network topology structure and the training parameters of the large model to be trained through the configuration information of the user, generate a communication load matrix based on the network topology structure and the training parameters of the large model to be trained, perform traffic simulation according to the communication load matrix, simulate the network traffic in the model training process of the large model to be trained, and realize the simulation of the network traffic in the model training process, so as to optimize and adjust the network configuration in the large model training process according to the simulation results. By defining the network topology structure and the training parameters of the model through the configuration information of the user, it is possible for the user to flexibly adjust the network structure to adapt to the simulation of the network traffic of training clusters with different scales and structures, and improve the flexibility and applicability of the simulation of the network traffic in the model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is a flowchart of the network traffic simulation method for large model training provided by the present invention.

[0042] Figure 2 It is a flowchart of the simulation process of the network traffic provided by the present invention.

[0043] Figure 3 It is a schematic structural diagram of the network traffic simulation device for large model training provided by the present invention.

[0044] Figure 4 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0045] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts fall within the protection scope of the present invention.

[0046] An embodiment of the present invention provides a method, device, equipment and medium for simulating network traffic in large model training, which can solve the deficiencies existing in the simulation of network traffic in large-scale model training by existing general network simulation tools, and solve problems such as inflexible configuration of network topology structures, inaccurate simulation of specific traffic patterns, insufficient analysis dimensions of performance indicators, and limited support for algorithm testing.

[0047] Specifically, an embodiment of the present invention provides a method for simulating network traffic in large model training, Figure 1 which is a schematic flowchart of the method for simulating network traffic in large model training provided by the present invention. As Figure 1 shown, the method includes the following steps:

[0048] Step 100: Obtain the configuration information of the user;

[0049] Step 200: Define the network topology structure and the training parameters of the large model to be trained according to the configuration information;

[0050] Step 300: Generate a communication load matrix based on the training parameters and the network topology structure; the communication load matrix is used to characterize the computing time and data transmission requirements of each network node in the network topology structure;

[0051] Step 400: Perform traffic simulation according to the communication load matrix to simulate the network traffic during the model training of the large model to be trained.

[0052] Obtain the configuration information of the user, and define the network topology structure and the training parameters of the large model to be trained according to the configuration information of the user. Among them, the configuration information of the user includes the network configuration information of the network structure and the training configuration information of the large model to be trained.

[0053] Furthermore, when defining the network topology structure and the training parameters of the large model to be trained, specifically, the network topology structure is defined according to the network configuration information, and the training parameters of the large model to be trained are defined according to the training configuration information.

[0054] The network topology structure is a network structure that supports network communication functions during the distributed training process of the large model to be trained, and is responsible for data synchronization between computing nodes and the transmission of model parameters. The network topology structure includes multiple computing nodes, which are also referred to as network nodes or communication nodes. The data transmission requirements of each network calculation are determined based on the training process of the large model to be trained, and the training process of the large model to be trained is determined according to the training parameters.

[0055] Generate a communication load matrix based on the training parameters and the network topology structure. This communication load matrix is used to characterize the computing time and data transmission requirements of each network node in the network topology structure. Among them, the computing time of the network node includes the forward computing time, communication duration, and backward computing time. The forward computing time is the computing duration corresponding to the forward computing process in the model training process, and the backward computing time is the computing duration corresponding to the backward computing process in the model training process. Optionally, the generated communication load matrix includes the communication load matrix between each network node in the network topology structure.

[0056] The forward computing time and the backward computing time are the computing durations obtained by abstractly modeling the model computing amount, and the communication duration is the communication duration obtained by modeling the bandwidth delay of the network link.

[0057] Furthermore, the forward computing process mainly includes data preparation, model application, layer-by-layer computing, and result output. The backward computing process mainly includes loss calculation, gradient calculation, parameter update, and iterative loop.

[0058] Among them, data preparation is to preprocess the sample data required for model training, such as normalization and format conversion, etc.; model application is to transfer the training sample data obtained in the data preparation stage to the model; layer-by-layer computing means that each layer in the model processes the input data to generate intermediate outputs; result output refers to the process of obtaining prediction values when the model performs tasks such as classification or regression. The output result of the model is usually one or more vectors, and this vector represents the prediction result of the model.

[0059] Further, the loss calculation is performed after the forward propagation of the model. The model network calculates the loss value, i.e., the error, based on the model output and the true labels. This loss value is the basis for measuring the gap between the model's predicted value and the true value. The gradient calculation uses the loss obtained from the forward calculation to calculate the gradients of all model parameters, i.e., the derivatives of the loss value with respect to the model parameters. This process is backpropagation. Parameter update is performed after calculating the gradients of each model parameter. The model weights and biases are updated through the gradient descent algorithm. The basic principle of gradient descent is that if the gradient of the loss value (i.e., the error) points in a certain direction, then the model parameters are adjusted in the opposite direction to reduce the loss value of the model. The iterative loop is that during the model training process, the model performs multiple forward propagations, calculates the loss, backpropagates the error, and updates the model weights and biases. A complete process of forward propagation and backpropagation is called one iteration, i.e., one "epoch". The model training process usually goes through multiple epochs until the model converges to a lower loss value or reaches the preset maximum number of iterations.

[0060] Perform traffic simulation according to the generated communication load matrix to simulate the network traffic during the model training process of the large model to be trained. The process of performing traffic simulation includes traffic generation and injection, data collection and analysis, and result verification and optimization phases.

[0061] In the traffic generation and injection phase, mainly according to the data synchronization and parameter transfer processes during the training of the large model to be trained, corresponding data packets and sending intervals are generated, and the data packets are sent according to this sending interval, so as to simulate the network traffic of data synchronization and parameter transfer during the training process of the large model to be trained.

[0062] In the data collection and analysis phase, the performance metrics during the training process of the large model to be trained are collected to determine whether the network performance and model performance meet the requirements, so as to adjust the network configuration according to the simulation results during the actual training process of the large model.

[0063] In this embodiment, the network topology structure and the training parameters of the large model to be trained are defined through the user's configuration information. A communication load matrix is generated based on the network topology structure and the training parameters of the large model to be trained. Traffic simulation is performed according to this communication load matrix to simulate the network traffic during the model training process of the large model to be trained, realizing the simulation of the network traffic during the model training process, so as to optimize and adjust the network configuration during the large model training process according to the simulation results. Defining the network topology structure and the training parameters of the model through the user's configuration information allows the user to flexibly adjust the network structure to adapt to the simulation of the network traffic of training clusters with different scales and structures, improving the flexibility and applicability of the simulation of the network traffic during model training.

[0064] Optionally, the user's configuration information is determined according to the user's requirements for large model training, including network configuration information and model configuration information. The network configuration information is used to configure the network topology structure, and the model configuration information is used to configure the training parameters of the large model to be trained. Based on this, step 200 may further include:

[0065] Step 201, determine the number of nodes according to the network configuration information in the configuration information, and set network nodes according to the number of nodes; the network configuration information includes node configuration information;

[0066] Step 202, configure the key parameters of the network nodes based on the node configuration information to define the network topology structure; the key parameters include the connection method between the network nodes, link bandwidth, and link delay;

[0067] Step 203, define the training parameters of the large model to be trained according to the model configuration information in the configuration information; the training parameters include the scale of model parameters, training parallel strategy, and training batch size.

[0068] The network configuration information includes the number of nodes and the node configuration information of each node. When defining the network topology structure, the number of nodes is determined according to the network configuration information in the configuration information, and then network nodes are set according to the number of nodes. Among them, the network configuration information also includes the node configuration information of each set node. Based on this node configuration information, the key parameters of each network node are configured, thereby defining the network topology structure.

[0069] Optionally, the key parameters of the network nodes at least include the connection method between the network nodes, link bandwidth, and link delay. The network topology structure includes multiple network nodes, and the key parameters of different network nodes may be the same or different.

[0070] Furthermore, define the training parameters of the large model to be trained according to the model configuration information in the configuration information, and the training parameters include the scale of model parameters, training parallel strategy, and training batch size.

[0071] Optionally, the training parameters further include the hidden layer size, the number of model layers, the number of attention heads of the attention mechanism, the hidden layer size of the feed-forward network, the vocabulary size, and the sequence length. The training parallel strategy includes data parallelism, tensor parallelism, and pipeline parallelism, and the training batch size includes the global batch size and the micro-batch size.

[0072] General network simulation tools lack sufficient accuracy and flexibility when simulating complex and specific communication patterns during the training process of large models, such as parameter synchronization and gradient transmission, resulting in the simulation results being difficult to comprehensively reflect the network communication requirements in actual training. In this embodiment, a communication load matrix is generated according to the traffic patterns of parameter synchronization and gradient transmission during the model training process, so that the simulation of network traffic can reflect the network communication requirements in the actual model training process.

[0073] Based on this, step 300 may further include:

[0074] Step 301, determining the traffic patterns of model parameter synchronization and gradient transmission during the training process of the large model to be trained based on the training parameters and the network topology;

[0075] Step 302, generating a communication load matrix according to the traffic patterns; the communication load matrix includes the forward calculation duration, the forward communication duration, and the backward calculation duration.

[0076] Based on the training parameters of the large model to be trained and the defined network topology, determine the traffic patterns of model parameter synchronization and gradient transmission during the training process of the large model to be trained, and then generate a communication load matrix according to the traffic patterns.

[0077] The generated communication load matrix includes the forward calculation duration, the forward communication duration, and the backward calculation duration. Among them, the communication load matrix includes the forward calculation duration of each traffic flow, the forward communication duration of each traffic flow, and the backward calculation duration of each traffic flow. The communication load matrix is generated according to the network topology and the training parameters of the large model to be trained. The communication load matrix represents the calculation duration of each network node and the data transmission requirements between network nodes, reflecting the specific traffic patterns of model parameter synchronization and gradient transmission during the model training process.

[0078] Optionally, when simulating the network traffic during the model training process of the large model to be trained, it is implemented based on the global pipeline scheduler to allocate the forward calculation process and the communication process during the model training process. Step 400 may further include:

[0079] Step 401, determining a set of calculation processes and a set of communication processes according to the communication load matrix; the set of calculation processes includes a set of forward calculation processes and a set of backward calculation processes;

[0080] Step 402, inputting the set of calculation processes and the set of communication processes into the global pipeline scheduler for allocation; the global pipeline scheduler determines the pipeline stages according to the forward execution times and the backward execution times of the network nodes;

[0081] Step 403: Execute the target communication process corresponding to the pipeline stage to simulate the network traffic during the model training process of the large model to be trained.

[0082] When simulating network traffic, determine the set of computing processes and the set of communication processes according to the generated communication load matrix, and input the set of computing processes and the set of communication processes into the global pipeline scheduler for allocation. The global pipeline scheduler determines the pipeline stage according to the number of forward executions and backward executions of network nodes.

[0083] Among them, the set of computing processes includes the set of forward computing processes and the set of backward computing processes. The set of computing processes also includes the execution status of each network node, which is used to represent whether the next step of the network node is to perform forward communication or backward communication. The global pipeline scheduler determines the number of forward executions and backward executions of network nodes according to the set of computing processes, so as to determine the pipeline stage. The number of forward executions of a network node can be determined according to the number of forward computing processes in the set of computing processes executed by the network node, and the number of backward executions of a network node can be determined according to the number of backward computing processes in the set of computing processes executed by the network node.

[0084] Furthermore, execute the target communication process corresponding to the pipeline stage to simulate the network traffic during the model training process of the large model to be trained. The target communication process is the communication process corresponding to the pipeline stage in the set of communication processes.

[0085] Optionally, the forward computing set contains the forward computing processes of each traffic flow, the communication process set contains the communication processes of each traffic flow, and the global pipeline scheduler allocates the forward computing processes and communication processes in sequence according to the model training process of the large model to be trained. The pipeline stage includes data loading, warm-up, single forward and single backward, cooling, and gradient synchronization, etc.

[0086] General network simulation tools are mainly based on general communication protocols, do not support customizing training parameters and batches of large models, are difficult to capture the communication characteristics of high-frequency words and small data volumes in large model training, and lack built-in support for training-related performance indicators. In order to simulate specific communication requirements in large-scale training, secondary development and customization are usually required, which increases the difficulty and development cost of network traffic simulation. The network traffic of large models is affected by training parallel strategies, and there are some special collective communication modes and traffic patterns that cannot be reproduced by general network simulation tools. Moreover, some congestion control algorithms or load balancing algorithms and other network tuning algorithms cannot be tested after traffic simulation. In terms of performance analysis dimensions, it mainly focuses on the network layer, lacks comprehensive performance evaluation during training, and the testing and evaluation of different tuning algorithms are also limited, which restricts the optimization of traffic simulation in the network transport layer.

[0087] Based on this, in the network traffic simulation method provided by the embodiments of the present invention, during the network traffic simulation process, the test and evaluation of different algorithms can be realized. When executing the communication process, congestion control algorithms and load balancing algorithms are adopted. Specifically, in step 403, when executing the target communication process corresponding to the pipeline stage, it may further include:

[0088] Step 413, obtain a preset algorithm set, and select a target algorithm combination from the algorithm set; the algorithm set includes multiple algorithm combinations, and each of the algorithm combinations is obtained by combining a congestion control algorithm and a load balancing algorithm; the target algorithm combination is any one of the multiple algorithm combinations;

[0089] Step 423, based on the target algorithm combination, execute the target communication process corresponding to the pipeline stage;

[0090] Step 433, return and execute the step of selecting a target algorithm combination from the algorithm set; until the target algorithm combination is the last one in the algorithm set.

[0091] When executing the communication process to simulate the network traffic in the model training process, first obtain a preset algorithm set, which includes multiple algorithm combinations, and any one of the algorithm combinations is obtained by combining at least one congestion control algorithm and one load balancing algorithm. Select a target algorithm combination from the algorithm set, and the target algorithm combination is any one of the multiple algorithm combinations in the algorithm set.

[0092] In one embodiment, based on multiple different congestion control algorithms and multiple different load balancing algorithms, the congestion control algorithms and the load balancing algorithms are combined pairwise to obtain multiple algorithm combinations, forming an algorithm set.

[0093] Further, based on the selected target algorithm combination, execute the target communication process corresponding to the pipeline stage, return and execute the step of selecting a target algorithm combination from the algorithm set until the selected target algorithm combination is the last one in the algorithm set.

[0094] Optionally, based on the same communication process, based on different algorithm combinations, execute the same communication process respectively. In this way, for each communication process, execute each communication process based on different algorithm combinations respectively; or, based on the same algorithm combination, execute all communication processes, and then select the next algorithm combination and repeat to execute all communication processes. In this way, for each algorithm combination, execute all communication processes based on this algorithm combination respectively.

[0095] In the large-scale distributed training process, due to memory limitations, it is necessary to introduce three-dimensional hybrid parallel technology. Affected by the three-dimensional hybrid parallel technology, the traffic pattern of the training iteration is solidified and divided into stages such as data loading, preheating, single forward and single reverse, cooling and gradient synchronization. In the network traffic simulation, this embodiment implements the pipeline operation process of calculating each batch of data blocks in turn as the batch data is input during the training iteration, and realizes tensor parallel communication during the calculation process, and executes point-to-point communication to pass to the next computing node. In this process, the end-to-end network performance of each flow is collected and visualized to evaluate the network performance. In addition, the analysis dimension of the performance indicators is mainly concentrated at the network level, lacking the problem of comprehensive performance evaluation of the training process. This embodiment can realize the comprehensive evaluation of the model training process based on network performance and model performance.

[0096] Based on this, in step 423, after executing the target communication process corresponding to the pipeline stage based on the target algorithm combination, the following may also be included:

[0097] Step 404, collecting performance indicators during the target communication process; the performance indicators include network performance indicators and training performance indicators; the network performance indicators include link utilization, end-to-end delay and network throughput, and the training performance indicators include single training iteration time, training efficiency and communication overhead ratio;

[0098] Step 405: Evaluate the network topology and the training process of the large model to be trained according to the performance index to obtain performance evaluation information.

[0099] The performance indicators of the target communication process are collected, including network performance indicators and training performance indicators. The network performance indicators further include link utilization, end-to-end delay and network throughput. The training performance indicators include single training iteration time, training efficiency and communication overhead ratio.

[0100] The network topology and the training process of the large model to be trained are evaluated according to the collected performance indicators to obtain corresponding performance evaluation information. Optionally, based on the performance evaluation information, the user's configuration information can be optimized to achieve the optimization of the network topology for model training.

[0101] In one embodiment, referring to Figure 2 The simulation process of network traffic shown in the figure first obtains the user's configuration information, including network configuration information and training configuration information, defines the network topology according to the network configuration information, including setting the number of nodes, and configuring the link bandwidth and link delay between nodes, etc., and configures the training parameters of the large model to be trained according to the training configuration information, including setting the model parameter scale, training parallel strategy and training batch size.

[0102] Based on the configured network topology and training parameters, a communication load matrix is generated, and then the simulation environment is initialized. According to the algorithm combination of the congestion control algorithm and the load balancing algorithm, network traffic simulation is performed to simulate the network traffic during the model training of the large model to be trained. Monitor and collect the performance metrics of the network traffic during the simulation process. The performance metrics include network performance metrics and training performance metrics. The network performance metrics are used to evaluate the network performance of model training, and the training performance metrics are used to evaluate the training performance of model training.

[0103] In one embodiment, the network configuration information includes the number of nodes and further includes the configuration information of the network topology, specifically including the number of GPUs within a single server of the network topology, the number of servers, the number of network switches, the number of network links, the server type, and the switch number. Generally, for the setting of network nodes, 8 network nodes are connected to an in-chassis switch, and these 8 nodes are respectively connected to the leaf switches. The link bandwidth is also configurable within a certain range to meet the simulation of network environments from small clusters to ultra-large-scale data centers. The link latency is also configurable within a certain range, for example, it is configurable between 0.001 ms and 10 ms, which is used to simulate links of different lengths and thus reproduce the cluster distribution under different geographical locations.

[0104] In some embodiments, the discrete event simulation (DES) method is used to accurately simulate the transmission process of data packets in the network, and record the end-to-end completion time of the traffic, link occupancy, bandwidth utilization, etc. When the pipeline executes the communication process, it determines which stage of the pipeline it is currently in according to the number of forward executions and backward executions of the current network node, and then executes the corresponding communication process at that stage. In any communication process, obtain the communication mode and the next state of the current network node, as well as the next network node. Among them, the communication mode includes no communication, one-way communication, two-way communication, and sleep waiting.

[0105] Furthermore, when executing the communication process, if the communication mode of the current network node is no communication, the next node of the current network node is backward calculation; if the communication mode of the current network node is one-way communication, a data packet is sent to the next network node, and the next state of the current network node is backward calculation; if the communication mode of the current network node is two-way communication, the current network node and the previous network node send data packets to each other, and the next state of the current network node is forward calculation; if the communication mode of the current network node is sleep waiting, wait for the data packet to arrive, and the next state of the current network node is backward calculation. At any moment during the model training process, the communication modes of different nodes in the network topology can be the same or different.

[0106] Optionally, after the simulation of network traffic is completed, a simulation report can be generated. The simulation report includes the collected performance metrics and the evaluation information of the user configuration information based on the performance metrics. The evaluation information can be used to guide the user to adjust the configuration information.

[0107] In this embodiment, by customizing the configuration of the network topology and the training parameters of the model, it is possible to support diversified simulation of network traffic during the model training process. Moreover, based on the training parallel strategy, a three-dimensional hybrid parallel training process for large model training can be realized, which conforms to the traffic pattern of large model training, is conducive to the accurate and flexible simulation of complex and specific communication patterns during large model training, and ensures that the simulation results can comprehensively reflect the network communication requirements in actual model training.

[0108] Furthermore, during the simulation of network traffic, the end-to-end performance metrics of each traffic flow during the simulation process can be visually presented, so as to obtain accurate network performance metrics and model training performance metrics during the model training process without a real cluster. And it supports the testing of load balancing algorithms and congestion control algorithms in the network transport layer, as well as the switching test of different algorithms, so as to obtain the performance metrics of different algorithms during large-scale distributed training, with strong adaptability.

[0109] The following describes the network traffic simulation device for large model training provided by the present invention. The network traffic simulation device for large model training described below can be correspondingly referred to the network traffic simulation method for large model training described above.

[0110] Refer to Figure 3 , the network traffic simulation device for large model training provided by the embodiment of the present invention includes:

[0111] An information acquisition module 10, configured to acquire the configuration information of the user;

[0112] A simulation configuration module 20, configured to define a network topology structure and training parameters of a large model to be trained according to the configuration information;

[0113] A load generation module 30, configured to generate a communication load matrix based on the training parameters and the network topology structure; the communication load matrix is used to characterize the computing time and data transmission requirements of each network node in the network topology structure;

[0114] A simulation module 40, configured to perform traffic simulation according to the communication load matrix and simulate the network traffic during the model training process of the large model to be trained.

[0115] In one embodiment, the simulation configuration module 20 is further configured to:

[0116] Determine the number of nodes according to the network configuration information in the configuration information, and set network nodes according to the number of nodes; the network configuration information includes node configuration information;

[0117] Configure the key parameters of the network nodes based on the node configuration information to define the network topology; the key parameters include the connection method between the network nodes, link bandwidth, and link delay;

[0118] Define the training parameters of the large model to be trained according to the model configuration information in the configuration information; the training parameters include the model parameter scale, training parallel strategy, and training batch size.

[0119] In one embodiment, the training parameters further include the hidden layer size, the number of model layers, the number of attention heads of the attention mechanism, the hidden layer size of the feed-forward network, the vocabulary size, and the sequence length;

[0120] The training parallel strategy includes data parallelism, tensor parallelism, and pipeline parallelism;

[0121] The training batch size includes the global batch size and the micro-batch size.

[0122] In one embodiment, the load generation module 30 is further configured to:

[0123] Based on the training parameters and the network topology, determine the traffic pattern of model parameter synchronization and gradient transmission during the training of the large model to be trained;

[0124] Generate a communication load matrix according to the traffic pattern; the communication load matrix includes the forward calculation duration, forward communication duration, and backward calculation duration.

[0125] In one embodiment, the simulation module 40 is further configured to:

[0126] Determine a set of calculation processes and a set of communication processes according to the communication load matrix; the set of calculation processes includes a set of forward calculation processes and a set of backward calculation processes;

[0127] Input the set of calculation processes and the set of communication processes into a global pipeline scheduler for allocation; the global pipeline scheduler determines the pipeline stage according to the number of forward executions and backward executions of the network nodes;

[0128] Execute the target communication process corresponding to the pipeline stage to simulate the network traffic during the model training process of the large model to be trained.

[0129] In one embodiment, the simulation module 40 is further configured to:

[0130] Obtain a preset algorithm set, and select a target algorithm combination from the algorithm set; the algorithm set includes multiple algorithm combinations, and each algorithm combination is obtained by combining a congestion control algorithm and a load balancing algorithm; the target algorithm combination is any one of the multiple algorithm combinations;

[0131] Based on the target algorithm combination, execute the target communication process corresponding to the pipeline stage;

[0132] Return and execute the step of selecting the target algorithm combination from the algorithm set until the target algorithm combination is the last one in the algorithm set.

[0133] In one embodiment, the simulation module 40 is further configured to:

[0134] Collect performance metrics in the target communication process; the performance metrics include network performance metrics and training performance metrics; the network performance metrics include link utilization, end-to-end delay, and network throughput, and the training performance metrics include single training iteration time, training efficiency, and communication overhead ratio;

[0135] Evaluate the network topology structure and the training process of the large model to be trained according to the performance metrics to obtain performance evaluation information.

[0136] Figure 4 Illustrates a schematic physical structure diagram of an electronic device, as Figure 4 shown. The electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the steps of the network traffic simulation method for large model training, such as including:

[0137] Obtain the configuration information of the user;

[0138] Define the network topology structure and the training parameters of the large model to be trained according to the configuration information;

[0139] Generate a communication load matrix based on the training parameters and the network topology structure; the communication load matrix is used to characterize the computing time and data transmission requirements of each network node in the network topology structure;

[0140] Execute traffic simulation according to the communication load matrix to simulate the network traffic in the model training process of the large model to be trained.

[0141] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0142] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the steps of the network traffic simulation method for large model training provided by the above-mentioned various methods, for example, including:

[0143] Obtain the configuration information of the user;

[0144] Define the network topology structure and the training parameters of the large model to be trained according to the configuration information;

[0145] Generate a communication load matrix based on the training parameters and the network topology structure; the communication load matrix is used to characterize the calculation time and data transmission requirements of each network node in the network topology structure;

[0146] Perform traffic simulation according to the communication load matrix to simulate the network traffic during the model training process of the large model to be trained.

[0147] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the steps of the network traffic simulation method for large model training provided by the above-mentioned various methods, for example, including:

[0148] Obtain the configuration information of the user;

[0149] Define the network topology structure and the training parameters of the large model to be trained according to the configuration information;

[0150] Generate a communication load matrix based on the training parameters and the network topology structure; the communication load matrix is used to characterize the calculation time and data transmission requirements of each network node in the network topology structure;

[0151] Perform traffic simulation according to the described communication load matrix to simulate the network traffic during the model training process of the large model to be trained.

[0152] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0153] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A network traffic simulation method for large model training, characterized in that It includes: Obtain the configuration information of the user; Define the network topology structure and the training parameters of the large model to be trained according to the configuration information; The training parameters include the model parameter scale, the training parallel strategy, and the training batch size; Generate a communication load matrix based on the training parameters and the network topology structure; the communication load matrix is used to characterize the computing time and data transmission requirements of each network node in the network topology structure; Perform traffic simulation according to the communication load matrix to simulate the network traffic during the model training process of the large model to be trained; The generating of the communication load matrix based on the training parameters and the network topology structure includes: Based on the training parameters and the network topology structure, determine the traffic patterns of model parameter synchronization and gradient transmission during the training process of the large model to be trained; Generate a communication load matrix according to the traffic pattern; the communication load matrix includes the forward computing duration, the forward communication duration, and the backward computing duration.

2. The network traffic simulation method for large model training according to claim 1, characterized in that, The defining of the network topology structure and the training parameters of the large model to be trained according to the configuration information includes: Determine the number of nodes according to the network configuration information in the configuration information, and set network nodes according to the number of nodes; the network configuration information includes node configuration information; Configure the key parameters of the network nodes based on the node configuration information to define the network topology structure; the key parameters include the connection method between the network nodes, the link bandwidth, and the link delay; Define the training parameters of the large model to be trained according to the model configuration information in the configuration information.

3. The network traffic simulation method for large model training according to claim 1, wherein, The training parameters further include the hidden layer size, the number of model layers, the number of attention heads of the attention mechanism, the hidden layer size of the forward feedback network, the vocabulary size, and the sequence length; The training parallel strategy includes data parallelism, tensor parallelism, and pipeline parallelism; The training batch size includes the global batch size and the micro batch size.

4. The network traffic simulation method for large model training according to claim 1, characterized in that The performing of the traffic simulation according to the communication load matrix to simulate the network traffic during the model training process of the large model to be trained includes: Determine the set of computing processes and the set of communication processes according to the communication load matrix; the set of computing processes includes the set of forward computing processes and the set of backward computing processes; Input the set of computing processes and the set of communication processes into the global pipeline scheduler for allocation; the global pipeline scheduler determines the pipeline stages according to the number of forward executions and the number of backward executions of the network nodes; Execute the target communication process corresponding to the pipeline stage to simulate the network traffic during the model training process of the large model to be trained.

5. The network traffic simulation method for large model training according to claim 4, characterized in that, The executing of the target communication process corresponding to the pipeline stage includes: Obtain a preset set of algorithms, and select a target algorithm combination from the set of algorithms; the set of algorithms includes multiple algorithm combinations, and each algorithm combination is obtained by combining a congestion control algorithm and a load balancing algorithm; the target algorithm combination is any one of the multiple algorithm combinations; Based on the target algorithm combination, execute the target communication process corresponding to the pipeline stage; Return to and execute the step of selecting a target algorithm combination from the set of algorithms until the target algorithm combination is the last one in the set of algorithms.

6. The method for simulating network traffic in large model training according to claim 4, wherein, After executing the target communication process corresponding to the pipeline stage, it further includes: Collect performance metrics during the target communication process; the performance metrics include network performance metrics and training performance metrics; the network performance metrics include link utilization, end-to-end delay, and network throughput, and the training performance metrics include single training iteration time, training efficiency, and communication overhead ratio; Evaluate the network topology structure and the training process of the large model to be trained according to the performance metrics to obtain performance evaluation information.

7. A network traffic simulation device for large model training, characterized in that, It includes: An information acquisition module for acquiring the configuration information of the user; A simulation configuration module for defining the network topology structure and the training parameters of the large model to be trained according to the configuration information; The training parameters include the model parameter scale, training parallel strategy, and training batch size; A load generation module for generating a communication load matrix based on the training parameters and the network topology structure; the communication load matrix is used to characterize the computing time and data transmission requirements of each network node in the network topology structure; A simulation module for performing traffic simulation according to the communication load matrix to simulate the network traffic during the model training process of the large model to be trained; The load generation module is further used for: The generating a communication load matrix based on the training parameters and the network topology structure includes: Based on the training parameters and the network topology structure, determining the traffic pattern of model parameter synchronization and gradient transmission during the training process of the large model to be trained; Generating a communication load matrix according to the traffic pattern; the communication load matrix includes forward calculation duration, forward communication duration, and backward calculation duration.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the network traffic simulation method for large model training as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the network traffic simulation method for large model training as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Network system simulation method and related device

    CN115618532A