Network traffic modeling method and simulation system for large model training cluster
Through network traffic modeling methods for large-model training clusters, precise modeling and simulation analysis, the problem of incomplete network traffic modeling in large-model distributed training is solved, and effective optimization and cost reduction of network scheduling strategies are achieved.
Patent Information
- Application Number
- CN202510578519.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-07
AI Technical Summary
In the existing technology, network traffic modeling in large-scale distributed training is incomplete, resulting in high cost of verification of traffic scheduling strategies, and poor comparability and consistency of experimental results, which cannot fully reflect the actual performance under different network conditions.
It provides a network traffic modeling method for large-scale training clusters. By accurately modeling the generation and transmission characteristics of network traffic during training, combining the cluster's computing resources and network topology, it provides efficient quantitative analysis and simulation tools for network communication traffic. This method includes obtaining basic parameter information, process modeling, computing modeling and communication traffic modeling, and constructing a computing and communication quantization model of distributed training processes.
Through precise modeling and simulation analysis, it can provide a theoretical basis for optimizing network scheduling strategies in large-scale model training, effectively verify different network architectures and traffic scheduling algorithms, realize the quantification of distributed training process calculation and communication and the evaluation of traffic scheduling strategies, significantly reducing verification costs.
Smart Images

Figure CN120128489A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data center networks, and particularly relates to a network traffic modeling method and simulation system for large model training clusters. Background Art
[0002] During the training process of large-scale deep learning models, the demand for computing resources is increasing day by day. Especially in a distributed training environment, the communication requirements between nodes have a crucial impact on the overall training efficiency. As the model scale expands, the communication bottleneck in the training process becomes more prominent. Therefore, building a high-performance network for intelligent computing has become an urgent problem to be solved.
[0003] Currently, network traffic scheduling algorithms mostly rely on real environments for experimental verification. However, this method not only consumes a huge amount but may also affect the universality of the results due to the complexity of actual deployment conditions. To solve this problem, some research has turned to simulation verification. However, due to differences between different simulation tools and platforms, the comparability and consistency of experimental results are poor. In addition, most of the current network traffic involved in large model training uses intuitive estimation or simple modeling methods, lacking accurate network traffic quantification and simulation tools. This makes the strategies for optimizing training efficiency often have certain limitations in actual applications and cannot fully reflect the actual performance under different network conditions. And various parallel strategies mainly relied on in large model distributed training often require collaborative training through multiple computing nodes. These training tasks often involve large-scale data synchronization and transmission, frequent parameter or gradient synchronization, and network load imbalance, which makes the scheduling and optimization of network traffic a research hotspot. In view of the above challenges, it is particularly important to propose a network traffic modeling method and simulation system for large model training clusters. By accurately modeling the characteristics such as the generation and transmission of network traffic during the training process, it can provide a theoretical basis for optimizing network scheduling strategies in large model training. Through simulation analysis, researchers can effectively verify different network architectures and traffic scheduling algorithms, thereby realizing the quantification of computing and communication in the distributed training process and the evaluation of the effectiveness of traffic scheduling strategies. Summary of the Invention
[0004] The purpose of the present invention is to provide a network traffic modeling method and simulation system for large model training clusters in view of the deficiencies of the prior art. It aims to solve problems such as imperfect network traffic modeling and high verification costs of traffic scheduling strategies in large model distributed training in the prior art; by accurately modeling the characteristics such as the generation and transmission of network traffic during the training process, combined with the computing resources and network topology of the cluster, it provides an efficient quantitative analysis and simulation tool for network communication traffic.
[0005] To achieve the above object, the present invention provides a network traffic modeling method for a large model training cluster, including the following steps: (1) Obtain basic parameter information related to model training, including the number of input samples and the model structure; obtain performance parameter information of the devices to be adopted; (2) Process modeling: Model the process of the target distributed training strategy to construct the training process of the large model in a hybrid parallel environment of data parallelism, tensor parallelism, and pipeline parallelism; (3) Computational modeling: Model each computing node in the distributed training of the large model, analyze the computational processes of each layer during training, including the input layer, Transformer decoder, and output layer; meanwhile, quantify the computational load and computational time of each computing node based on the number of input samples, the scale of model parameters, and the training process. (4) Communication traffic modeling: Analyze the data transmission requirements of each training process according to the hybrid parallel configuration, and quantify the data exchange volume between communication nodes; combine different parallel strategies and communication modes to quantify the communication volume and communication time of each communication node.
[0006] Further, in step (2), according to the set hybrid parallel strategy, simulate the distributed training process: Implement data parallelism in multiple rows of servers, and each row of servers represents the training process of a single model replica; Implement tensor parallelism inside each server, and a single server contains multiple GPUs, and the training tasks are split and distributed to these GPUs according to the tensor parallelism degree; Implement pipeline parallelism on each row of servers, and a single server in each row of servers represents a pipeline stage.
[0007] Further, in step (3), the computational amount in the distributed training process of the large model is concentrated in the multi-head attention layer, the feed-forward network layer, and the fully connected layer.
[0008] Further, the calculation of the multi-head attention layer includes query, key, value matrix mapping, attention calculation, and multi-head splicing; the calculation of the feed-forward network layer includes three linear transformations and bias term addition, one Swish transformation, and one element-wise multiplication; the calculation of the fully connected layer includes one matrix multiplication and one softmax transformation; the model computational amount is: (computational amount of the multi-head attention layer + computational amount of the feed-forward network layer) × number of Transformer block layers + computational amount of the fully connected layer; estimate the computational time for model training according to the model computational amount.
[0009] Further, step (4) includes the following sub-steps: (4.1)Quantization from the perspective of data parallelism: Copy the model parameters and optimizer states to multiple GPUs, and then evenly distribute the training data to these GPUs; The communication generated by data parallelism mainly occurs during the gradient and parameter synchronization process between different model replicas. Use the AllReduce communication mode to calculate the single communication volume between GPU cards; (4.2)Quantization from the perspective of tensor parallelism: Split the parameter matrices of the Transformer block layer and perform parallel calculations on multiple GPUs. The communication caused by tensor parallelism mainly involves synchronizing the parameter or gradient matrices of different model parts; Use the AllReduce mode for data exchange and calculate the single communication volume between GPU cards; (4.3)Quantization from the perspective of pipeline parallelism: Allocate the parameters of different layers of the model to different GPUs. The communication generated by pipeline parallelism mainly occurs during the context interaction process between different stages of the model. Use the point-to-point communication mode to calculate the single communication volume between GPU cards; (4.4)Based on the above communication volume and combined with the bandwidth information of relevant devices, estimate the single communication time consumption of model training.
[0010] To achieve the above object, the present invention also provides a network traffic simulation system for large model training clusters, including: Input component module: including a parameter input component and a cluster configuration component; The parameter input component is used to obtain the total number of model parameters and related training parameters, including the weight matrices of each Transformer layer and the optimizer state; The cluster configuration component is used to set the computing power standard and training strategy of the cluster, including the configuration of each dimension in hybrid parallelism; Computing and communication quantization module: used to perform quantization analysis on each computing node and communication node during the distributed training process by using the information obtained and set by the input component module, so as to estimate the computing requirements and communication requirements of each node; Network optimization module: including a network topology simulation sub-module, a network status monitoring sub-module, and a traffic scheduling sub-module; The network topology simulation sub-module is used to simulate the network topology of the distributed training cluster; The network status monitoring sub-module is used to monitor the network status; The traffic scheduling sub-module is used to execute the traffic scheduling strategy to achieve congestion control in the cluster network; Effect evaluation module: used to evaluate the training and communication time consumption based on the quantization results of computing and communication and the traffic scheduling strategy implemented by the network optimization module, and analyze the optimization effect.
[0011] To achieve the above object, the present invention also provides a network traffic modeling device for a large model training cluster, including a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, the network traffic modeling method for the large model training cluster as described above is implemented.
[0012] To achieve the above object, the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the network traffic modeling method for the large model training cluster as described above is implemented.
[0013] To achieve the above object, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the network traffic modeling method for the large model training cluster as described above is implemented.
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: By accurately modeling the characteristics of network traffic generation and transmission during the training process, it can provide a theoretical basis for optimizing the network scheduling strategy in large model training. Through simulation analysis, different network architectures and traffic scheduling algorithms can be effectively verified, so as to realize the quantification of calculation and communication in the distributed training process and the evaluation of the effect of the traffic scheduling strategy. Generally speaking, the present invention can not only provide important technical support for large model training, but also provide a powerful tool for application optimization in related fields. Compared with the traditional physical machine training method, the present invention provides a powerful tool for evaluating the network traffic scheduling algorithm in large model training, and significantly reduces the high economic cost of verifying the traffic scheduling algorithm in the actual large model distributed training cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is a flowchart of the method of the present invention; Figure 2 is an architecture diagram of the present invention; Figure 3 is a schematic diagram of the distributed training process of the present invention; Figure 4 is a schematic diagram of the system of the present invention; Figure 5 is a structural diagram of the device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The present invention will be described in detail below with reference to the accompanying drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0018] In the scenario of large model distributed training, the present invention combines the Transformer architecture with a hybrid parallel strategy to design a modeling quantization method for computing and communication, so as to implement a network traffic simulation system for an intelligent computing cluster for large model distributed training.
[0019] As Figure 1 and Figure 2 shown, the present invention provides a network traffic modeling method for a large model training cluster, and the method includes the following steps: (1) Obtain relevant information such as cluster scale, model parameters, and training configuration: Obtain basic parameters related to model training, including the number of input samples, model structure, vocabulary dimension, and dimension information of each parameter matrix (such as batch size, embedding dimension, input sequence length, etc.). At the same time, obtain the performance parameter information of the devices to be used and the hybrid parallel strategy adopted; among them, the performance parameter information of the devices to be used includes GPU card computing power (i.e., the floating-point computing power of the GPU (Graphics Processing Unit)), cluster network topology, link bandwidth (i.e., the end-to-end communication bandwidth of the interconnected devices), etc.; the hybrid parallel strategy adopted includes data parallelism, tensor parallelism, and pipeline parallelism; and obtain information such as the expected overall cluster scale and GPU parallel dimensions in each parallel dimension. This information will be used for subsequent computing and communication modeling and support the quantitative analysis of the modeling results.
[0020] (2) Process modeling: Perform process modeling on the target distributed training strategy to construct the training process of a large-scale model in a hybrid parallel environment such as data parallelism, tensor parallelism, and pipeline parallelism.
[0021] Refer to Figure 3 , perform process modeling on the distributed training strategy to be adopted. According to the set hybrid parallel strategy, simulate the distributed training process: Implement data parallelism in multiple rows of servers, and each row of servers represents the training process of a single model copy; Implement tensor parallelism inside each server, and a single server contains multiple GPUs, and the training tasks are split according to the tensor parallelism degree and assigned to these GPUs; Implement pipeline parallelism on each row of servers, and a single server in each row of servers represents a pipeline stage.
[0022] (3)Computational Modeling: Analyze the computational processes of each layer during the training process based on the model-related information, especially the computations of the input layer, Transformer decoder, and output layer. At the same time, quantify the computational load and computational time of each computing node based on factors such as the number of input samples, model parameter scale (i.e., the neuron scale of each layer of the neural network, the size of the weight matrix, etc.), and training steps (i.e., the task division and training process of the large model based on the hybrid parallel strategy in the distributed intelligent computing cluster, etc.).
[0023] Specifically, model each computing node in the distributed training of the large model, analyze the specific computational processes of each layer in the entire training process, including the input layer, Transformer decoder, and output layer. Estimate the computational power requirements of each computing node during the training process based on the basic parameter information related to model training obtained in step (1). The computational volume in the distributed training process of the large model is concentrated in the multi-head attention layer (Multi-Head Attention, MHA), feed-forward network layer (Feed Forward Network, FFN), and fully connected layer. Here, the Decoder architecture (the Decoder part in the Transformer architecture) model is used as an example for analysis.
[0024] (3.1)The computation of the multi-head attention layer includes query, key, value matrix mapping, attention calculation, and multi-head concatenation.
[0025] Query, key, value matrix mapping , where represents the input matrix of the multi-head attention layer,[[]] represents the query weight matrix in the multi-head attention layer,[[]] represents the key weight matrix in the multi-head attention layer,[[]] represents the value weight matrix in the multi-head attention layer,[[]] , , the operations involved are b, s, h matrix multiplication of dimensions and h, d dimensions, and the amount of computation is 3×2bshd = 6bshd; b represents the input batch size batch size, that is, the number of sequences included in a batch; s represents the input sequence length, h represents the embedding dimension; d represents the dimension of a single head in the multi-head attention layer, (d = h / n), where n represents the number of heads in the multi-head attention layer.
[0026] Single-head attention calculation , the calculations involved include 、 And , where The computational cost of scaling and transformation is 4 bs 2 , which is much smaller than other computational costs and can be ignored. Additionally, the computational cost of the two multiplication operations involving query, key, and value matrices is 4bs 2 d; represents the tensor parallelism degree, represents the data parallelism degree.
[0027] Multi - head concatenation , where represents the mapping weight matrix after concatenating multiple attention heads, , and the computational cost involved in matrix concatenation can be ignored. The operations involved are b, s, h - dimensional matrix multiplication with h, h - dimensional matrix, and the computational cost is 2bsh 2 .
[0028] Then the total computational cost of the multi - head attention layer is: (6bshd + 4bs 2 d)×n + 2bsh 2 = 8bsh 2 + 4bs 2 h.
[0029] (3.2) The calculation of the feed - forward network layer FFN includes three linear transformations, bias term addition, one Swish transformation, and one element - wise multiplication.
[0030]
[0031] Among them, represent the weight matrix and bias term of the first linear transformation of the SwiGLU activation function respectively, represent the weight matrix and bias term of the second linear transformation of the SwiGLU activation function respectively, represent the weight matrix and bias term of the second - layer linear transformation of the feed - forward network layer respectively, , represents the hidden layer dimension of the feed - forward network layer. The computational costs of bias term addition, Swish transformation, and element - wise multiplication are 2bsh’ + bsh, 3bsh’, and bsh’ respectively, which can be ignored. The computational cost of the three linear transformations occupies the main part of the computational cost of the feed - forward network layer, specifically matrix multiplication, and the total computational cost is 3×2bshh’ = 6bshh’.
[0032] (3.3) The fully - connected layer calculation includes one matrix multiplication and one softmax transformation.
[0033]
[0034] Among them, represents the input matrix of the last layer, represents the mapping weight matrix that maps the input matrix to the vocabulary dimension, , and the involved operation is b, s, h matrix multiplication of dimensions and h, V dimensions, and the amount of computation is 2bshV; V represents the size of the input mapping vocabulary.
[0035] (3.4) According to the above analysis of the amount of computation, it can be known that the total amount of computation of the model is (8bsh 2 + 4bs 2 h + 6bshh’) × L + 2bshV, where L represents the number of Transformer blocks. Combining with the computing power information of relevant devices, the computing time consumption of model training is initially estimated; the computing time consumption is equal to the total amount of computation divided by the total computing power, where the total computing power is equal to the product of the GPU utilization rate, the computing power of a single GPU, and the number of GPUs.
[0036] (4) Communication traffic modeling: Analyze the data transmission requirements of each training step according to the hybrid parallel configuration, and quantify the data exchange volume between communication nodes, including parameter transmission, gradient synchronization, model update, etc. Combining different parallel strategies and communication modes, quantify the communication volume and communication time of each communication node.
[0037] Analyze the data transmission requirements of each training step according to the hybrid parallel configuration, and analyze its specific communication volume from the perspective of different parallel methods. Here, it can be quantified from the perspectives of data parallelism, pipeline parallelism, and tensor parallelism respectively. The specific analysis process is as follows: (4.1) Data parallelism is a method to improve training throughput. It copies the model parameters and optimizer states to multiple GPUs, and then evenly distributes the training data to these GPUs. The communication generated by data parallelism mainly occurs during the gradient and parameter synchronization between different model replicas, and the communication mode adopted is the AllReduce communication mode. The single communication volume between GPU cards can be expressed as: , where G = D at this time. Among them, 4h 2 represents the number of parameters of the 4 weight matrices (W Q , W Q , W Q and W O ) to be updated in the MHA layer; 3hh’ represents the number of parameters of the 3 weight matrices (W U , W G and W D ) to be updated in the FFN layer; Indicates the number of Transformer blocks in a single card; Indicates the data volume transmitted by Ring AllReduce data sharding; 2Bytes indicates that the data is transmitted with 16-bit precision, and the size of each parameter is 2Bytes. Among them, Indicates the hidden layer dimension of the feed-forward network layer, Indicates the tensor parallelism, Indicates the pipeline parallelism, Indicates the number of Transformer blocks, Indicates the number of GPUs in the current parallel mode (usually equal to the parallelism), and D represents the data parallelism.
[0038] (4.2) Tensor parallelism divides the parameter matrix of the Transformer block layer and performs parallel computing on multiple GPUs, thus making more efficient use of the computing power of GPUs. The communication caused by tensor parallelism mainly involves synchronizing the parameter or gradient matrices of different model parts, and the communication method uses the AllReduce mode for data exchange. The single communication volume between GPU cards can be expressed as: , where G = T. Among them, Indicates the size of the activation values (forward propagation) / gradients (backward propagation) parameters to be synchronized between GPU cards; Indicates the number of Stages in a single server; Indicates the data volume transmitted by Ring AllReduce data sharding; 2Bytes indicates that the size of each parameter is 2Bytes. Among them, Indicates the micro-batch size divided in pipeline parallelism.
[0039] (4.3) Pipeline parallelism aims to allocate the parameters of different layers of the large language model to different GPUs. In practice, consecutive layers of the Transformer can be loaded onto the same GPU to reduce the cost of transmitting hidden states or gradients between GPUs. The communication generated by pipeline parallelism mainly occurs during the context interaction between different stages of the model, and the communication method uses a point-to-point communication mode. The single communication volume between GPU cards can be expressed as: . Among them Indicates the size of the activation value parameters to be transmitted between pipeline Stages; Indicates the tensor parallelism within a single pipeline Stage; 2Bytes indicates that the size of each parameter is 2Bytes.
[0040] (4.4) Based on the above traffic analysis and combined with the bandwidth information of relevant devices, preliminarily estimate the communication time consumption for each single model training; the communication time consumption is equal to the traffic divided by the bandwidth.
[0041] Build a simulation system based on the quantization characteristics of calculation and communication in steps (3) and (4), and modularize each quantization process, so as to simulate the calculation and communication processes of large model distributed training.
[0042] Traffic scheduling strategy evaluation: Simulate the effects of different traffic scheduling strategies in the cluster network environment through the simulation system, set multiple scheduling strategies and obtain corresponding quantization results, so as to analyze and evaluate the optimization effects of each strategy.
[0043] Specifically, based on the built simulation system, add a network optimization module to further simulate the traffic scheduling strategy in the cluster network environment. This module includes a network topology simulation sub-module, a network status monitoring sub-module, and a traffic scheduling sub-module, which are respectively used to simulate the network topology, monitor the real-time traffic characteristics in the network, and implement different traffic scheduling strategies. By real-time monitoring the congestion situation in the network, this module can execute the traffic scheduling strategy in a timely manner, so as to achieve effective congestion control.
[0044] Corresponding to the embodiment of the network traffic modeling method for large model training clusters described above, the present application also provides a network traffic simulation system for large model training clusters.
[0045] Figure 4 It is a schematic structural diagram of a network traffic simulation system for large model training clusters. Refer to Figure 4 , this system may include: Input component module: including a parameter input component and a cluster configuration component. The parameter input component is used to obtain the total number of model parameters and related training parameters, including the weight matrix of each Transformer layer and the optimizer state, etc.; the cluster configuration component is used to set the computing power standard and training strategy of the cluster, including the computing power of GPU (Graphics Processing Unit) cards, link bandwidth, and configurations of each dimension in hybrid parallelism, etc.; Calculation and communication quantization module: Use the information obtained and set by the input component module to perform quantization analysis on each computing node and communication node in the distributed training process, so as to estimate the computing requirements and communication requirements of each node; Network optimization module: including a network topology simulation sub-module, a network status monitoring sub-module, and a traffic scheduling sub-module. The network topology simulation sub-module is used to simulate the network topology of the distributed training cluster; the network status monitoring sub-module is responsible for monitoring the network status, recording interface statistics, bandwidth usage, and network activity flows, etc.; the traffic scheduling sub-module is used to execute various traffic scheduling strategies to achieve congestion control in the cluster network.
[0046] Effect evaluation module: Based on the quantization results of computing and communication and the traffic scheduling strategy implemented by the network optimization module, evaluate the training and communication time consumption, and analyze the optimization effect.
[0047] Regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0048] See Figure 5 , an apparatus for network traffic modeling for a large model training cluster provided by an embodiment of the present invention includes a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement the method for network traffic modeling for a large model training cluster in the above embodiments.
[0049] An embodiment of the apparatus for network traffic modeling for a large model training cluster provided by the present invention can be applied to any device with data processing capabilities. The any device with data processing capabilities can be a device or apparatus such as a computer. The apparatus embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful apparatus, it is formed by the processor of any device with data processing capabilities where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for running. From the hardware level, as Figure 5 shown, it is a hardware structure diagram of any device with data processing capabilities where the apparatus for network traffic modeling for a large model training cluster provided by the present invention is located. In addition to Figure 5 the shown processor, memory, network interface, and non-volatile memory, any device with data processing capabilities where the apparatus in the embodiment is located usually further includes other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated herein.
[0050] The implementation processes of the functions and roles of each unit in the above apparatus are specifically detailed in the implementation processes of the corresponding steps in the above method, and will not be elaborated herein.
[0051] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative work.
[0052] An embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the network traffic modeling method for a large model training cluster in the above embodiments is implemented.
[0053] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.
[0054] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the network traffic modeling method for a large model training cluster described above is implemented.
[0055] The above embodiments are used to explain the present invention, rather than limit the present invention. Any modifications and changes made within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A network traffic modeling method for a large model training cluster, characterized in that: The following steps are involved: (1) Obtain basic parameter information related to model training, including the number of input samples and model structure; Obtain performance parameter information of the equipment to be adopted; (2) Process modeling: Perform process modeling on the target distributed training strategy and construct the training process of the large model in a hybrid parallel environment of data parallelism, tensor parallelism, and pipeline parallelism. (3) Computational modeling: Model each computing node in the distributed training of large models and analyze the computational process of each layer during the training process, including the input layer, Transformer decoder, and output layer. At the same time, quantify the computational load and computational time of each computing node based on the number of input samples, model parameter scale, and training process. (4) Communication traffic modeling: Analyze the data transmission requirements of each training process based on the hybrid parallel configuration and quantify the amount of data exchange between each communication node. Combine different parallel strategies and communication modes to quantify the communication volume and communication time of each communication node.
2. The network traffic modeling method for large model training cluster according to claim 1 is characterized in that: In step (2), according to the set hybrid parallel strategy, the distributed training process is simulated: data parallelism is implemented in multiple rows of servers, and each row of servers represents the training process of a single model replica; Tensor parallelism is implemented within each server. A single server contains multiple GPUs, and the training tasks are divided according to tensor parallelism and distributed to these GPUs; pipeline parallelism is implemented in each row of servers. A single server in each row of servers represents a pipeline stage.
3. The network traffic modeling method for large model training cluster according to claim 1 is characterized in that: In step (3), the amount of computation in the large model distributed training process is concentrated in the multi-head attention layer, feedforward network layer and fully connected layer.
4. The network traffic modeling method for large model training cluster according to claim 3 is characterized in that: The calculation of the multi-head attention layer includes query, key, value matrix mapping, attention calculation and multi-head splicing; the calculation of the feedforward network layer includes three linear transformations and bias term addition, one Swish transformation and one element-by-element multiplication; the calculation of the fully connected layer includes one matrix multiplication and one softmax transformation; the model calculation amount is: (the calculation amount of the multi-head attention layer + the calculation amount of the feedforward network layer) × the number of Transformer block layers + the calculation amount of the fully connected layer; based on the model calculation amount, the computational time of model training is estimated.
5. The network traffic modeling method for large model training cluster according to claim 1 is characterized in that: Step (4) includes the following sub-steps: (4.1) Quantification from the perspective of data parallelism: copy the model parameters and optimizer states to multiple GPUs, and then evenly distribute the training data to these GPUs; the communication generated by data parallelism is mainly the process of synchronizing the gradients and parameters between different model copies, and the AllReduce communication mode is used to calculate the single communication volume between GPU cards; (4.2) Quantification from the perspective of tensor parallelism: The parameter matrix of the Transformer block layer is split and calculated in parallel on multiple GPUs. The communication caused by tensor parallelism mainly involves synchronizing the parameters or gradient matrices of different model parts; the AllReduce mode is used for data exchange, and the single communication volume between GPU cards is calculated; (4.3) Quantification from the perspective of pipeline parallelism: The parameters of different layers of the model are assigned to different GPUs. The communication generated by pipeline parallelism is mainly in the context interaction process between different stages of the model. The point-to-point communication mode is used to calculate the single communication volume between GPU cards. (4.4) Based on the above communication volume and the bandwidth information of related devices, estimate the single communication time of model training.
6. A network traffic simulation system for large model training clusters, characterized in that: include: Input component module: includes parameter input component and cluster configuration component; The parameter input component is used to obtain the total model parameters and related training parameters, including the weight matrix and optimizer status of each Transformer layer; the cluster configuration component is used to set the computing power standard and training strategy of the cluster, including the configuration of each dimension in hybrid parallelism; Computation and communication quantification module: used to use the information obtained and set by the input component module to perform quantitative analysis on each computing node and communication node in the distributed training process, so as to estimate the computing demand and communication demand of each node; Network optimization module: including a network topology simulation submodule, a network status monitoring submodule and a traffic scheduling submodule; the network topology simulation submodule is used to simulate the network topology of the distributed training cluster; the network status monitoring submodule is used to monitor the network status; the traffic scheduling submodule is used to execute the traffic scheduling strategy to achieve congestion control in the cluster network; Effect evaluation module: It is used to evaluate the training and communication time consumption and analyze the optimization effect based on the quantitative results of computing and communication and the traffic scheduling strategy implemented by the network optimization module.
7. A network traffic modeling device for a large model training cluster, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, it implements the network traffic modeling method for a large model training cluster as described in any one of claims 1-5.
8. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the network traffic modeling method for a large model training cluster as described in any one of claims 1 to 5 is implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the network traffic modeling method for a large model training cluster as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Distributed training method for large-scale deep neural network
CN113515370A
Parallel strategy search method for efficient training of artificial intelligence large model
CN116680301A
Distributed training task allocation method based on communication demand and network resource matching
CN117636004A
Resource allocation method, device and equipment for large model cluster, storage medium and program product
CN119597368A
Method for estimating large language model training and reasoning time consumption
CN119719708A
Cited By
Heterogeneous cluster resource allocation method and storage medium
CN121542065A