Network topology optimization selection method and system for large model training scene

By extracting the structural parameters and deployment information of the big model, performing simulation optimization, and dynamically adjusting the network topology selection, the problem of extended training time in big model training is solved, and training efficiency and resource utilization are improved.

CN120200918APending Publication Date: 2025-06-24SHENZHEN INST OF ADVANCED TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510270783.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

During the training process of large-scale model, the increase in model parameter index leads to an extended training time. The existing technology lacks effective methods to optimize training topology selection, affecting training efficiency and resource utilization.

Method used

By extracting the structural parameters and deployment information of the large model, coarse-grained and fine-grained simulations are performed, training time reference values ​​under different network topology are generated, and topology selection strategies are dynamically adjusted to optimize training efficiency.

Benefits of technology

It significantly improves the efficiency and speed of model training, rationally allocates computing resources and network resources, improves overall resource utilization, and adapts to different application scenarios and training needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200918A_ABST
    Figure CN120200918A_ABST
Patent Text Reader

Abstract

The invention discloses a large model training scene-oriented network topology optimization selection method and system, and belongs to the technical field of large models. The method comprises the following steps: extracting structural parameters of a target large model, and analyzing to obtain model information; determining a distributed training strategy and a batch size, and analyzing requirements for network bandwidth and time delay to obtain deployment information; performing coarse-grained simulation on the target large model to generate training time reference values under different network topologies; performing fine-grained simulation based on the training time reference value, and verifying topology performance to obtain a simulation result; and comparing simulation results according to dynamic changes of the model information and / or deployment information, selecting an optimal network topology, and quantifying influence weights of different factors on topology selection. By systematically analyzing the influence of model information and deployment information on network topology selection in a model training scene, an effective method and way are provided for optimizing the training efficiency, improving the resource utilization rate and enhancing the system flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a method and system for optimizing network topology selection for large model training scenarios, belonging to the technical field of large models. Background Art

[0002] In recent years, due to the processing, processing, and application requirements of large amounts of data such as text and streaming media from all walks of life, the development of unsupervised, self-learning, fast-response, and high-precision generative large language models has been promoted, and they have been gradually and widely applied to task scenarios in fields such as computer vision, natural language processing, and text dialogue. The requirements for the accuracy and response time of the generated results of large models have become increasingly strict, resulting in the continuous expansion of large models and an exponential increase in model parameters, posing challenges to model training. How to optimize the model training time has become a major bottleneck in the development of large models.

[0003] The development of large models starts from the basic model (i.e., a smaller model parameter) and gradually expands the model parameters (i.e., model information includes: hidden layer dimension, context length of the input tensor, number of model layers, etc.). There are already relatively reliable experiments showing that different large models have a selection tendency for the underlying network topology due to their model structure and the sparsity of the model structure during the training scenario, but there are few studies on the influence of changes in model information on model training topology selection. At the same time, the model training time is affected by model deployment information (i.e., batch size Batch / Mirco-batch and distributed strategies such as TP / PP). Summary of the Invention

[0004] In view of the limitations of the prior art, the present application proposes a method and system for optimizing network topology selection for large model training scenarios, which provides effective methods and ways to optimize training efficiency, improve resource utilization, and enhance system flexibility by systematically analyzing the influence of model information and deployment information on network topology selection in model training scenarios.

[0005] To solve the above technical problems, the technical solutions adopted in the present application are as follows: In a first aspect, the present application provides a method for optimizing network topology selection for large model training scenarios, including: Extracting the structural parameters of the target large model, including the hidden layer dimension, context length of the input tensor, and number of model layers, quantifying the influence of the structural parameters on network communication requirements, and analyzing to obtain model information; Determining the distributed training strategy and batch size, analyzing the requirements for network bandwidth and latency, and obtaining deployment information; Performing coarse-grained simulation on the target large model to generate training time reference values under different network topologies; performing fine-grained simulation based on the training time reference values to verify the topology performance and obtain simulation results; According to the dynamic changes of model information and / or deployment information, compare the simulation results and select the optimal network topology.

[0006] As a further improvement of this application, the target large model includes a dense model and a sparse model: The dense model includes a large model based on the Transformer architecture; The sparse model includes a mixture-of-experts large model; The extraction of the structural parameters of the target large model is achieved by extending the network topology optimization tool to support the Embedding operator and the MOE model, and extracting the structural parameters.

[0007] As a further improvement of this application, the coarse-grained simulation includes: constructing a task graph of model operators by simulating the data flow of distributed training; dynamically generating communication tasks according to the parallel strategy and adding routing dependencies for the Allreduce operation; The fine-grained simulation includes: simulating the impact of packet transmission delay, packet loss rate, and congestion control on the training time based on the TCP protocol.

[0008] As a further improvement of this application, in the extraction of the structural parameters of the target large model, the priority order is: Number of model layers > Hidden layer dimension > Context length; When the number of model layers exceeds the threshold, preferentially select the Fattree-like topology; In the determination of the distributed training strategy and batch size, the priority order is: Degree of parallelism > Batch size; When the degree of parallelism exceeds the threshold, preferentially select the TopoOpt-like topology.

[0009] As a further improvement of this application, it further includes: when both the model information and / or deployment information change, calculate the comprehensive influence coefficient of the two by affecting the weights, and dynamically adjust the topology selection strategy; output the optimal network topology configuration scheme under the current training scenario.

[0010] As a further improvement of this application, perform coarse-grained simulation on the target large model to generate reference training time values under different network topologies; perform fine-grained simulation based on the reference training time values to verify the topology performance and obtain the simulation results, including: According to the type of the target large model, measure the forward and backward propagation times of the target large model operators, measure the propagation times according to possible parallel strategies, use the propagation times as candidates for the optimal parallel strategies of the operators, and generate the first result file after measuring all the operators; After setting the topology type of the cluster, calculate the operator calculation time of the first result file, randomly select a parallel scheme for each operator, calculate according to the data flow direction to obtain the current simulation training time, use the training time as the optimal training time reference time of the current topology, and generate a second result file at the same time; Use the first result file and the second result file as the input of the simulator. The second result file generated by the selected topology type corresponds to the topology selection type, and perform fine-grained training time simulation to verify the topology performance and obtain the simulation result.

[0011] In a second aspect, the present application provides a network topology optimization selection method system for large model training scenarios, including: A model information analysis module, configured to extract the structural parameters of the target large model, including the hidden layer dimension, the input tensor context length, and the number of model layers, quantify the impact of the structural parameters on the network communication requirements, and analyze to obtain model information; A deployment information analysis module, configured to determine the distributed training strategy and batch size, analyze the requirements for network bandwidth and latency, and obtain deployment information; A coarse / fine-grained simulation module, configured to perform coarse-grained simulation on the target large model to generate training time reference values under different network topologies; perform fine-grained simulation based on the training time reference values to verify the topology performance and obtain the simulation result; An optimal network topology selection module, configured to compare the simulation results according to the dynamic changes of the model information and / or deployment information, select the optimal network topology, and quantify the influence weights of different factors on the topology selection.

[0012] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the network topology optimization selection method for large model training scenarios.

[0013] In a fourth aspect, the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the network topology optimization selection method for large model training scenarios.

[0014] In a fifth aspect, the present application provides a computer program product, where the computer program product includes computer instructions, and the computer instructions direct a computer to execute the network topology optimization selection method for large model training scenarios.

[0015] The beneficial effects of the present application compared with the prior art are: A network topology optimization selection method for large model training scenarios provided by this application can significantly improve the efficiency and speed of model training by determining the optimal training topologies under different model information and deployment information. This is particularly important for large model training and real-time applications. The selection of the optimal topology means that computing resources and network resources can be more reasonably allocated, thereby improving the overall resource utilization rate. This method allows for flexible adjustment of the training topology according to different models and requirements, so as to adapt to different application scenarios and training needs. Through the method of controlling variables and simulation experiments, the performance of different topologies can be effectively evaluated without actually deploying a large amount of hardware resources. This greatly reduces the experimental cost and time cost. This method is not only applicable to specific models such as GPT2, Llama2, and MOE, but can also be extended to other types of models and training scenarios. In addition, with the emergence of new models and information technologies, this method can also be adjusted and optimized accordingly. Therefore, this application provides an effective method and approach for optimizing training efficiency, improving resource utilization rate, and enhancing system flexibility by systematically analyzing the influence of model information and deployment information on network topology selection in model training scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following introduces the accompanying drawings related to the technical solutions in the embodiments of this application or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in this application. For those skilled in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of a network topology optimization selection method for large model training scenarios provided by this application; Figure 2 It is a flowchart of the network topology optimization selection method for large model training scenarios given in the embodiments of this application; Figure 3 It is part of the experimental results of this application; Figure 4 It is a device for a network topology optimization selection method for large model training scenarios provided by this application; Figure 5 It is a schematic diagram of an electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0019] In the description of the present application, unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.

[0020] The present application mainly proposes a method for determining the optimal network topology based on the model information or (and) deployment information of the large language model based on the fine-grained simulation training time. It develops based on the TopoOpt open-source project, which in turn is based on the Flexflow project (the technology focuses on searching for the optimal parallel strategy of deep learning models) for several deep learning networks such as DLRM, VGG, ResNet, NCF, CANDLE, etc. By iteratively optimizing the network topology and parallel strategy of the cluster, the training time of the model is minimized.

[0021] The deep learning models mainly targeted by the TopooOpt project are not difficult to analyze. Their network structures are relatively simple, the number of model layers is small, and the number of model parameters is small, resulting in relatively small network traffic. Now, deep learning technology has developed to large models for various processing and application scenarios based on massive data. Whether its iterative search optimization is applicable to the characteristics of complex model structures, massive parameters, and ultra-large traffic remains to be verified.

[0022] On the other hand, the vertical update of the large model includes but is not limited to changes in the model structure such as changing operators, optimizing operators, fusing operators, adjusting the operator order, and changes in model information such as changing the context length; the horizontal update is mainly changes in model information. Taking Llama1 and Llama2 as examples for vertical updates, Llama2 designed the Rotary Position Embedding (RoPE) encoding and the RMSNorm operator for computational simplification, and at the same time adjusted the order of the RMSNorm operator; taking the Llama2 model as an example for horizontal updates, it changed the model parameters and provided options of sizes such as 7B, 13B, 64B, etc. The larger the number of parameters, the higher the accuracy of the model generation results, but at the same time, the training time of the model is longer. Currently, there are few or even no methods for determining the optimal topology selection in the large model training scenario for changes in model information.

[0023] At the same time, there are mature distributed training technology frameworks. For example, Microsoft's Deepspeed framework adopts the ZERO technology, Google's DistBelief framework adopts the DP MP distributed strategy, and the TeraPipe framework adopts the PP distributed strategy training framework, etc. It shows that the distributed strategy will affect the training time. Based on this phenomenon, when searching, there are few or even no methods found for determining the optimal topology selection in the large model training scenario based on the changes in model deployment information.

[0024] As Figure 1 shown, a network topology optimization selection method for large model training scenarios of the present application includes: S100, extracting the target large model structure parameters, including the hidden layer dimension, the input tensor context length, and the number of model layers, quantifying the influence of the structure parameters on the network communication requirements, and analyzing to obtain the model information; S200, determining the distributed training strategy and batch size, analyzing the requirements for network bandwidth and latency, and obtaining the deployment information; S300, performing coarse-grained simulation on the target large model to generate training time reference values under different network topologies; performing fine-grained simulation based on the training time reference values to verify the topology performance and obtain the simulation results; S400, according to the dynamic changes of the model information and / or deployment information, comparing the simulation results, and selecting the optimal network topology. It is also possible to quantify the influence weights of different factors on the topology selection.

[0025] In the above solution, by systematically analyzing the influence of model information and deployment information on the network topology selection in the model training scenario, the training efficiency is optimized. In the first part of the experiment, by fixing the model information except for the hidden layer dimension, the context length of the input tensor, and the number of model layers, these variables are changed one by one to observe their influence on the training topology selection. In the second part of the experiment, the deployment information except for the batch size (Batch / Micro-batch (micro-batch)) and the distributed strategy (such as the parallelism size of TP (tensor parallelism) / PP (pipeline parallelism)) is fixed, and these variables are changed one by one, and their influence on the training topology selection is also observed.

[0026] Furthermore, experiments are carried out on three different models, GPT2, Llama2, and MOE, to ensure the wide applicability of the results. Different models may have different computing requirements and memory occupations, so their training topology selections may be different. Use TopoOpt for coarse-grained simulation to provide a preliminary reference for the training time under different topology structures. Use ffsim-opera (traffic simulation tool) to perform fine-grained simulation based on TCP to obtain more accurate training time data, so as to more accurately determine the optimal training topology.

[0027] Furthermore, in both parts of the experiment, care was taken to avoid cross - effects between experimental variables. For example, in the first part of the experiment, the distributed training strategy was kept unchanged to focus on the impact of model information on topology selection; in the second part of the experiment, the model parameters were kept unchanged to focus on the impact of deployment information on topology selection.

[0028] The following details the content of this application with specific embodiments.

[0029] The basic content of this application is divided into two major parts: one is to determine the impact of changes in model information on network topology selection in the model training scenario. The other is to determine the impact of changes in deployment information on network topology selection in the model training scenario. The specific description is as follows: In the experiment of the first part, the control variable method was used to conduct experimental tests based on TopoOpt. For model information: hidden layer dimension, context length of the input tensor, and number of model layers, one variable was changed while the others remained unchanged under each of the three models, GPT2, Llama2, and MOE, and at the same time, the influence of the experimental variables in the second part was avoided. The distributed training strategy was also kept the same. The training time of different topological structures simulated under the coarse - grained TopoOpt was used as a reference for the training time of the current topology, and the fine - grained training time was simulated based on TCP in ffsim - opera to determine a more accurate training time. By comparing the experimental data, the optimal training topology of the model under a certain change in model information and the degree of its influence on topology selection were determined.

[0030] As an optional solution, in the extraction of the structural parameters of the target large - model, the priority order is: Number of model layers > Hidden layer dimension > Context length; When the number of model layers exceeds the threshold, Fattree - like topologies are preferentially selected; In the determination of the distributed training strategy and batch size, the priority order is: Degree of parallelism > Batch size; When the degree of parallelism exceeds the threshold, TopoOpt - like topologies are preferentially selected.

[0031] In the experiments of the second part, the control variable method is adopted to conduct experimental tests based on TopoOpt. Similarly, under the three models of GPT2, Llama2, and MOE respectively, the deployment information: batch size Batch / Mirco-batch and distributed strategies such as TP / PP parallelism are controlled variables one by one. One of them is changed while the other variables remain unchanged. At the same time, the influence of the experimental variables in the first part is avoided, and the model parameters are also kept the same. The training time of different topological structure simulations under the coarse-grained TopoOpt is used as a reference for the training time of the current topology. The fine-grained training time simulation is carried out based on TCP in ffsim-opera to determine a more accurate training time. By comparing the experimental data, the optimal training topology of the model under the change of a certain deployment information and the degree of its influence on topology selection are determined.

[0032] In the above solution, a solution for determining the optimal network topology selection in the training scenario based on model information for large models, not limited to simple deep learning models, is proposed. At the same time, a solution for determining the optimal topology selection in the large model training scenario based on model deployment information is proposed. At the same time, when both model information and deployment information change, whether different factors have the same or opposite effects on the optimal topology selection effect of model training is proposed.

[0033] Furthermore, there is currently no method for network topology selection in the large model training scenario. This application extends TopoOpt, adds Embedding (word embedding) operators and MOE model-related operators in large models, and expands the dimensions of the operators according to the dimensions of the input data, with a wider applicability. The factors affecting network topology selection in the large model training scenario are determined, considering both the model's own structural information such as model parameters and distributed training information such as distributed strategies and batch size. Both internal and external factors are considered comprehensively, which is very persuasive. Under the fine-grained TCP Packet-level simulation, the simulation results of large model training are closer to the published training data, increasing the persuasiveness and credibility of the experimental conclusions.

[0034] This application gives specific steps, such as Figure 2 shown, including the following steps: The first step is to determine the target large model for experimental testing, and select three models: GPT2, Llama2, and MOE. The reason is that GPT2 and Llama belong to the Decoder type with the same underlying structure and are dense network structures, while MOE belongs to a sparse network structure. Comparing different dimensions can make the results more universal. The experiment is based on the TopoOpt project, and its search object needs to be extended to large models to achieve the expansion of Embedding operators and data processing dimensions, and write the implementations of GPT2, Llama2, and MOE models.

[0035] In the second step, according to the target large model type, such as GPT2 or Llama2 or MOE, measure the forward and backward propagation time of its operators. According to possible parallel strategies, such as DP (Data Parallelism), TP (Tensor Parallelism, etc.), measure its propagation time under this parallel strategy as a candidate for the optimal parallel strategy of the operator. After measuring all operators, generate the measure.json file of the results.

[0036] In the third step, after the Simulator class sets the topology type of the cluster, calculate the operator computation time from the measure.json file imported for the model of this class. Randomly select a parallel scheme for each operator and calculate according to the data flow direction. Finally, obtain the current simulated training time. Define the training time at this time as the reference time for the optimal training time of the current topology. At the same time, generate a Taskgraph file in flatbuffer format.

[0037] In the fourth step, select control variables, change at least one of the model information or deployment information, and keep other variables unchanged, then repeat the third step.

[0038] Among them, through experiments, the result directory when testing model factors can be obtained. The result directory is the file directory where the test results are located, showing the variables changed in the experiment and other parameter information.

[0039] In the fifth step, use the Taskgraph files in flatbuffer format generated in the third and fourth steps as the input of the ffsim-Opera simulator. The Taskgraph generated by the topology type selected in the third step should correspond to the current topology selection type, and conduct fine-grained TCP Packet-level training time simulation.

[0040] In the sixth step, perform the second to fifth experimental steps for each large model and organize the data.

[0041] Through the above steps, The following conclusions can be drawn for this application : 1. For the large model training scenario, both model information (hidden layer dimension, context length of the input tensor, number of model layers) and deployment information (batch size Batch / Mirco-batch and distributed strategies such as TP / PP parallelism size) will affect the selection of the model array center network topology. 2. For the model information factor, the increase or decrease in the number of model layers has a greater impact on the selection of the model (logical array center) network topology than the increase or decrease in the hidden layer dimension and the context of the input tensor. When the number of model layers is set to a relatively large value, the simulation running time is shorter under the Fattree class topology structure, and there is a tendency to select the Fattree class data center network topology. Some experimental results are as Figure 3 shown.

[0042] 3. For the deployment information factor, the increase or decrease in the parallelism of the distributed strategy has a greater impact on the selection of the network topology of the model (logical array center) than the increase or decrease in the batch size (Batch / Mirco-batch). When the parallelism is set to a relatively large value, the simulation running time is shorter under the TopoOpt class of topology structures, and there is a tendency to select the TopoOpt class of data center network topologies.

[0043] Therefore, the main purpose of this application is to find the optimal network topology according to the model information or (and) deployment information training scenarios of the current type of large model, and it can also be applied to the following scenarios: 1) When the model structures are similar and the model parameter sizes are close, it can guide the selection of the underlying network topology during training; 2) It can also guide the network topology selection according to similar distributed strategies; 3) Or select the optimal distributed training strategy according to the underlying network topology structure.

[0044] Furthermore, the feasibility of the method of this application is verified through experiments. This application is implemented in CUDA and Cpp languages and tested through Shell scripts. CUDA is used to measure the forward and backward propagation times of the model operators. The Simulator class is designed to construct the Taskgraph (task flow chart) of the operators based on the definition of the model structure. Data dependencies are added through the addition of dependencies to the Taskgraph. For the Allreduce operation, a new data path is added by constructing the routing of the source node and the target node, thereby adding new communication-related tasks, and the training time is simulated at a coarse-grained level. The design of the CLOS class of Data Center Network (DCN) structures includes, but is not limited to, AbstractSwitch (ideal switch topology), Fattree (fat tree topology), Sip-ML (DNN training cluster with Tbps-level bandwidth per GPU, allocate d wavelengths (each wavelength has a bandwidth of B) to the GPUs in each SiP-ML topology cluster, and find the topology with a reconfiguration delay of 25 μs according to the SiP-Ring algorithm), TopoOpt (multi-way topology based on the Ring structure), Expander (Expander topology where each Server, i.e., the device providing network services, has d NIC cards and the network bandwidth is B), OCS-Reconfig (Optical Cicruit Switch optical line switch topology) structures, etc., to achieve TCP Packet-level simulation. On this basis, more topology types can be extended to make the experimental conclusions have a wider topology universality.

[0045] Such asFigure 4 As shown in the figure, the third objective of the present application is to provide a network topology optimization selection system for large model training scenarios, including: A model information analysis module 100, which is used to extract the structural parameters of the target large model, including the hidden layer dimension, the input tensor context length, and the number of model layers, quantify the impact of the structural parameters on the network communication requirements, and analyze to obtain model information; A deployment information analysis module 200, which is used to determine the distributed training strategy and batch size, analyze the requirements for network bandwidth and latency, and obtain deployment information; A coarse / fine-grained simulation module 300, which is used to perform coarse-grained simulation on the target large model to generate reference training time values under different network topologies; perform fine-grained simulation based on the reference training time values to verify the topology performance and obtain simulation results; An optimal network topology selection module 400, which is used to compare the simulation results according to the dynamic changes of the model information and / or deployment information, select the optimal network topology, and quantify the influence weights of different factors on the topology selection.

[0046] The network topology optimization selection system for large model training scenarios of the present application is based on the above-mentioned network topology optimization selection method for large model training scenarios.

[0047] As Figure 5 shown, the fourth objective of the embodiment of the present application is to provide an electronic device, including a memory 701, a processor 702, and a computer program stored in the memory 701 and executable on the processor. When the processor executes the computer program, the above-mentioned network topology optimization selection method for large model training scenarios is implemented. It also includes a communication interface 703 and a bus 704.

[0048] The fifth objective of the embodiment of the present application is to provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned network topology optimization selection method for large model training scenarios is implemented.

[0049] The sixth objective of the embodiment of the present application is to provide a computer program product. The computer program product includes computer instructions, and the computer instructions instruct the computer to execute the above-mentioned network topology optimization selection method for large model training scenarios.

[0050] These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements in the process Figure 1 one process or multiple processes and / or boxes Figure 1The functions specified in one or more boxes.

[0051] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in Figure 1 one process or more processes and / or boxes Figure 1 the functions specified in one box or more boxes.

[0052] This application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, readable storage media, optical storage, etc.) that contain computer-usable program code.

[0053] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or box in the flowchart and / or block diagram, as well as the combination of flows and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in Figure 1 one process or more processes and / or boxes Figure 1 the functions specified in one box or more boxes.

[0054] Obviously, the described embodiments are only partial embodiments of this application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts should belong to the scope protected by this application.

[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and not to limit them. Although this application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of this application or make equivalent replacements, and any modification or equivalent replacement that does not deviate from the spirit and scope of this application should be covered within the scope of the claims of this application.

Claims

1. A network topology optimization selection method for large model training scenarios, characterized in that: include: Extract the structural parameters of the target large model, including the hidden layer dimension, input tensor context length, and number of model layers, and quantify the impact of the structural parameters on network communication requirements to obtain model information through analysis; Determine the distributed training strategy and batch size, analyze the requirements for network bandwidth and latency, and obtain deployment information; Perform coarse-grained simulation on the target large model to generate training time reference values ​​under different network topologies; Perform fine-grained simulation based on training time reference values ​​to verify topology performance and obtain simulation results; According to the dynamic changes of model information and / or deployment information, the simulation results are compared and the optimal network topology is selected.

2. According to a network topology optimization selection method for large model training scenarios according to claim 1, it is characterized in that: The target large model includes a dense model and a sparse model: Dense models include large models based on the Transformer architecture; Sparse models include mixture of experts large models; The extraction of the structural parameters of the target large model is achieved by extending the network topology optimization tool to support the Embedding operator and the MOE model, and extracting the structural parameters.

3. According to a network topology optimization selection method for large model training scenarios according to claim 1, it is characterized in that: The coarse-grained simulation includes: building a task graph of the model operator by simulating the data flow of distributed training; dynamically generating communication tasks according to the parallel strategy and adding routing dependencies of the Allreduce operation; The fine-grained simulation includes: simulating the impact of data packet transmission delay, packet loss rate and congestion control on training time based on the TCP protocol.

4. According to a network topology optimization selection method for large model training scenarios according to claim 1, it is characterized in that: In the extraction of the target large model structure parameters, the priority order is: Model number of layers > hidden layer dimension > context length; When the number of model layers exceeds the threshold, Fattree topology is preferred; In determining the distributed training strategy and batch size, the priority order is: Parallelism size > batch size; When the degree of parallelism exceeds the threshold, the TopoOpt topology is preferred.

5. According to a network topology optimization selection method for large model training scenarios according to claim 1, it is characterized in that: Also includes: When model information and / or deployment information change simultaneously, the combined influence coefficient of the two is calculated by the influence weight, and the topology selection strategy is dynamically adjusted; Output the optimal network topology configuration solution for the current training scenario.

6. According to a network topology optimization selection method for large model training scenarios according to claim 1, it is characterized in that: Perform coarse-grained simulation on the target large model to generate training time reference values ​​under different network topologies; Fine-grained simulation is performed based on the training time reference value to verify the topology performance and obtain simulation results, including: According to the target large model type, forward and backward propagation time of the target large model operator is measured, the propagation time is measured according to possible parallel strategies, the propagation time is used as the optimal parallel strategy candidate of the operator, and a first result file is generated after measuring all operators; After setting the topology type of the cluster, by importing the operator calculation time of the first result file, randomly selecting a parallel scheme for each operator, calculating according to the data flow direction, obtaining the current simulation training time, taking the training time as the optimal training time reference time for the current topology, and generating the second result file at the same time; The first result file and the second result file are used as inputs of the simulator. The second result file generated by the selected topology type corresponds to the topology selection type. A fine-grained level training time simulation is performed to verify the topology performance and obtain simulation results.

7. A network topology optimization selection system for large model training scenarios, characterized in that: include: Model information analysis module, used to extract the structural parameters of the target large model, including hidden layer dimensions, input tensor context length, number of model layers, and quantify the impact of structural parameters on network communication requirements, and analyze and obtain model information; Deployment information analysis module, used to determine the distributed training strategy and batch size, analyze the requirements for network bandwidth and latency, and obtain deployment information; The coarse / fine-grained simulation module is used to perform coarse-grained simulation on the target large model and generate training time reference values ​​under different network topologies; Perform fine-grained simulation based on training time reference values ​​to verify topology performance and obtain simulation results; The optimal network topology selection module is used to compare simulation results and select the optimal network topology according to the dynamic changes of model information and / or deployment information.

8. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the network topology optimization selection method for a large model training scenario according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the network topology optimization selection method for a large model training scenario according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising computer instructions, characterized in that: The computer instructions instruct the computer to execute the network topology optimization selection method for large model training scenarios as described in any one of claims 1-6.