Parallel training simulation method and device for large language model

CN122819306APending Publication Date: 2026-09-25NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511350500.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]本发明主要解决如何对大语言模型的训练过程的计算消耗和通信消耗进行准确模拟的问题,本发明公开了一种大语言模型的并行训练模拟方法和装置

Benefits of technology

[0042]本发明设计了首个专为ZeRO策略下参数分片大语言模型训练设计的执行驱动仿真器。与先前基于合成数据或仅重放通信轨迹的工具不同,ZeRO-Emu将真实GPU内核性能分析与CPU集群上的实际集合通信重放相结合,并通过通信锚点机制显式建模计算-通信重叠。这使其能高精度仿真ZeRO-1/2/3配置,在无需GPU硬件的情况下实现迭代级精度和真实延迟捕捉。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819306A_ABST
    Figure CN122819306A_ABST
Patent Text Reader

Abstract

The application discloses a parallel training simulation method and device for a large language model, and the method comprises the following steps: acquiring a large language model information set; performing computation graph generation processing on the large language model information set to obtain computation graph information; and performing model training time simulation simulation based on the computation graph information to obtain the total training time of the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence processors and large language model training simulation, specifically to a parallel training simulation method and apparatus for large language models. Background Technology

[0002] Training large language models (LLMs) is increasingly constrained by GPU memory, prompting the adoption of sharding strategies such as ZeRO and FSDP. These strategies reduce memory footprint by partitioning the model state and replacing AllReduce with finer-grained operations such as ReduceScatter and AllGather. However, validating these techniques on real GPU clusters is costly, and existing simulators, which rely on analytical models, cannot capture the runtime perturbations and computation-communication overlap changes introduced by sharding, hindering rapid iteration of algorithm-system co-design. Summary of the Invention

[0003] This invention primarily addresses the problem of accurately simulating the computational and communication overhead of training large language models. It discloses a parallel training simulation method and apparatus for large language models.

[0004] In a first aspect, this invention discloses a parallel training simulation method for a large language model, comprising:

[0005] S1, obtain the large language model information set;

[0006] S2, Perform computation graph generation processing on the large language model information set to obtain computation graph information;

[0007] S3. Based on the computation graph information, perform a simulation of the model training time to obtain the total training time of the large language model.

[0008] The large language model information set includes large language model structure information, parallel training strategy information, and training device information;

[0009] The large language model structure information includes the functional modules contained in the large language model and the connection relationships between the functional modules;

[0010] The parallel training strategy information includes the number of training iterations of the large language model, the communication node sequence, and the computing node sequence; the communication node sequence is a sequence of communication node information used in chronological order during a training process; the computing node sequence is a sequence of computing node information used in chronological order during a training process.

[0011] The communication node information includes the communication node serial number, communication task information, and information interaction relationship with other nodes;

[0012] The computing node information includes the computing node number, computing task information, and information interaction relationships with other nodes; the computing task information refers to the model layer operator operations performed on the computing node, including convolution, pooling, sampling, and loss function calculation.

[0013] The training device information includes the hardware composition information of the training device; the hardware composition information includes the number of CUDA cores contained in the GPU of the training device and the CUDA core connection relationship.

[0014] The computation graph information includes nodes and edges; the nodes are used to represent model layer operator operations or communication operations in a training process of a large language model, and the edges are used to represent the connection relationship between nodes and the direction of data flow between nodes.

[0015] The simulation of model training time based on the computation graph information yields the total training time of the large language model, including:

[0016] S31, Based on the computation graph information, perform computation performance statistics to obtain the total computation time and computation communication graph;

[0017] S32, based on the calculated communication graph, perform communication time capture to obtain the total communication time;

[0018] S33, add the total computation time and the total communication time to obtain the training time;

[0019] S34 multiplies the training time of one session by the number of training sessions of the large language model to obtain the total training time of the large language model, thus completing the parallel training simulation of the large language model.

[0020] The step of performing computational performance statistics based on the computation graph information to obtain the total computation time and computational communication graph includes:

[0021] S311, Deploy the large language model structure information on the training device;

[0022] S312, Based on computation graph information, train the large language model once on a single GPU of the training device;

[0023] S313, using the CUPTI tool, extract the set of training information for a single session from all CUDA cores of the single GPU;

[0024] The single training information set includes a CUDA operator information set, a CUDA operator time set, and operator layer mapping information; the operator layer mapping information represents the correspondence between CUDA operators and model layer operators; the CUDA operator information set includes several CUDA operator information sets; the CUDA operator information sets are the names of the CUDA operators performed on each CUDA core; the CUDA operator time set includes CUDA operator times; the CUDA operator times are the time required to run CUDA operators on each CUDA core.

[0025] S314, Based on the single training information set, update the computation graph information to obtain the computation communication graph;

[0026] S315, sum up the times of all CUDA operators in the single training information set to obtain the total computation time.

[0027] The step of updating the computation graph information based on the single training information set to obtain a computation communication graph includes:

[0028] S3141, Obtain operator layer mapping information from the single training information set;

[0029] S3142, Based on the operator layer mapping information, find the CUDA operator corresponding to the model layer operator for the edge in the computation graph information;

[0030] S3143, using the CUDA operator corresponding to the model layer operator of each edge of the computation graph information, the model layer operator of the edge is replaced to obtain the computation communication graph.

[0031] In a second aspect, the present invention discloses a parallel training simulation device for a large language model, used to implement the parallel training simulation method for the large language model, comprising: a graph generator module, a performance model module, and a communication capture module;

[0032] The graph generator module is used to acquire a large language model information set, perform computation graph generation processing on the large language model information set, and obtain computation graph information.

[0033] The performance model module is used to perform computational performance statistics based on the computation graph information to obtain the total computation time and computational communication graph; send the computational communication graph to the communication capture module; receive the total communication time sent by the communication capture module; add the total computation time and the total communication time to obtain the training time; and multiply the training time by the number of training iterations of the large language model to obtain the total training time of the large language model.

[0034] The communication capture module is used to capture communication time based on the computational communication graph, obtain the total communication time, and send the total communication time to the performance model module.

[0035] A third aspect of the present invention discloses a parallel training simulation device for a large language model, the device comprising:

[0036] Memory containing executable program code;

[0037] A processor coupled to the memory;

[0038] The processor calls the executable program code stored in the memory to execute the parallel training simulation method for the large language model.

[0039] In a fourth aspect of this invention, a computer-readable storage medium is disclosed, the computer-readable storage medium storing computer instructions, which, when invoked by a computer, are used to execute the parallel training simulation method for the large language model.

[0040] A fifth aspect of this invention discloses an information data processing terminal, which is used to implement the parallel training simulation method for the large language model.

[0041] The beneficial effects of this invention are as follows:

[0042] This invention designs the first execution-driven simulator specifically for training large language models with parameter fragmentation under the ZeRO strategy. Unlike previous tools based on synthetic data or simply replaying communication trajectories, ZeRO-Emu combines real GPU kernel performance analysis with actual ensemble communication replay on CPU clusters and explicitly models computation-communication overlap through a communication anchoring mechanism. This enables high-precision simulation of ZeRO-1 / 2 / 3 configurations, achieving iterative-level accuracy and realistic latency capture without requiring GPU hardware.

[0043] Through empirical evaluation on node clusters and various GPT / BERT models, ZeRO-Emu achieves an average prediction error of less than 4.51%, accurately reproducing the iterative process decomposition and communication scaling effects, even under memory-constrained ZeRO-3 configurations. Notably, our analysis reveals how ZeRO sharding exacerbates communication bottlenecks during large-scale scaling and validates the ability of our ZeRO-Emu to expose such performance trends without requiring GPU hardware. Attached Figure Description

[0044] Figure 1 This is a flowchart illustrating the implementation of the method of the present invention;

[0045] Figure 2This is a communication flowchart of the method of the present invention;

[0046] Figure 3 This is a simulation result of the method of the present invention on a single node.

[0047] Figure 4 These are experimental results of the method of this invention under various ZeRO stage configurations (ZeRO 1 / 2 / 3) and different cluster sizes (2 / 4 / 6 nodes). Detailed Implementation

[0048] To better understand the content of this invention, an embodiment is provided here.

[0049] Figure 1 This is a flowchart illustrating the implementation of the method of the present invention.

[0050] In a first aspect, this invention discloses a parallel training simulation method for a large language model, comprising:

[0051] S1, obtain the large language model information set;

[0052] S2, Perform computation graph generation processing on the large language model information set to obtain computation graph information;

[0053] S3. Based on the computation graph information, perform a simulation of the model training time to obtain the total training time of the large language model.

[0054] The large language model information set includes large language model structure information, parallel training strategy information, and training device information;

[0055] The large language model structure information includes the functional modules contained in the large language model and the connection relationships between the functional modules;

[0056] The parallel training strategy information includes the number of training iterations of the large language model, the communication node sequence, and the computing node sequence; the communication node sequence is a sequence of communication node information used in chronological order during a training process; the computing node sequence is a sequence of computing node information used in chronological order during a training process.

[0057] The communication node information includes the communication node serial number, communication task information, and information interaction relationship with other nodes;

[0058] The computing node information includes the computing node number, computing task information, and information interaction relationships with other nodes; the computing task information refers to the model layer operator operations performed on the computing node, including convolution, pooling, sampling, loss function calculation, etc.

[0059] The training device information includes the hardware composition information of the training device; the hardware composition information includes the number of CUDA cores contained in the training GPU and the CUDA core connection relationship.

[0060] The computation graph information includes nodes and edges; the nodes are used to represent model layer operator operations or communication operations in a training process of a large language model, and the edges are used to represent the connection relationship between nodes and the direction of data flow between nodes.

[0061] The simulation of model training time based on the computation graph information yields the total training time of the large language model, including:

[0062] S31, Based on the computation graph information, perform computation performance statistics to obtain the total computation time and computation communication graph;

[0063] S32, based on the calculated communication graph, perform communication time capture to obtain the total communication time;

[0064] S33, add the total computation time and the total communication time to obtain the training time;

[0065] S34 multiplies the training time of one session by the number of training sessions of the large language model to obtain the total training time of the large language model, thus completing the parallel training simulation of the large language model.

[0066] The step of performing computational performance statistics based on the computation graph information to obtain the total computation time and computational communication graph includes:

[0067] S311, Deploy the large language model structure information on the training device;

[0068] S312, Based on computation graph information, train the large language model once on a single GPU of the training device;

[0069] S313, using the CUPTI tool, extract a single training information set from all CUDA cores of the single GPU; the single training information set includes a CUDA operator information set, a CUDA operator time set, and operator layer mapping information; the operator layer mapping information represents the correspondence between CUDA operators and model layer operators; the CUDA operator information set includes CUDA operator information, which is the CUDA operator information performed on each CUDA core; the CUDA operator time set includes CUDA operator time, which is the time required to run CUDA operators on each CUDA core.

[0070] S314, Based on the single training information set, update the computation graph information to obtain the computation communication graph;

[0071] S315, sum up the times of all CUDA operators in the single training information set to obtain the total computation time.

[0072] The step of updating the computation graph information based on the single training information set to obtain a computation communication graph includes:

[0073] S3141, Obtain operator layer mapping information from the single training information set;

[0074] S3142, Based on the operator layer mapping information, find the CUDA operator corresponding to the model layer operator for the edge in the computation graph information.

[0075] S3143, using the CUDA operator corresponding to the model layer operator of each edge of the computation graph information, the model layer operator of the edge is replaced to obtain the computation communication graph.

[0076] The extraction of the operator layer mapping information can be performed using the Daydream algorithm proposed by Hongyu Zhu et al.

[0077] The step of capturing communication time based on the calculated communication graph to obtain the total communication time includes:

[0078] Based on the computation-communication graph, communication operations from the parallel training strategy information are performed on the CPU cluster of the training device, and the total communication time is statistically obtained. Specifically, based on the parallel training strategy information, the Gloo library is used directly to execute real communication primitives on the CPU cluster, which enables accurate simulation of inter-node communication behavior without relying on synthetic latency formulas. To ensure accurate measurement of ensemble operations (such as AllReduce, AllGather, and ReduceScatter), careful synchronization between ranks is required to avoid deviations in start time. To address this issue, this invention implements a two-stage synchronization mechanism. First, all ranks are paused, awaiting a scheduling signal from the rank 0 node (CPU), which traverses the computation-communication graph (CCG) to determine when each ensemble operation should be triggered. Once synchronization is established, the ensemble operation is executed concurrently on all participating ranks. Upon completion, the rank 0 node (CPU) issues a second synchronization barrier, stopping the communication process of all communication nodes and ensuring consistency in timestamp collection. This protocol is crucial for accurately capturing the runtime cost of ensemble operations and avoiding underestimation due to asynchronous start time drift. During the communication simulation, each communication bucket is assigned a globally unique identifier. By explicitly coordinating collection calls and capturing runtime behavior, this invention provides a realistic analysis of communication overhead, even in large-scale, memory-constrained ZeRO configurations. This execution-based approach enables the invention to simulate real bandwidth contention, message granularity effects, and phase-specific communication intensity—all crucial for modeling ZeRO-style parallelism under memory constraints.

[0079] Another possible implementation of S315 includes:

[0080] S3151, obtain all CUDA operator times and CUDA computational task parameter information sets in a single training information set at several historical moments; the computational task parameter information set includes computational data volume, CUDA storage space value, and CUDA computational speed value.

[0081] S3152, for each CUDA, the CUDA operator time is the dependent variable, and the set of CUDA computation task parameter information is the multivariate independent variable;

[0082] S3153, perform fusion fitting on the multivariate independent and dependent variables to obtain the time prediction model for each CUDA;

[0083] S3154: Collect the set of CUDA computation task parameter information at the current moment, and use the time prediction model of each CUDA to calculate the corresponding set of computation task parameter information to obtain the computation time of each CUDA.

[0084] S3155 sums up the computation time of all CUDA operations to obtain the total computation time.

[0085] The fusion fitting process includes:

[0086] The first regression model is obtained by performing multiple linear regression on the multivariate independent and dependent variables.

[0087] The multivariate independent and dependent variables are subjected to nonlinear regression to obtain the second regression model.

[0088] Fuzzy regression was performed on the multivariate independent and dependent variables to obtain the third regression model;

[0089] Using each regression model, the computational task parameter information set of CUDA at the aforementioned historical moments is processed to obtain the predicted value;

[0090] The set of residual values ​​for each regression model is calculated; the residual values ​​are obtained by subtracting the predicted values ​​from the CUDA operator time.

[0091] All regression models are fused and calculated to obtain the time prediction model rf(x) for each CUDA model;

[0092] The expression for the fusion calculation is:

[0093]

[0094] Where Ei is the exponential integral, x is the input multivariate independent variable, and f i (x) represents the i-th regression model, μ i and δ i denoted as the mean and variance of the set of residual values ​​for the i-th regression model, respectively.

[0095] In Large Language Model (LLM) training optimization scenarios, the core objective of S2's "computation graph generation and processing" is to transform the unstructured set of model information (structure, parallel strategies, devices) into a structured, executable computation graph representation, providing a visual and computable "blueprint" for subsequent training task scheduling, resource allocation, and performance optimization. This process revolves around four core stages: "information parsing - logical mapping - constraint fusion - optimization verification," with specific steps and details as follows:

[0096] 1. Prerequisites for computation graph generation: Information standardization and analysis

[0097] The prerequisite for generating the computation graph is to transform the "large language model information set" (structure, parallel strategy, device) acquired by S1 into a standardized data format that can be recognized by machines (such as JSON, Protocol Buffers) to avoid mapping deviations caused by chaotic information formats. This step requires parsing and structuring three types of information: LLM training usually relies on distributed parallelism (such as data parallelism, tensor parallelism, pipelined parallelism), and this step needs to transform the "number of training sessions, communication node sequence, and computation node sequence" into a time-node association table to clarify "when to compute, when to communicate, and which node to use".

[0098] 2. Hierarchical Construction of the Computation Graph

[0099] Based on the standardized parsed information, the computation graph generation adopts a "layered construction" strategy—gradually refining from the "logic layer computation graph" to the "physical layer computation graph" to ensure "model logic correctness" and "hardware resource adaptability".

[0100] The large language model information set is stored as a JSON file.

[0101] The ZeRO-Emu disclosed in this invention is a lightweight simulator for large-scale model training. It uses a single GPU to extract the computation graph and replays the corresponding communications on a CPU cluster, thereby achieving high-fidelity performance prediction without accessing the GPU cluster. Figure 1 As shown, ZeRO-Emu takes user profiles as input and automatically constructs a computation graph for distributed execution based on the specified parallelization strategy, while accurately capturing the actual execution latency of each communication primitive.

[0102] The simulator consists of three main components: a graph generator, a performance model, and a communication capture module.

[0103] Based on the computation graph, ZeRO-Emu performs a single-GPU forward and backward training process (i.e., the entire training process) on rank 0 (the only node equipped with a physical GPU). This execution is implemented using the PyTorch framework and is only used to extract the structure and runtime metadata required for the simulation.

[0104] From this real-world run, ZeRO-Emu uses CUPTI to capture operator sequences, tensor shapes, and inter-layer dependency information from a single GPU's CUDA core, in order to construct a unified representation of model execution.

[0105] Based on this analysis process, a task-level computation-communication graph is constructed. A task-level computation-communication graph (CCG) encodes the computational steps and communication interactions required for distributed execution.

[0106] Each model layer is decomposed into fine-grained task nodes, including forward, backward, and parameter update operations. These nodes are further annotated with parallel-aware metadata reflecting tensor partitioning, gradient slicing, and optimizer state distribution, depending on the selected ZeRO stage.

[0107] To realistically simulate actual synchronization behavior, communication nodes (such as AllReduce_bucket, Reduce_Scatter_bucket, and All_Gather_bucket) are inserted into the graph. These nodes are scheduled according to the semantics of each ZeRO stage. All communication dependencies are explicitly encoded to preserve the causal relationships between computation and communication, enabling the simulator to accurately simulate overlapping patterns, bucket scheduling, and inter-node traffic.

[0108] In the performance model, in order to simulate GPU workloads on a CPU cluster, ZeRO-Emu performs an analysis step to build a performance mapping table.

[0109] This table, based on runtime traces collected from actual GPU training, links each high-level model operator (e.g., feedforward network, multi-head self-attention) to its corresponding execution time. We use a CUPTI-based tool to break down each operator into its underlying CUDA kernels and measure their execution time in a fine-grained manner.

[0110] To capture operator-level variations under different workloads, analytics tracks are collected in various configurations, including batch size, sequence length, and model depth.

[0111] ZeRO-Emu integrates awareness of the memory partitioning behavior introduced by the ZeRO strategy. The simulator considers these partitioning strategies by adjusting tensor shapes and communication boundaries accordingly, ensuring that timing and memory models are consistent with the actual data layout on each device. Analysis results are aggregated into a unified function-delay map. During simulation, each compute node in the task graph is assigned a delay value based on its operator type and context.

[0112] This design allows ZeRO-Emu to simulate real computation timings without repeatedly executing kernels, thus efficiently approximating GPU workloads on the CPU. The performance mapping module is modular and scalable, supporting future integration with hardware-specific latency models or compiler-level performance predictors.

[0113] After constructing a complete CCG at rank 0 and labeling the latency values ​​obtained from the analysis for each node, ZeRO-Emu initiates a time-driven simulation via the predict module to simulate a complete training iteration. Within each GPU rank, operations are sequentially mapped to CUDA streams, strictly adhering to inter-layer and intra-layer dependency constraints.

[0114] like Figure 2 As shown, the black arrows indicate the execution order of GPU-side kernels, which are organized into logical flows for computation and communication. This separation allows ZeRO-Emu to model overlapping behaviors where possible, such as concurrently issuing ReduceScatter or AllGather operations with computation kernels on separate flows. The blue arrows represent the mapping between CPU-side operator definitions and their corresponding GPU kernel launches, highlighting the boundary between graph construction and runtime scheduling. These links are crucial for aligning CPU-issued operations with their underlying GPU counterparts. Furthermore, the green arrows represent the CPU-level control flow across model layers, driving the order of operator launches and communication-triggered scheduling. Communication events are injected according to a specified parallel strategy (e.g., ZeRO-1 / 2 / 3) and the system's gradient bucketing strategy. These events are placed at precisely computed points on the timeline to reflect real-world synchronization behavior and communication latency.

[0115] For each rank, the duration of each iteration is measured from the timestamp of the first operator's start to the time when the last task on all streams completes. Global synchronization is modeled by ordering rank execution using a FIFO (First-In-First-Out) strategy, respecting any cross-rank dependencies caused by optimizer state partitioning or parameter swapping. The critical path (defined as the longest cumulative execution path across all ranks) determines the iteration time of the simulation.

[0116] This fine-grained scheduling process, based on operator latency and real communication behavior, enables ZeRO-Emu to reconstruct the true execution timeline while accurately capturing the effects of parallel configuration, communication overlap, and system heterogeneity.

[0117] We implemented ZeRO-Emu using Python 3.10.5 and torch 2.6.1 (with CUDA 12.1 support). To obtain accurate computational analysis data, we integrated CUPTI (CUDA Profiling Tools Interface) to collect fine-grained traces of CUDA API calls and GPU kernel execution. To model communication behavior, our simulator supports NCCL (v2.17.1) and Gloo as backends; all CPU cluster-based simulations used Gloo. All experiments were performed on a dedicated GPU cluster, and the configuration is summarized in Table 1. Iteration time was measured after a warm-up phase to eliminate initialization artifacts, and results were averaged over 10 consecutive training steps to reduce noise. The evaluated models included encoder-only and decoder-only Transformer architectures (GPT3 and BERT), covering different ZeRO parallel stages and cluster sizes.

[0118] To systematically evaluate the simulation accuracy of ZeRO-Emu, we used DeepSpeed ​​to conduct experiments under various models and scales (BERT 0.336B / 0.11B, GPT-3 0.76B / 1.3B / 2.7B / 6.7B), various ZeRO stage configurations (ZeRO 1 / 2 / 3), and different cluster sizes (2 / 4 / 6 nodes). For each setting, we collected the iteration times of real training and compared them with the prediction results of ZeRO-Emu. The results are as follows: Figure 3 and Figure 4 As shown. Figure 3 These are the experimental results from a single node; Figure 4 These are experimental results conducted under various ZeRO phase configurations (ZeRO 1 / 2 / 3) and different cluster sizes (2 / 4 / 6 nodes).

[0119] We first evaluate ZeRO-Emu in the simplest but most basic scenario: single-GPU training with a batch size of 128. Figure 3 The measured and predicted iteration times of BERT and GPT-3 variants were compared on V100 and A100 GPUs. The mean absolute error was 3.96% on V100 and 3.64% on A100 across all models; none exceeded 5% on any single run. This high fidelity stems from the absence of inter-node communication: the only remaining uncertainty is GPU-level microarchitectural noise (e.g., dynamic frequency scaling, SM occupancy jitter), which is captured by our CUPTI-based tracking. After warm-up and a 10-step averaging, these fluctuations were reduced to within the reported statistical interval, confirming the accuracy of the ZeRO-Emu computational latency model when communication variables are eliminated.

[0120] To evaluate the accuracy of ZeRO-Emu in distributed training scenarios, we conducted experiments using GPT-3 and BERT model variants, covering ZeRO-1, ZeRO-2, and ZeRO-3 strategies, under 2×4 2×4, 4×4 4×4, and 6×4 6×4 GPU cluster configurations. Figure 4 As shown, the simulator achieved an average prediction error of 4.51% per iteration in all experiments, with a maximum error of less than 8%, demonstrating its robustness under different cluster sizes and topologies.

[0121] In the calculation expressions of this invention, the variables involved have all been dimensionless before calculation.

[0122] In a second aspect, the present invention discloses a parallel training simulation device for a large language model, used to implement the parallel training simulation method for the large language model, comprising: a graph generator module, a performance model module, and a communication capture module;

[0123] The graph generator module is used to acquire a large language model information set, perform computation graph generation processing on the large language model information set, and obtain computation graph information.

[0124] The performance model module is used to perform computational performance statistics based on the computation graph information to obtain the total computation time and computational communication graph; send the computational communication graph to the communication capture module and receive the total communication time sent by the communication capture module; add the total computation time and the total communication time to obtain the training time; and multiply the training time by the number of training iterations of the large language model to obtain the total training time of the large language model.

[0125] The communication capture module is used to capture communication time based on the calculated communication graph to obtain the total communication time.

[0126] A third aspect of the present invention discloses a parallel training simulation device for a large language model, the device comprising:

[0127] Memory containing executable program code;

[0128] A processor coupled to the memory;

[0129] The processor calls the executable program code stored in the memory to execute the parallel training simulation method for the large language model.

[0130] In a fourth aspect of this invention, a computer-readable storage medium is disclosed, the computer-readable storage medium storing computer instructions, which, when invoked by a computer, are used to execute the parallel training simulation method for the large language model.

[0131] A fifth aspect of this invention discloses an information data processing terminal, which is used to implement the parallel training simulation method for the large language model.

[0132] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A parallel training simulation method for a large language model, characterized in that, include: S1, obtain the large language model information set; S2, Perform computation graph generation processing on the large language model information set to obtain computation graph information; S3. Based on the computation graph information, perform a simulation of the model training time to obtain the total training time of the large language model.

2. The parallel training simulation method for large language models as described in claim 1, characterized in that, The large language model information set includes large language model structure information, parallel training strategy information, and training device information; The large language model structure information includes the functional modules contained in the large language model and the connection relationships between the functional modules; The parallel training strategy information includes the number of training iterations of the large language model, the communication node sequence, and the computing node sequence; the communication node sequence is a sequence of communication node information used in chronological order during a training process; the computing node sequence is a sequence of computing node information used in chronological order during a training process. The communication node information includes the communication node serial number, communication task information, and information interaction relationship with other nodes; The computing node information includes the computing node number, computing task information, and information interaction relationships with other nodes; the computing task information refers to the model layer operator operations performed on the computing node, including convolution, pooling, sampling, and loss function calculation. The training device information includes the hardware composition information of the training device; the hardware composition information includes the number of CUDA cores contained in the GPU of the training device and the CUDA core connection relationship.

3. The parallel training simulation method for large language models as described in claim 1, characterized in that, The computation graph information includes nodes and edges; the nodes are used to represent model layer operator operations or communication operations in a training process of a large language model, and the edges are used to represent the connection relationship between nodes and the direction of data flow between nodes.

4. The parallel training simulation method for large language models as described in claim 2, characterized in that, The simulation of model training time based on the computation graph information yields the total training time of the large language model, including: S31, Based on the computation graph information, perform computation performance statistics to obtain the total computation time and computation communication graph; S32, Based on the calculated communication graph, perform communication time capture to obtain the total communication time; S33, add the total computation time and the total communication time to obtain the training time; S34 multiplies the training time of one session by the number of training sessions of the large language model to obtain the total training time of the large language model, thus completing the parallel training simulation of the large language model.

5. The parallel training simulation method for a large language model as described in claim 4, characterized in that, The step of performing computational performance statistics based on the computation graph information to obtain the total computation time and computational communication graph includes: S311, Deploy the large language model structure information on the training device; S312, Based on computation graph information, train the large language model once on a single GPU of the training device; S313, using the CUPTI tool, extract the set of training information for a single session from all CUDA cores of the single GPU; The single training information set includes a CUDA operator information set, a CUDA operator time set, and operator layer mapping information; the operator layer mapping information represents the correspondence between CUDA operators and model layer operators; the CUDA operator information set includes several CUDA operator information sets; the CUDA operator information sets are the names of the CUDA operators performed on each CUDA core; the CUDA operator time set includes CUDA operator times; the CUDA operator times are the time required to run CUDA operators on each CUDA core. S314, Based on the single training information set, update the computation graph information to obtain a computation communication graph; S315, sum up the times of all CUDA operators in the single training information set to obtain the total computation time.

6. The parallel training simulation method for a large language model as described in claim 5, characterized in that, The step of updating the computation graph information based on the single training information set to obtain a computation communication graph includes: S3141, Obtain operator layer mapping information from the single training information set; S3142, Based on the operator layer mapping information, find the CUDA operator corresponding to the model layer operator for the edge in the computation graph information; S3143, using the CUDA operator corresponding to the model layer operator of each edge of the computation graph information, the model layer operator of the edge is replaced to obtain the computation communication graph.

7. A parallel training simulation device for a large language model, characterized in that, A parallel training simulation method for implementing a large language model as described in any one of claims 1 to 6 includes: a graph generator module, a performance model module, and a communication capture module; The graph generator module is used to acquire a large language model information set, perform computation graph generation processing on the large language model information set, and obtain computation graph information. The performance model module is used to perform computational performance statistics based on the computation graph information to obtain the total computation time and computational communication graph; send the computational communication graph to the communication capture module; receive the total communication time sent by the communication capture module; add the total computation time and the total communication time to obtain the training time; and multiply the training time by the number of training iterations of the large language model to obtain the total training time of the large language model. The communication capture module is used to capture communication time based on the computational communication graph, obtain the total communication time, and send the total communication time to the performance model module.

8. A parallel training simulation device for a large language model, characterized in that, The device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the parallel training simulation method for a large language model as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, which, when invoked by a computer, are used to execute the parallel training simulation method for a large language model as described in any one of claims 1 to 7.

10. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the parallel training simulation method for large language models as described in any one of claims 1 to 7.