Method, device and system for executing AI training simulation in batches

By designing a batch execution AI training simulation system including system simulator, in-machine communication simulator and cross-machine network simulator, using SPME parallel execution strategy and ECS abstract model, the problem of low parallel execution efficiency of AI training simulation experiments in the existing technology is solved, and efficient simulation experiment execution and cache efficiency is achieved.

CN120046702AInactive Publication Date: 2025-05-27BEIJING YUSUN NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510191043.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to efficiently perform a large number of AI training simulation experiments in parallel, resulting in too long time exploring the design space, especially in large-scale GPU clusters.

Method used

Design a system for batch execution of AI training simulation, including system simulator, in-machine communication simulator and cross-machine network simulator, model the AI ​​training process through ECS abstraction, realize the SPME parallel execution strategy, uniformly process the steps of the simulation process, and package data to obtain the advantages of parallelism and cache efficiency.

Benefits of technology

Through the SPME parallel execution strategy, the process synchronization overhead is reduced, the cache hit rate is improved, the cross-experiment batching and parallelization opportunities are fully utilized, and the efficiency of simulation experiments is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046702A_ABST
    Figure CN120046702A_ABST
Patent Text Reader

Abstract

The invention discloses a method, device and system for executing AI training simulation in batches, the system for executing AI training simulation in batches comprises a system simulator, a built-in communication simulator and a cross-machine network simulator, the system simulator is used for controlling and scheduling the AI training simulation process, and the built-in communication simulator is used for controlling and scheduling the AI training simulation process. The built-in communication simulator is used for carrying out collective communication operation in the server between the GPUs, and the cross-machine network simulator is used for registering a plurality of point-to-point cross-network communication in the system simulator. According to the method, the process synchronization overhead is reduced, the cache hit rate is increased, cross-experiment batch processing and parallelization opportunities neglected by an existing method can be fully utilized, the system can package data in a mode of uniformly processing the steps of the simulation process so as to obtain the advantages of parallelism and cache efficiency, and during the system execution period, the system can be quickly executed. And the adjacent threads of the continuous entities in the processing table keep consistent access to the component data, so that the advantage of cache efficiency is brought.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of AI training simulation, and specifically to a method, device, and system for batch execution of AI training simulation. Background Art

[0002] The growth of AI model scale requires a huge training system; currently, relevant companies are building clusters with more than 24,000 GPUs and will soon launch O(100k) GPU clusters. Therefore, the design space for LLM training has become wider and deeper, including parallelization strategies, collective communication primitive parameters, congestion control algorithms and parameters, architecture topology design, etc.; since all these options can interact with each other in complex and unpredictable ways, finding the best design point of the training system requires a large number of simulation experiments using different option combinations to fully search the design space. For example, exploring the optimal parallel group size requires nearly 100 experiments, while determining the best topology for connecting large-scale GPUs requires more than 10k experiments; how to quickly complete these experiments has become very challenging.

[0003] The running of these experiments is usually strictly parallel, which can be utilized to reduce the end-to-end exploration time. The problem lies in how to efficiently execute these experiments in parallel? The obvious answer, running n experiments on n CPU cores, cannot meet the current performance requirements, especially considering the competition for shared resources such as memory capacity and CPU low-level caches means that scaling is sublinear. Therefore, the present invention designs a method, device, and system for batch execution of AI training simulation to consider how to optimize all experiments together. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, device, and system for batch execution of AI training simulation to solve the problems raised in the above background art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A system for batch execution of AI training simulation, including a system simulator, an in-machine communication simulator, and an inter-machine network simulator. The system simulator is used to control and schedule the process of AI training simulation. The in-machine communication simulator is used for collective communication operations within the server between GPUs. The inter-machine network simulator is used to register a number of point-to-point cross-network communications in the system simulator.

[0006] Preferably, the input of the system simulator is the model workload, which characterizes the computational graph of each GPU; the nodes of the system simulator are computational operations or collective communication operations.

[0007] Preferably, for the collective communication operation between servers, the system emulator generates a series of point-to-point send and receive streams by implementing the collective communication algorithm of NCCL, and the system emulator hijacks the NCCL API to analyze the start and end times of each communication in the CCA operation.

[0008] A method for batch-executing AI training simulations, comprising the following steps:

[0009] S1: The user writes application code, sets up the AI training system, and models the AI training process using ECS abstraction. The modeling includes the modeling of entities, components, and systems, and also includes the construction of the system execution graph;

[0010] S2: Specify multiple experiments to be run; the AI training system automatically performs multi-experiment simulations on the GPU.

[0011] Preferably, the modeling of the entities and components: In the context of AI training, the network entities in DONS are retained, and at the same time, key new entities are introduced. The state of an entity is characterized by the values of its components. The components defined for the task entity include type, load, predecessor node, and successor node; the key new entities are distinguished by their state, congestion control variables, and send buffer components. Entities of the same type with the same components are considered to share a prototype.

[0012] Preferably, the modeling of the system: The system represents data parallel computing performed on a set of entities. The system is described by a query that specifies the component data included in the input and the functions to be performed on this data. The query is designed to select entities that possess a set of predetermined components; at each step, the system checks the completion status of the predecessor nodes. If completed, it will activate the subsequent task nodes. For computational tasks, the clock is incremented to advance the simulation process; for inter-server communication tasks, a new flow is registered in the packet-level network simulator; for intra-server communication tasks, the system executes an analytical model to estimate the communication time.

[0013] Preferably, the system constructs an execution graph: The system execution graph defines the entire set of ECS systems required to be executed in the simulation step, which executes eight systems: Schedule, AnalyticalSys, SendSys, NICSndSys, ForwardSys, TransmitSys, NICRcvSys, and ACKSys. After the Task entity executes the Schedule system, AnalyticalSys simulates the in-server communication, and then injects new flows into the network simulator. The SndFlow entity executes the Send system in the network simulator to send data packets to the corresponding destinations. Then, the data packets pass through the NIC and the forwarding path composed of consecutive IngressPort entities and EgressPort entities. Finally, the data packets reach the RcvFlow entity, which, if necessary, sends back an acknowledgment. After this round ends, the next simulation step begins, and all eight systems are repeatedly executed until the simulation ends.

[0014] Preferably, the AI training system compiles the system into the GPU using the NVIDIA CUDA C++ compiler and uses the single-instruction multiple-threading programming model in CUDA to directly map each system call to a single GPU thread.

[0015] An apparatus for batch-executing AI training simulations includes a system emulator, an in-machine communication emulator, and a cross-machine network emulator.

[0016] Compared with the prior art, the beneficial effects of the present invention are:

[0017] 1) SPME parallel execution strategy: Reduces the process synchronization overhead and improves the cache hit rate;

[0018] 2) ECS modeling for large model training systems: Can make full use of the cross-experiment batch processing and parallelization opportunities ignored by existing methods. The system can package data (such as data packets) in a unified way of processing the steps of the simulation process (for example, running all the in-ports of all switches simultaneously) to obtain the advantages of parallelism and cache efficiency;

[0019] 3) Unified component memory management mechanism across simulation experiments; During system execution, adjacent threads of consecutive entities in the processing table maintain consistent access to component data, bringing the advantage of cache efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Are four parallel execution strategy diagrams;

[0021] Figure 2 Is the system architecture diagram of the present invention;

[0022] Figure 3Examples of entities, components, and systems of the present invention;

[0023] Figure 4 Execution diagram of the system of the present invention. Detailed implementation manners

[0024] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0025] Please refer to Figures 1-4 , the present invention provides a technical solution: a system for batch execution of AI training simulation, including a system emulator, an in-machine communication emulator, and a cross-machine network emulator. The system emulator is used to control and schedule the process of AI training simulation. The in-machine communication emulator is used for collective communication operations within the server between GPUs. The cross-machine network emulator is used to register a number of point-to-point cross-network communications in the system emulator. Users only need to write application codes, set the AI training system, including workload, CCL parameters, topology, etc., and specify multiple experiments to be run; then this patent automatically performs multi-experiment simulation on the GPU.

[0026] In the present invention, the input of the system emulator is the model workload, which represents the computational graph of each GPU; the nodes of the system emulator are computational operations or collective communication operations. This patent assumes that the computational operations have been marked with computational times as the workload generated by Chakra. This patent supports typical parallel strategies (such as TP, PP, and DP). For collective communication operations between servers, the system emulator generates a series of point-to-point send and receive streams by implementing the collective communication algorithm (CCA) of NCCL. However, within a given CCA, the inherent overhead of the NCCL software stack will affect the startup time of each stream. To improve the simulation accuracy, this patent hijacks the NCCL API to analyze the start and end times of each communication in the CCA operation. Therefore, this patent calibrates the simulation of collective communication between servers by introducing the measured overhead.

[0027] In-machine communication emulator: For collective communication operations within the server between GPUs (such as TP communication), this patent directly performs simulation based on the proposed analysis model, which has different empirical parameters according to the operator type and GPU type. This model provides fast and accurate in-machine communication simulation.

[0028] Cross-machine network emulator: The system emulator registers many point-to-point cross-network communication requirements in this network emulator, which will perform discrete event simulation (DES) to strictly execute packet-level events and ensure correctness, such as NS-3 and DONS. When the process is completed, the system emulator will receive a notification.

[0029] A method for batch-executing AI training simulation, comprising the following steps:

[0030] S1: The user writes application code, sets up the AI training system, and models the AI training process using ECS abstraction. The modeling includes the modeling of entities, components, and systems, and also includes the construction of the system execution graph.

[0031] S2: Specify multiple experiments to be run; the AI training system automatically performs multi-experiment simulation on the GPU.

[0032] In the present invention, the modeling of the entities and components: In the context of AI training, this patent retains the network entities in DONS, such as Sender, IngressPort, EgressPort, and Receiver, while introducing key new entities, such as training Task and Flow. The state of an entity is characterized by the values of its components. For example, as Figure 3 shown, the components defining a task entity include type (computation or communication), load (computation time or traffic), predecessor nodes, and successor nodes. A Flow entity is distinguished by its components such as state, congestion control variables, and send buffer. Entities of the same type with the same components are considered to share a prototype. Figure 3 Illustrates three prototypes of the AI training system and the example component values of 4 entities for each prototype.

[0033] In the present invention, the modeling of the system: The system represents data parallel computing performed on a set of entities. The system is described by a query that specifies the component data included in the input and the functions to be performed on this data. The query aims to select entities that possess a set of predetermined components; for example, the scheduling system (as Figure 3 shown) is executed by task. In each step, the system checks the completion status of the predecessor nodes. If completed, it will activate the subsequent task nodes. For computing tasks, the clock is incremented to advance the simulation process; for inter-server communication tasks, a new flow is registered in the packet-level network simulator; for intra-server communication tasks, the system executes an analytical model to estimate the communication time.

[0034] In the present invention, the system constructs an execution graph: the system execution graph defines the entire set of ECS systems required for execution in the simulation step, which executes eight systems: Schedule, AnalyticalSys, SendSys, NICSndSys, ForwardSys, TransmitSys, NICRcvSys, and ACKSys. After the Task entity executes the Schedule system, AnalyticalSys simulates the communication within the server, and then injects a new flow into the network simulator. The SndFlow entity executes the Send system in the network simulator to send the data packet to the corresponding destination. Then, the data packet passes through the NIC and the forwarding path composed of consecutive IngressPort entities and EgressPort entities. Finally, the data packet reaches the RcvFlow entity, which, if necessary, sends back an acknowledgment. After this round ends, the next simulation step begins, and all eight systems are repeatedly executed until the simulation ends.

[0035] In the present invention, the AI training system compiles the system into the GPU using the NVIDIA CUDA C++ compiler and directly maps each system call to a single GPU thread using the single instruction multiple threads programming model in CUDA.

[0036] The present invention: By analyzing all possible parallel execution strategies, the most effective parallel execution strategy is found. First, the parallelization strategies for multiple independent simulation experiments can be classified into four types along two axes: using a single process or multiple processes to execute a simulation program, and running a single experiment or multiple experiments in a simulation program. 1. Single-Process Single-Experiment (SPSE): This is the most commonly used method in the design space exploration process, where multiple processes are launched, and each process executes one experiment. For a given experiment, NS-3 and OMNeT++ default to using a single process and a single thread, while DONS and UNISON utilize multi-threading (multi-core) to accelerate a single experiment. 2. Multi-Process Single-Experiment (MPSE): This method uses multiple processes to run a single simulation experiment, such as NS-3 or OMNeT++ with MPI. 3. Single-Process Multiple-Experiments (SPME): In one program, this strategy extends the single experiment in SPSE to multiple experiments to save the inherent overhead of processes. 4. Multi-Process Multiple-Experiments (MPME): Similarly, this method uses multiple processes to execute one program in SPME. For example, NS-3 utilizes various processes to accelerate a single program that contains multiple independent topologies running different experiments. First, MPSE and MPME require multiple processes to accelerate one or more experiments, resulting in frequent inter-process synchronization, which brings a very large context switching overhead, further leading to very low parallel efficiency. Second, SPSE executes a single experiment in a single independent process, facing problems of high process scheduling overhead and low cache hit rate. Finally, SPME has the most potential because all experiments share one process, reducing the process scheduling overhead.

[0037] Meanwhile, SPME can support the DOD design concept. The DOD design concept allows separating the process from the data and operating on the same type of data in all entities simultaneously, which enables SPME+DOD to identify cross-experiment batch processing and parallelization opportunities ignored by existing methods. The system can package data (such as data packets) in a way that uniformly processes the steps of the simulation process (e.g., running all incoming ports of all switches simultaneously) (e.g., time) to obtain parallelism and cache efficiency advantages. This patent believes that the strategy of a user running multiple experiments in a single simulation process (referred to as SPME) extends the SIMD abstraction to multi-experiment simulation. Similar to the advantages of deep learning and DONS, the advantage of this method is that it significantly reduces the inherent overhead of processes and repeated cache / memory usage. A major additional benefit is that SPME is more suitable for deployment on GPUs, whose extremely large number of cores and customization optimizations for SIMD tasks provide a considerable performance advantage for SPME tasks.

[0038] The present invention centrally manages the storage of all component data for all simulation experiments; specifically, this patent creates an in-memory table to store the component data of the same prototype entities in all experiments; to allow efficient access to consecutive components, these tables are columnar storage, and the component data is stored consecutively in memory, following standard ECS practices; since the system may access the state of each experiment when processing entities, this patent adds an implicit ExpID component to each table; the ExpID allows this patent to quickly look up the data of each experiment for each entity and provide it to the ECS system as needed; this storage scheme has two main performance advantages; first, during system execution, adjacent GPU threads processing consecutive entities in the table maintain consistent access to the component data, even if these entities belong to different experiments; if this patent maintains separate tables for each experiment, incoherent data access will occur when only a few entities in each experiment match the ECS query, because adjacent threads in the GPU warp need to store data across different tables. The second benefit is that it can reduce the memory footprint when executing a large number of experiments. Since entities can be created dynamically, the over-allocation of table storage (to accommodate the efficient addition of new entities) will be amortized across all experiments; this patent also uses a unified table design to amortize the cost of ECS metadata; for example, the query object that stores the table ID and column ID of each matching prototype only needs to be stored once across all experiments.

[0039] At a higher level, the present invention uses the NVIDIA CUDA C++ compiler to compile these systems into the GPU and uses the single-instruction multiple-thread (SIMT) programming model in CUDA to directly map each system call to a single GPU thread. To allow the present invention to manage the parallelization of entity updates and data flows in the entire system execution graph, the ECS system is implemented as a function that receives the components of a single entity as parameters. When the ECS system is added to the system execution graph, the present invention uses the provided ECS query to find all matching prototypes and the corresponding column indices for each accessed component. Using this information, the present invention generates code to manage component access and calls the ECS system function across multiple GPU threads for each matching entity; for example, the sending system needs the data in columns 3 and 4 of the flow table. Using this information, the present invention generates a system entry function that maps GPU threads to table row indices and passes the obtained component data into the ECS system function. Given a table with N rows, the present invention will call the system entry function N times, where each call causes a single GPU thread to execute the ECS system function. In the case of the sending system, the query matches N rows in the flow table, so the present invention will call the entry function N times. This patent provides the first N call column pointers for the Flow table.

[0040] The present invention adopts an SPME parallel execution strategy: reducing process synchronization overhead and improving cache hit rate; modeling ECS for large model training systems: being able to fully utilize cross-experiment batch processing and parallelization opportunities ignored by existing methods. The system can pack data (such as data packets) in a way that uniformly processes the steps of the simulation process (for example, running all ingress ports of all switches simultaneously) (time) to obtain the advantages of parallelism and cache efficiency. A unified component memory management mechanism across simulation experiments; during system execution, adjacent threads of consecutive entities in the processing table maintain consistent access to component data, bringing the advantage of cache efficiency.

[0041] The content not described in detail in this specification belongs to the prior art well-known to those skilled in the art. Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A system for batch execution of AI training simulation, characterized in that: It includes a system simulator, an intra-machine communication simulator and a cross-machine network simulator. The system simulator is used to control and schedule the AI ​​training simulation process, the intra-machine communication simulator is used to perform collective communication operations within the server between GPUs, and the cross-machine network simulator is used to register several point-to-point cross-network communications in the system simulator.

2. A system for batch execution of AI training simulation according to claim 1, characterized in that: The input of the system simulator is a model workload, which represents the computation graph of each GPU; the nodes of the system simulator are computation operations or collective communication operations.

3. A system for batch execution of AI training simulation according to claim 1, characterized in that: For inter-server collective communication operations, the system simulator generates a series of point-to-point sending and receiving flows by implementing the collective communication algorithm of NCCL. The system simulator hijacks the NCCL API to analyze the start and end time of each communication in the CCA operation.

4. A method for batch execution of AI training simulation according to any one of claims 1 to 3, characterized in that: The steps include: S1: The user writes application code, sets up the AI ​​training system, and uses ECS abstraction to model the AI ​​training process. The modeling includes the modeling of entities, components, and systems, as well as the construction of the system execution graph. S2: Specify multiple experiments to run; the AI ​​training system automatically performs multi-experiment simulations on the GPU.

5. The method for batch execution of AI training simulation according to claim 4, characterized in that: Modeling of the entities and components: In the context of AI training, the network entities in DONS are retained, and key new entities are introduced. The state of an entity is characterized by the values ​​of its components. The components that define a task entity include type, load, predecessor node, and successor node; key new entities are distinguished by their state, congestion control variables, and send buffer components. Similar entities with the same components are considered to share prototypes.

6. A method for batch execution of AI training simulation according to claim 4, characterized in that: Modeling of the system: The system represents a data-parallel computation performed on a collection of entities. The system is described by a query that specifies the component data contained in the input and the function performed on this data. The query aims to select entities with a set of predetermined components. At each step, the system checks the completion status of the predecessor node. If completed, it activates the subsequent task node. For the computational task, the clock is increased to advance the simulation process. For inter-server communication tasks, new flows are registered in a packet-level network simulator; for intra-server communication tasks, the system executes analytical models to estimate communication time.

7. The method for batch executing AI training simulation according to claim 4, characterized in that: Construction of the system execution graph: The system execution graph defines the entire set of ECS systems required to be executed in the simulation step, which executes eight systems: Schedule, AnalyticalSys, SendSys, NICSndSys, ForwardSys, TransmitSys, NICRcvSys and ACKSys. After the Task entity executes the Schedule system, AnalyticalSys simulates the communication within the server, and then injects the new flow into the network simulator. The SndFlow entity executes the Send system in the network simulator to send the data packet to the corresponding destination, and then the data packet passes through the NIC and the forwarding path composed of consecutive IngressPort entities and EgressPort entities. Finally, the data packet arrives at the RcvFlow entity. If necessary, the entity will send back a confirmation. After this round ends, the next simulation step starts, and all eight systems will be repeatedly executed until the simulation ends.

8. The method for batch executing AI training simulation according to claim 4, characterized in that: The AI ​​training system uses the NVIDIA CUDA C++ compiler to compile the system into the GPU and uses the single instruction multi-threaded programming model in CUDA to map each system call directly to a single GPU thread.

9. A device for batch execution of AI training simulation, characterized in that: It includes the system simulator, the intra-machine communication simulator and the inter-machine network simulator as described in any of claims 1-3.