Parallel strategy search method for efficient training of artificial intelligence large model

Through the method of NeRF model and hybrid integer planning combined with two-layer policy network, the problems of low coordination efficiency of multi-dimensional parallel strategy and poor dynamic adaptability of hardware topology in artificial intelligence large model training are solved, efficient communication prediction and strategy optimization are achieved, and training efficiency and system robustness are improved.

CN120012879AActive Publication Date: 2025-05-16西安圣瞳科技有限公司

Patent Information

Application Number
CN202510488064.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the efficient training of artificial intelligence large models, the existing parallel strategy search method ignores the physical topological effect of heterogeneous hardware, resulting in large deviations in communication overhead estimates, and it is difficult to model multi-strategy synergy effects, which are prone to falling into local optimization, especially under the interweaving of video memory and communication bottlenecks in large-scale clusters, the robustness of the strategy is limited.

Method used

By entering the physical topology parameters of the cluster, a three-dimensional hardware topology field is generated using the NeRF model, annotating the calculation diagram, building a hybrid integer planning model and a two-layer policy network, outputting a collection of parallel policies, and verifying and adjusting in the simulator, and finally deploying to the real cluster and feedback updating the model.

Benefits of technology

It significantly improves the communication time accuracy, reduces the single iteration time of GPT-3-level model training, avoids local optimal traps, reduces memory fragmentation rate, improves GPU utilization, shortens strategy re-optimization time, and reduces system throughput fluctuations, achieving multi-objective balance between training efficiency, memory utilization rate and system robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012879A_ABST
    Figure CN120012879A_ABST
Patent Text Reader

Abstract

The invention discloses a parallel strategy search method for efficient training of an artificial intelligence large model, and relates to the technical field of artificial intelligence, and the method comprises the steps: generating a three-dimensional hardware topology field through training a NeRF model; analyzing the calculation graph of the large model to be trained, mapping operator nodes into the three-dimensional hardware topology field, and generating a labeling calculation graph; constructing a mixed integer programming model based on the annotation calculation graph, and outputting a candidate strategy set; inputting the candidate strategy set into a double-layer strategy network, and outputting a parallel strategy set; loading the three-dimensional hardware topology field in the simulator, verifying and adjusting the parallel strategy set, and outputting an optimal strategy; and deploying the optimal strategy to the real cluster, and triggering a re-screening strategy. According to the method, a physical-level hardware communication field is modeled based on a neural radiation field, mixed integer programming and deep reinforcement learning are fused, parallel strategy global-local collaborative optimization is realized, and multi-target balance of training efficiency, video memory utilization rate and system robustness is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a parallel strategy search method for efficient training of large artificial intelligence models. Background Art

[0002] As the parameter scale of large artificial intelligence models continues to expand, parallel training based on distributed clusters has become a key technical path. Existing technologies mainly optimize basic strategies such as data parallelism, model parallelism, and pipeline parallelism, and implement strategy search through heuristic algorithms or reinforcement learning. For example, Google proposed GPipe to reduce video memory usage through micro-batch pipeline division, and NVIDIA Megatron-LM is based on tensor parallel optimization model segmentation. However, such methods have significant limitations: first, the physical topology effects of heterogeneous hardware (such as cross-rack communication attenuation and NVLink bandwidth heterogeneity) are ignored when optimizing traditional parallel strategy combinations, resulting in large deviations in communication overhead estimates; second, the existing hierarchical optimization framework (such as dividing the pipeline first and then optimizing data parallelism) is difficult to model multi-strategy synergy effects and is prone to local optimality, especially under the conditions of interweaving video memory and communication bottlenecks in large-scale clusters, the robustness of the strategy is severely limited. Summary of the invention

[0003] In view of the above existing problems, the present invention is proposed.

[0004] Therefore, the present invention provides a parallel strategy search method for efficient training of large artificial intelligence models to solve the problems of low coordination efficiency of multi-dimensional parallel strategies and poor dynamic adaptability of hardware topology in large-scale distributed training.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a parallel strategy search method for efficient training of large artificial intelligence models, which includes inputting physical topology parameters of a cluster, training a NeRF model through multi-view bandwidth and delay data, and generating a three-dimensional hardware topology field including bandwidth density and communication attenuation coefficient; Parse the computational graph of the large model to be trained, map the operator nodes to the three-dimensional hardware topology field, annotate each operator node with a dynamic communication cost label according to the bandwidth density and communication attenuation coefficient of the operator node, and generate an annotated computational graph; Construct a mixed integer programming model based on the annotated computational graph and output a set of candidate strategies that satisfy the communication field constraints; The candidate strategy set is input into the two-layer strategy network, and the parallel strategy set including the global scheduling framework and local adaptive rules is output; Load the 3D hardware topology field in the simulator, verify and adjust the parallel strategy set, and output the optimal strategy; Deploy the optimal strategy to the real cluster and collect training data, provide feedback to update the NeRF model, and trigger the mixed integer programming model to re-screen the strategy.

[0006] As a preferred solution of the parallel strategy search method for efficient training of large artificial intelligence models described in the present invention, the physical topology parameters include rack layout, connection relationship between nodes and physical distance.

[0007] As a preferred solution of the parallel strategy search method for efficient training of large artificial intelligence models described in the present invention, the operator node includes the operator's computing amount, video memory requirements and dependencies.

[0008] As a preferred solution of the parallel strategy search method for efficient training of large artificial intelligence models described in the present invention, the mixed integer programming model constructed based on the labeled computational graph is to integrate the computational time, communication time and constraints through the objective function, and to enumerate the pipeline stage division and data parallel sharding strategies through the quantum annealing algorithm.

[0009] As a preferred solution of the parallel strategy search method for efficient training of large artificial intelligence models described in the present invention, wherein: the candidate strategy set is input into a two-layer strategy network, and the parallel strategy set output including a global scheduling framework and local adaptive rules is generated by using a Transformer encoder to perform global attention modeling on the spatiotemporal features of the annotated calculation graph to generate a preliminary scheduling strategy, the cellular automaton network performs local rule adjustment based on the states of adjacent nodes, and the Transformer encoder and the cellular automaton network are jointly optimized through a differentiable interface.

[0010] As a preferred solution of the parallel strategy search method for efficient training of large artificial intelligence models described in the present invention, the verification and adjustment of the parallel strategy set refers to using the low-fidelity mode to screen potential high-quality strategies, and the high-fidelity mode to simulate cross-node communication paths through ray tracing, and dynamically adjust task allocation in combination with the cellular automaton rule base.

[0011] As a preferred solution of the parallel strategy search method for efficient training of large artificial intelligence models described in the present invention, the training data includes communication efficiency, video memory fluctuation data and hardware abnormal events.

[0012] As a preferred solution of the parallel strategy search method for efficient training of large artificial intelligence models described in the present invention, the training data is divided into two feedback paths, the first path inputs the NeRF model to incrementally update the three-dimensional hardware topology field parameters, and the second path inputs the two-layer strategy network. By comparing the deviation between the simulated prediction value and the measured value, the dependence weight of the Transformer attention mechanism on the hardware field characteristics is adjusted, and the local rule base of the cellular automaton is evolved.

[0013] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the parallel strategy search method for efficient training of large artificial intelligence models as described in the first aspect of the present invention.

[0014] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the parallel strategy search method for efficient training of large artificial intelligence models as described in the first aspect of the present invention.

[0015] The beneficial effects of the present invention are: based on the neural radiation field (NeRF) modeling of physical-level hardware communication field, the integration of mixed integer programming and deep reinforcement learning, the global-local collaborative optimization of parallel strategies is realized, the three-dimensional bandwidth density field is constructed through the NeRF model, the cross-rack link attenuation effect is quantified, and the accuracy of predicting communication time is significantly improved. In the GPT-3 level model training, the cluster TPI (single iteration time) is reduced, and the mixed integer programming (MIQP) is used to jointly optimize the pipeline stage division, data parallel sharding and calculation-communication overlap factor to avoid the local optimal trap caused by traditional hierarchical optimization. The actual measurement shows that the video memory fragmentation rate is significantly reduced, and the GPU utilization rate is improved. Based on the Transformer-cellular automaton architecture, combined with the online incremental learning mechanism, the strategy re-optimization after the hardware topology change (such as node expansion and contraction) is completed, the fault recovery time is shortened, and the system throughput fluctuation is reduced compared with the static strategy. This solution breaks through the limitation of the traditional method of only optimizing a single performance indicator and achieves a multi-objective balance of training efficiency, video memory utilization and system robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0017] Figure 1 This is an operation diagram of the parallel strategy search method for efficient training of large artificial intelligence models in this embodiment.

[0018] Figure 2 This is a flow chart of the parallel strategy search method for efficient training of large artificial intelligence models in this embodiment.

[0019] Figure 3 This is an operation diagram of the calculation diagram annotation and optimization model in this embodiment.

[0020] Figure 4 This is a diagram of the double-layer strategy network structure in this embodiment. DETAILED DESCRIPTION

[0021] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0022] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0023] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0024] In this embodiment, refer to Figure 1~Figure 4 This embodiment provides a parallel strategy search method for efficient training of large artificial intelligence models, including the following steps: S1. Input the physical topology parameters of the cluster, train the NeRF (neural radiation field) model through multi-view bandwidth and delay data, and generate a three-dimensional hardware topology field including bandwidth density and communication attenuation coefficient.

[0025] Physical topology parameters include rack layout, node connection relationships, and physical distances; It should be noted that the rack layout uses a laser rangefinder or RFID tag positioning system (radio frequency identification tag) to obtain the rack number, location coordinates, size, and internal GPU node distribution through the data center infrastructure management system or the physical planning drawings during cluster deployment. The connection relationship between nodes is obtained through cluster management tools (such as the Kubernetes node topology application interface) or network switch configuration tables (such as the link layer discovery protocol). Use commands such as lshw (list hardware) and ibstat (InfiniBand status tool) to query the InfiniBand / NVLink (InfiniBand / NVIDIA high-speed interconnect technology) connection status and generate a topology diagram of direct connections between nodes or cross-switch connections. Physical distance: The distance between nodes in a rack is determined by the rack design drawing (e.g., the distance between adjacent GPU nodes is 0.1 meters), and the distance across racks is measured by a laser rangefinder. Physical distance directly affects the signal attenuation model (e.g., the attenuation coefficient is proportional to the square of the distance), and the rack layout determines the electromagnetic interference area (e.g., the signal attenuation at the edge of a metal rack increases by 30%). The node connection relationship determines the communication path selection (e.g., the NVLink direct connection bandwidth is 256GB / s, and the cross-switch bandwidth is reduced to 100GB / s), and the bandwidth density field is generated in combination with the physical distance. According to the physical topology, high attenuation paths (e.g., cross-rack paths with a distance of >3 meters) are avoided, and low-latency links in the same rack are preferred, reducing the communication time by 40%. When a node fails, it quickly switches to the backup node in the same rack based on the connection relationship (e.g., migrating to an adjacent GPU with a distance of 0.2 meters), and the recovery time is shortened to seconds. Physical distance limits the data parallel sharding strategy (e.g., cross-rack sharding must meet the bandwidth density ≥100Gbps / m³), avoiding communication bottlenecks caused by long-distance transmission.

[0026] Based on the physical topology parameters of the cluster, a distributed probe program is deployed in the cluster to perform cross-node bandwidth testing (test tool: Network Bandwidth Testing Tool Version 3), covering all connection paths (same rack, across racks, and across switches), obtaining multi-perspective bandwidth data, and obtaining one-way delays between nodes through Ping (network connectivity testing tool) testing, and collecting delay data; Integrate multi-view bandwidth data and delay data to form a multi-view bandwidth-delay data set; The cluster physical space is divided into a three-dimensional grid (resolution 0.1 m), and each grid point is associated with bandwidth density and communication attenuation coefficient; Input a multi-view bandwidth-delay dataset and train the NeRF model to learn the nonlinear mapping relationship from spatial coordinates to communication performance. The optimization goal is to minimize the mean square error between the predicted bandwidth density and the actual value. It should be noted that during the training process, the NeRF model learns how to map the spatial coordinates to the expected communication performance values. In particular, the optimization process focuses on minimizing the mean square error between the bandwidth density predicted by the NeRF model and the bandwidth density actually observed. This means that the NeRF model needs to accurately capture the changes in communication quality corresponding to each point in space, so that it can provide accurate communication performance predictions for unobserved spatial coordinates. By continuously iteratively updating the NeRF model parameters, the prediction results are made as close to the true values ​​as possible, thereby optimizing network planning and performance evaluation.

[0027] Compare the NeRF model prediction value with the actual test data. If the NeRF predicted bandwidth density error exceeds 5%, add probe sampling points in the area with large error (e.g. error > 5%) (e.g. rack edge) and retrain the NeRF model. Outputs a three-dimensional hardware topology field (in HDF5 format), including bandwidth density, communication attenuation coefficient, and coordinate mapping table for each spatial point.

[0028] It should be noted that the NeRF model is trained with multi-view bandwidth and delay data to model the impact of physical topology. NeRF is used to convert parameters such as rack layout and physical distance into bandwidth density and communication attenuation coefficient of the three-dimensional field, and accurately quantify the nonlinear attenuation of cross-node communication (such as the error of cross-rack path attenuation coefficient <5%). Through the error calibration mechanism (such as adding sampling points when the error is >5%), the trained NeRF field can dynamically adapt to hardware changes (such as adding new nodes) to avoid policy failures caused by topology changes.

[0029] S2. Analyze the computational graph of the large model to be trained, map the operator nodes to the three-dimensional hardware topology field, label each operator node with a dynamic communication cost label according to the bandwidth density and communication attenuation coefficient of the operator node, and generate a labeled computational graph.

[0030] The operator node contains the operator's computational effort, video memory requirements, and dependencies; Enable the context manager in the PyTorch performance analyzer, start the advanced performance analysis mode, perform a single forward + back propagation for the large model to be trained (such as the Meta large language model), generate a trace file containing detailed operator execution information, and extract the metadata of each operator node, including the operator's computational workload, video memory requirements, and dependencies; It should be noted that the computational amount is calculated by directly reading the calculated value. If the computational graph is not provided, it is dynamically calculated by the operator type (such as two-dimensional convolution layer, matrix multiplication) and input dimension (such as the computational amount of the MatMul operator = 2×Batch×M×N×K, where M is the row of the left matrix, N is the column of the right matrix, K is the column of the left matrix, and Batch is the number of data samples), and normalized; the video memory demand is calculated by recording the input and output tensor dimensions of the operator, calculating the number of bytes according to the data type, and calculating the input video memory demand by the sum of the bytes of all input tensors, and calculating the output video memory demand by the sum of the bytes of all output tensors. The peak video memory demand is calculated using the maximum value of the video memory temporarily allocated during the execution of the operator, and the creation and release timestamps of each tensor are recorded. The video memory life cycle (unit: microseconds); the dependency relationship includes the following steps: load the trace file generated by the Profiler, extract all operator nodes and thread / stream IDs in the order of execution, including predecessor operators and successor operators. The predecessor operator is the operator that has not been completed in the same thread before the current operator, and the successor operator is the operator that starts executing earliest after the current operator in the same thread. For multi-GPU / multi-thread scenarios, distinguish different execution flows. If the output tensor of operator A is used by operator B (located on a different GPU), mark it as a cross-device dependency. Detect circular dependencies in the calculation graph through depth-first search, and throw an exception if they exist (large models are usually directed acyclic graphs). The longest path from input to output, mark all operators on the path, and output the dependency table, which includes the predecessor list, successor list, and whether each operator is on the critical path. The allocation strategy is cyclically allocated in layer order to bind model layers to physical GPU nodes (e.g., 64 layers of LLaMA are allocated to 8 GPUs, 8 layers per GPU); Query the bandwidth density and communication attenuation coefficient of the corresponding position of each operator node in the three-dimensional hardware topology field, add a communication cost label to each operator node, and calculate the communication cost using the following formula: Communication cost = amount of data transmitted / (bandwidth density × attenuation compensation factor), where attenuation compensation factor = 1 / (1+accumulated value of path attenuation coefficient), and the path attenuation coefficient comes from the accumulated attenuation value of the communication path in the three-dimensional hardware field; Construct a 32-dimensional joint feature vector using computational load, video memory requirements, and communication cost; It should be noted that the computing load (8 dimensions) includes operator type encoding (4 dimensions), logarithmic normalization value of computing amount (1 dimension), parallelism (1 dimension) and dependency depth (2 dimensions); the video memory requirement (8 dimensions) includes input / output tensor size (2 dimensions), video memory peak (1 dimension), life cycle (2 dimensions) and reuse times (1 dimension); the communication cost (16 dimensions) includes cross-device times (1 dimension), maximum single communication cost (1 dimension) and critical path marking (1 dimension), etc. Output annotated computation graph with hardware features (JSON file, including 32-dimensional joint feature vector and physical binding information).

[0031] It should be noted that the computational graph is parsed and hardware features are integrated to annotate the operator communication cost. Through the 32-dimensional joint feature vector (computational load + video memory requirement + communication cost), the mapping relationship between operators and hardware resources is established. In the GPT-3 level model, the scheduling matching degree of key path operators is improved by 45%. Based on the attenuation compensation factor (1 / (1+path attenuation)), the communication priority is dynamically adjusted to reduce the data transmission volume of high attenuation paths. Experiments show that the communication conflict rate is reduced by 38%.

[0032] S3. Construct a mixed integer programming model based on the annotated computational graph and output a set of candidate strategies that satisfy the communication field constraints.

[0033] The mixed integer programming model based on the annotated computational graph is constructed by integrating the computation time, communication time and constraints through the objective function, and enumerating the pipeline stage division and data parallel sharding strategy through the quantum annealing algorithm; Define discrete variables and continuous variables. Discrete variables include the number of pipeline stages and the number of data parallelism. The number of pipeline stages is an integer variable with a value range of [1,16], which means that the mixed integer programming model is divided into 1 to 16 pipeline stages. The number of data parallelism is an enumeration value {1,2,4,8}, which indicates the number of data shards (the enumeration value must satisfy ≤ the total number of GPUs). The continuous variable [0,1] indicates the time overlap ratio of computation and communication through the computation-communication overlap factor (0 means no overlap and 1 means complete overlap). The objective function is to minimize the single iteration time, which is expressed as: ; in, represents a single iteration time, Indicates The total amount of computation in the stage (extracted from the annotated computation graph), Indicates The amount of data communicated across devices during the phase (calculated based on the operator binding location), Represents the first The current path bandwidth density of the stage (unit: Gbps / m³), The index variable representing the phase division in the pipeline parallelism is used to traverse the time consumption of all stages. represents the total number of stages of parallel partitioning of the pipeline, The adjustment factor representing the overlap of computation and communication time directly affects the actual proportion of communication overhead. When it is 1, it means that the communication time is completely covered by the computing time (ideally, the communication overhead is zero). When it is 0, it means that the calculation and communication are completely serial (no overlap, and the communication time is all extra consumption). represents the attenuation compensation factor, =1 / (1+accumulated value of path attenuation coefficient), the path attenuation coefficient comes from the accumulated attenuation value of the communication path in the three-dimensional hardware field, The GPU computing power is the measured floating point computing power (FLOP / s) based on the MLPerf benchmark, and the real-time utilization is obtained by running standard operators such as GEMM (general matrix multiplication); The constraints include video memory constraints and communication path constraints. The video memory constraint is that the video memory usage of a single card is ≤40GB (based on the cumulative video memory requirements of the annotated computational graph); the communication path constraint is that the bandwidth density of the cross-shard communication path is ≥100Gbps / m³ (verified from the three-dimensional hardware topology field); The combinatorial optimization problem of discrete variables and continuous variables is converted into a quadratic unconstrained binary optimization model, embedded in a D-Wave quantum computer (quantum annealing hardware), and the quantum annealing algorithm is run to screen out candidate solutions that meet the video memory constraints and bandwidth constraints, and the top 50 quantum annealing solutions are extracted as the initial solutions of the classical optimizer; For each quantum annealing solution, the discrete variables and continuous variables are fixed, and the continuous variables are optimized to minimize the single iteration time. The algorithm is terminated when the change of the continuous variable is less than 0.01 or 100 iterations are reached. It should be noted that when performing quantum annealing calculations to solve optimization problems, a series of possible solutions will first be obtained. For each solution obtained from the quantum annealing process, we must first identify and fix the values ​​of discrete variables and continuous variables. The so-called fixation means that these variables are temporarily regarded as constants in the further optimization process. Next, the focus is on optimizing those continuous variables marked as adjustable, in order to reduce the time required for a single iteration, thereby speeding up the entire solution process. It is necessary to carefully analyze which continuous variables in the current solution have a significant impact on the iteration time, and find the optimal values ​​of these variables through appropriate methods (such as local search, gradient descent, etc.), so as to shorten the time consumption of each iteration as much as possible while maintaining the quality of the solution. The ultimate goal is to improve the overall efficiency and practicality of the quantum annealing algorithm. It should be noted that in actual operation, the optimization process is repeatedly performed to gradually approach the optimal solution while continuously optimizing the iteration time.

[0034] Eliminate video memory overlimit and communication timeout. Specifically, remove strategies with single card video memory usage > 40GB and single iteration time > 500ms (based on high-fidelity simulation prediction). Sort in ascending order by single iteration time, and select the strategy with the shortest single iteration time. If there are duplicate discrete variables or continuous variables in the top 200 strategies, add the suboptimal solution according to the single iteration time. Output a set of candidate strategies that satisfy the communication field constraints (CSV file, each line contains fields: number of pipeline stages, number of data parallelism, overlap factor, single iteration time (ms), video memory usage (GB), shard binding list).

[0035] It should be noted that a mixed integer programming model is constructed, and quantum annealing is used to solve the Nash optimal strategy. The quantum annealing algorithm supports traversing the combinatorial space of 16-stage pipelines and 8-shard data parallelism (traditional brute force search only supports 4 stages × 4 shards), and the coverage of candidate strategies is expanded by 16 times. The constraints of video memory occupancy ≤ 40GB and bandwidth density ≥ 100Gbps / m³ eliminate invalid strategies (such as video memory overflow or communication bottleneck strategies).

[0036] S4. Input the candidate strategy set into the two-layer strategy network, and output a parallel strategy set including the global scheduling framework and local adaptive rules.

[0037] The candidate strategy set is input into the two-layer strategy network, and the output is a parallel strategy set containing a global scheduling framework and local adaptive rules. The Transformer encoder (transformer model) generates a preliminary scheduling strategy by performing global attention modeling on the spatiotemporal features of the annotated computational graph. The cellular automaton network performs local rule adjustment based on the states of adjacent nodes. The Transformer encoder and the cellular automaton network are jointly optimized through a differentiable interface. Specifically, the operator nodes of the annotated computation graph are arranged in topological order to generate a sequence. The input of each operator is a 32-dimensional joint feature vector. The operator position encoding is added to indicate the execution order in the computation graph. Eight attention heads are used, each of which focuses on feature combinations of different dimensions (e.g., head 1 focuses on computational load, and head 2 focuses on communication cost). The attention weight is calculated and expressed as: ; in, represents the attention weight, Corresponding to query, key, and value matrices respectively, represents the dimension of the key vector (column vector of K), Indicates the similarity between the query (row of Q) and the key (row of K); Identify critical paths (e.g. strong dependencies between computationally intensive operators) through attention weights; It should be noted that the 32-dimensional feature vectors of the operators in the computational graph are mapped to query Q, key K, and value V matrices, and the feature association weights are calculated separately through the 8-head attention mechanism to generate a global dependency graph. After weight fusion (learnable parameter weighting), directed weighted edges are constructed to quantify the computation and communication association across operators. Based on the global dependency graph, the DFS algorithm is improved to search for the longest path, and the weight calculation integrates the dependency strength and operator computation. The path with the highest weight accumulation value is dynamically marked as the critical path. According to the training feedback, weight penalties are imposed on the high-computation operators in the path, and they are preferentially scheduled to the low-communication attenuation area. The hot zone of the three-dimensional hardware topology field is visualized to overlay the critical path (bandwidth density <100Gbps / m³). If the path overlaps with the high-attenuation area, the cellular automaton rule is triggered (such as migrating the Conv2d operator to the GPU in the same rack), and the attention weight is adjusted through backpropagation (such as reducing the operator dependency weight of the high-attenuation path). In the training of large models such as LLaMA-70B, the critical path communication delay was reduced by 42%, and the pipeline idle time was reduced by 35%, systematically solving the problems of resource contention and efficiency loss caused by path selection deviation in traditional solutions.

[0038] The output layer generates the mini-batch size for each pipeline stage (e.g. stage 1 mini-batch = 32, stage 2 = 16); Based on the attention weights, select the GPU with the highest bandwidth density as the parameter server (such as GPU0 in rack A); Define load balancing rules and memory recycling rules. Specifically, the triggering condition for the load balancing rule is that the node load is > 80% and the load of any neighboring node is < 50%. The execution action is to migrate 15% of the computing tasks of the current node to the low-load neighbor. The triggering condition for the memory recycling rule is that the node memory margin is < 10% (for example, the remaining memory is < 4GB under 40GB of memory). The execution action is to trigger the activation of recalculation, exchanging time for space to reduce memory usage. Load the 3D hardware topology field in a single-node simulator, bind the cell state (load and video memory) to each GPU node, detect the node state in each simulation clock cycle, and immediately execute the corresponding action if the rule triggering condition is met; After the cellular automaton executes the memory recycling rule, it records the memory change, converts the memory change into a gradient signal, back-propagates it to the attention weight matrix of the Transformer, reduces the scheduling weight of the high memory consumption node, and uses the reward function to calculate the strategy reward value. After each round of training, the 20% strategies with the lowest reward value are eliminated, and 40 new strategies are generated through random perturbations to maintain a total of 200.

[0039] It should be noted that Transformer global scheduling + cellular automation local rule base. Transformer identifies critical paths (such as matrix multiplication dependency chains) through multi-head attention. In LLaMA-70B training, the bandwidth density of critical path operator allocation increased by 50%, and the peak throughput increased by 28%. Cellular rules migrate loads in real time (such as migrating 15% of tasks when node load > 80%), reducing pipeline bubbles caused by resource contention, and the average GPU utilization rate increased from 68% to 89%.

[0040] S5. Load the three-dimensional hardware topology field in the simulator, verify and adjust the parallel strategy set, and output the optimal strategy.

[0041] Verifying and adjusting the parallel strategy set refers to the low-fidelity mode screening potential high-quality strategies, and the high-fidelity mode simulates the cross-node communication path through ray tracing and dynamically adjusts the task allocation in combination with the cellular automation rule base; Assume that all communication paths use a uniform bandwidth of 100 Gbps (ignoring differences in physical topology). Calculate the single iteration time for each strategy. Directly read the video memory usage field in the strategy. If it is > 40 GB, remove it immediately. Output the initially screened strategy set (40) as high-fidelity verification input. According to the GPU binding information in the strategy, the source node (such as GPU0 in rack A) and the target node (such as GPU4 in rack B) of the communication are determined, and light is emitted from the source node. The path is selected according to the bandwidth density distribution of the NeRF field, including the optimal path and the disaster recovery path. The optimal path is transmitted along the area with the highest bandwidth density (such as the direct path within the rack), and the disaster recovery path is automatically switched to the detour path (such as through the relay node) when the attenuation coefficient of the optimal path is greater than 0.5; Calculate the total communication time, expressed as: ; in, Indicates the total communication time (unit: seconds), Indicates the total amount of data. Represents nodes for segmented transmission in a communication path (such as routers and fiber segments in a network). The location of each node is mapped by the grid coordinates of the three-dimensional hardware topology field. Representation Node The bandwidth density, Representation Node The communication attenuation coefficient of In the simulator, load balancing rules and memory recycling rules are bound to each GPU node. The node status is detected every 10ms simulation clock cycle, and the load balancing rules and memory recycling rules are triggered. When migrating tasks, the path attenuation coefficient of the target node is recalculated. The high-fidelity single iteration time is calculated by time accumulation, and the strategy with the smallest single iteration time is output as the optimal strategy.

[0042] It should be noted that low-fidelity coarse screening + high-fidelity ray tracing simulation. The low-fidelity model (assuming fixed bandwidth) quickly eliminates the video memory over-limit strategy (taking 5 minutes, traditional full high-fidelity takes 2 hours), and the efficiency is improved by 24 times. High-fidelity simulation combined with ray tracing selects the optimal path (such as avoiding areas with attenuation > 0.5), the communication time estimation error is <3%, and the strategy deployment failure rate is reduced by 72%.

[0043] S6. Deploy the optimal strategy to the real cluster and collect training data, provide feedback to update the NeRF model, and trigger the mixed integer programming model to re-screen the strategy.

[0044] The training data includes communication efficiency, video memory fluctuation data, and hardware abnormal events; It should be noted that communication efficiency is achieved by deploying distributed probe programs (such as iperf3 and NVIDIA performance analyzer), performing cross-node bandwidth tests, and recording real-time bandwidth and latency (ms); obtaining point-to-point communication throughput through NVIDIA NVLink / InfiniBand performance counters, and calculating effective bandwidth utilization (actual bandwidth / theoretical bandwidth). Video memory fluctuation data is achieved by integrating PyTorch memory analysis hooks to record the video memory allocation and release timestamps of each operator; and counting peak video memory occupancy and video memory fragmentation rate. Hardware abnormal events are captured through the Kubernetes event monitoring interface and GPU health check tools, such as node failures, NVLink disconnection, and high temperature events; the role of communication efficiency data: compare the simulated communication time with the measured value. If the deviation is >10%, update the bandwidth density parameters of the three-dimensional hardware field (such as correcting the measured bandwidth attenuation coefficient from 0.2 to 0.25) to make the model more in line with the real physical environment. Dynamically adjust the attention weight of the Transformer to reduce the scheduling priority of high-fluctuation paths (such as the standard deviation of the bandwidth of the cross-rack path >15%), and reduce the communication conflict rate by 38%. The role of video memory fluctuation data: Through the video memory peak and fragmentation rate data, the video memory recycling rules of the cellular automaton are optimized (such as adjusting the threshold for triggering Checkpointing from 10% to 8%), and the video memory overflow events are reduced by 72%. Combined with the video memory life cycle data, long-lifetime tensors are migrated to low-fragmentation GPU nodes, and the video memory reuse rate is increased by 45%. The role of hardware abnormal events: After detecting a node failure, the cellular automaton is triggered to migrate tasks to the backup node (such as the rack backup GPU), and the gradient is synchronized through the NCCL consistency protocol, and the fault recovery time is <2 seconds. The frequency of abnormal events is counted (such as NVLink disconnection >3 times per hour), and the execution priority of the fault-tolerant rules is automatically increased, and the interruption rate is reduced by 94%.

[0045] The training data is divided into two feedback paths. The first path inputs the NeRF model to incrementally update the parameters of the three-dimensional hardware topology field. The second path inputs the two-layer policy network to adjust the dependence weight of the Transformer attention mechanism on the hardware field features by comparing the deviation between the simulated prediction value and the measured value, and evolve the local rule base of the cellular automaton. Parse the gradient synchronization node list in the optimal strategy, generate communication code, generate multi-stream parallelism according to the micro-batch division scheme, execute the compiled code on the test cluster, and verify the functional correctness (such as the consistency of global reduction operation results); It should be noted that the gradient synchronization node information (such as node IP, GPU UUID and NCCL communication port) is extracted from the optimal strategy, and the globally unique communication group ID is generated based on the node physical location sorting. The NCCL library (such as NCCL global reduction operation) is called to automatically generate the communication code. According to the micro-batch division scheme (such as micro-batch = 32 for each pipeline stage), an independent CUDA stream is allocated for each stage, and the CUDA stream wait event is inserted to achieve computation-communication synchronization. According to the number of pipeline stages (S=8), 8 CUDA stream wait events are created and bound to the corresponding GPU nodes. The asynchronous data transmission between stages is achieved through the CUDA asynchronous memory copy function. Kubernetes (container orchestration system) is used to generate a deployment list, inject environment variables (such as NCCL network interface binding environment variables), start distributed tasks in the test cluster, and monitor the NCCL communication status. Initialize random tensors (size 1024×1024), compare the results of each node after executing AllReduce, calculate the relative error, and judge failure if the relative error is >1e-6. Simulate node failure to verify whether the cellular automaton rule migrates tasks to the backup node within 5 seconds and restores gradient synchronization.

[0046] Use Kubernetes deployment tools (such as the Kubernetes command-line tool) to distribute the compiled package to the target GPU node, start the distributed training task, and inject the local rule library (such as load balancing rules); Monitor the status of cluster nodes. If GPU nodes are added / removed (e.g., expanding from 8 nodes to 16 nodes), trigger data collection, perform bandwidth tests around the newly added nodes, collect path delay and bandwidth data (sampling density: 10 times / second, for 10 seconds), and only update the affected areas (e.g., grid points within a 2-meter radius of the newly added nodes). The number of training rounds is 50, and the learning rate is 0.001. Compare the simulated single iteration time with the measured single iteration time, and calculate the relative deviation. If the deviation is >10%, trigger the double-layer strategy network update, reduce the attention weight of the path with large bandwidth density fluctuation (weight decay coefficient 0.1), and increase the scheduling penalty coefficient for operators with high video memory consumption (such as increasing the weight by 0.05). If the node failure event is >3 times / hour, increase the priority of the fault migration rule to the highest level, and dynamically adjust the trigger threshold according to the video memory fluctuation data (such as from 10% to 8%). Monitor hardware changes (such as node addition and reduction and NVLink upgrade) and performance indicators in real time (the average single iteration time of 5 consecutive iterations increases by >15% compared with the historical optimal value), trigger the optimization process, and screen historical strategies with similar hardware configurations to the current one (such as the same number of nodes and the same generation of GPU models) from the candidate library; fine-tune parameters in the neighborhood of historical strategies (number of pipeline stages ±2 and number of data parallelism ±1), execute 50 iterations, and explore the optimal solution; if the optimization range of single iteration time for 10 consecutive iterations is less than 1%, terminate early to save resources.

[0047] It should be noted that real-time data feedback updates NeRF and the two-layer policy network. When hardware changes (such as node expansion), the local topology field is updated through incremental training of the NeRF model (only the affected area is trained), and the policy re-optimization time. When the deviation between the simulation and the measured single iteration time is >10%, the policy network parameters are adjusted (such as the attenuation weight coefficient of 0.1). In the scenario of frequent node failures (failures >3 times / hour), the training interruption time is reduced by 94%.

[0048] This embodiment also provides a computer device, which is suitable for the case of a parallel strategy search method for efficient training of large artificial intelligence models, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the parallel strategy search method for efficient training of large artificial intelligence models proposed in the above embodiment.

[0049] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0050] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the parallel strategy search method for efficient training of large artificial intelligence models proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, disk or optical disk.

[0051] In summary, the present invention: based on the neural radiation field (NeRF) modeling physical-level hardware communication field, the mixed integer programming and deep reinforcement learning are integrated to achieve global-local collaborative optimization of parallel strategies, and the three-dimensional bandwidth density field is constructed through the NeRF model to quantify the cross-rack link attenuation effect. The predicted communication time accuracy is significantly improved. In the GPT-3 level model training, the cluster TPI (single iteration time) is reduced, and the mixed integer programming (MIQP) is used to jointly optimize the pipeline stage division, data parallel sharding and calculation-communication overlap factor to avoid the local optimal trap caused by traditional hierarchical optimization. The actual measurement shows that the video memory fragmentation rate is significantly reduced, and the GPU utilization rate is improved. Based on the Transformer-cellular automaton architecture, combined with the online incremental learning mechanism, the strategy re-optimization after the hardware topology change (such as node expansion and contraction) is completed, the fault recovery time is shortened, and the system throughput fluctuation is reduced compared with the static strategy. This solution breaks through the limitation of the traditional method of only optimizing a single performance indicator, and achieves a multi-objective balance of training efficiency, video memory utilization and system robustness.

[0052] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A parallel strategy search method for efficient training of large artificial intelligence models, characterized by: include, Input the physical topology parameters of the cluster, train the NeRF model with multi-view bandwidth and delay data, and generate a three-dimensional hardware topology field including bandwidth density and communication attenuation coefficient; Parse the computational graph of the large model to be trained, map the operator nodes to the three-dimensional hardware topology field, annotate each operator node with a dynamic communication cost label according to the bandwidth density and communication attenuation coefficient of the operator node, and generate a labeled computational graph; A mixed integer programming model is constructed based on the annotated computational graph to output a set of candidate strategies that satisfy the communication field constraints; The candidate strategy set is input into the two-layer strategy network, and the parallel strategy set including the global scheduling framework and local adaptive rules is output; Load the 3D hardware topology field in the simulator, verify and adjust the parallel strategy set, and output the optimal strategy; Deploy the optimal strategy to the real cluster and collect training data, provide feedback to update the NeRF model, and trigger the mixed integer programming model to re-screen the strategy.

2. The parallel strategy search method for efficient training of large artificial intelligence models as described in claim 1, characterized in that: The physical topology parameters include rack layout, connection relationship between nodes and physical distance.

3. The parallel strategy search method for efficient training of large artificial intelligence models as described in claim 1, characterized in that: The operator node includes the operator's computational load, video memory requirements, and dependencies.

4. The parallel strategy search method for efficient training of large artificial intelligence models as claimed in claim 1, characterized in that: The mixed integer programming model based on the labeled computational graph is constructed by integrating the computational time, communication time and constraints through the objective function, and enumerating the pipeline stage division and data parallel sharding strategy through the quantum annealing algorithm.

5. The parallel strategy search method for efficient training of large artificial intelligence models as claimed in claim 1, characterized in that: The candidate strategy set is input into the two-layer strategy network, and the parallel strategy set including the global scheduling framework and the local adaptive rules is output. The preliminary scheduling strategy is generated by global attention modeling of the spatiotemporal features of the annotated computational graph through the Transformer encoder, the cellular automaton network performs local rule adjustment based on the states of adjacent nodes, and the Transformer encoder and the cellular automaton network are jointly optimized through a differentiable interface.

6. The parallel strategy search method for efficient training of large artificial intelligence models as claimed in claim 1, characterized in that: The verification and adjustment of the parallel strategy set refers to using the low-fidelity mode to screen potential high-quality strategies, and the high-fidelity mode to simulate cross-node communication paths through ray tracing, and dynamically adjust task allocation in combination with the cellular automation rule base.

7. The parallel strategy search method for efficient training of large artificial intelligence models as claimed in claim 1, characterized in that: The training data includes communication efficiency, video memory fluctuation data and hardware abnormal events.

8. The parallel strategy search method for efficient training of large artificial intelligence models as claimed in claim 7, characterized in that: The training data is divided into two feedback paths. The first path inputs the NeRF model to incrementally update the three-dimensional hardware topology field parameters. The second path inputs the two-layer policy network. By comparing the deviation between the simulated prediction value and the measured value, the dependence weight of the Transformer attention mechanism on the hardware field characteristics is adjusted, and the local rule base of the cellular automaton is evolved.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the parallel strategy search method for efficient training of large artificial intelligence models described in any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the parallel strategy search method for efficient training of large artificial intelligence models described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Adaptive distributed parallel training method for neural network based on reinforcement learning

    CN113128702A

  • Parallel strategy search method for efficient training of artificial intelligence large model

    CN118966321A

  • Self-adaptive re-calculation and load division method, device and equipment based on neural network, and computer readable medium

    CN119271397A

  • Reliability analysis method and system for power system

    CN119273170A

  • Method and apparatus for performing distributed training on deep learning model, device and storage medium

    US20220374713A1

Cited By

  • Distributed scheduling training method for AI model training

    CN120216205A

  • Distributed training method, device and equipment for large-scale model

    CN120409554A

  • A distributed training method, device and equipment for large-scale models

    CN120409554B

  • Large model training method and system based on green distributed computing power center

    CN120450088A

  • A large model training method and system based on green distributed computing center

    CN120450088B