Intelligent computing cluster parallel training performance optimization method for large model training
By building a hardware topology map and adaptive routing algorithm to optimize communication paths, and combining it with a fault prediction model to achieve rapid recovery, the problems of unbalanced resource utilization and communication bottlenecks in intelligent computing clusters during large model training are solved, thereby improving training performance.
Patent Information
- Application Number
- CN202511032757.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing intelligent computing clusters have problems such as unbalanced GPU/NPU utilization, low computing power utilization, severe communication bottlenecks, and low fault recovery efficiency in large model training.
By collecting hardware characteristic indicators of heterogeneous computing nodes in real time, a hardware topology map is constructed, an adaptive routing algorithm is used to select the optimal communication path, and mixed precision compression technology is combined for data transmission. A fault prediction model is built for rapid recovery, and the communication path and storage mechanism are dynamically adjusted.
It improves the utilization of super node computing power, improves communication efficiency, shortens fault recovery time, and optimizes training performance.
Smart Images

Figure CN120804914A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence computing, in particular to a parallel training performance optimization method for a big model training-oriented intelligent computing cluster. BACKGROUND
[0002] The intelligent computing cluster is a high-performance computing cluster specially designed for artificial intelligence training and reasoning tasks. By integrating large-scale GPU / TPU acceleration cards, high-speed networks, and distributed software stacks, it provides efficient computing power support for deep learning, big model training, and other scenarios. However, the existing intelligent computing cluster has the following shortcomings in terms of training requirements for big model training: The traditional scheduler uses a fixed resource partitioning strategy, resulting in uneven utilization of heterogeneous computing units such as GPUs / NPUs, with some nodes overloaded and others idle. Multiple types of chips such as CPUs, GPUs, and NPUs coexist, and traditional scheduling strategies are difficult to dynamically adapt to hardware topology, resulting in low computing power utilization. For example, the resource void problem in GPU interconnection is prominent in tensor parallel training, and the interconnection efficiency within the super node is insufficient. Communication bottleneck is significant: In the All-to-All communication mode, the data transmission delay between nodes is high, and the GPU idle rate in expert parallel training is more than 30%. The existing network topology lacks dynamic routing optimization, and the cross-node bandwidth utilization is less than 60%. Low fault recovery efficiency: Lack of active health management mechanism, average hardware fault localization time is more than 15 minutes, task restart causes training progress loss of about 20%. To address the above technical deficiencies, the present application proposes a solution. SUMMARY
[0003] The present application aims to: collect hardware feature indicators of heterogeneous computing nodes in the intelligent computing cluster in real time, construct a hardware topology map, and calculate the optimal computing power combination suitable for the hardware topology map; use an adaptive routing algorithm to select the optimal communication path, and combine a mixed precision compression technology to encode and decode the transmission data, achieving efficient transmission of transmission data; construct a fault prediction model based on a deep learning algorithm, input real-time data from hardware monitoring logs into the fault prediction model to obtain prediction results, and store the prediction results based on a pre-set result grading standard, and train checkpoints based on the prediction results to achieve fast recovery, improving super node computing power utilization, and dynamically adjusting the communication path based on network congestion state to improve fault recovery speed.
[0004] To achieve the above purpose, the present application adopts the following technical solution: a parallel training performance optimization method for a big model training-oriented intelligent computing cluster, comprising the following steps: Step one, real-time collection of hardware characteristic indicators of heterogeneous computing nodes in the intelligent computing cluster, and construction of a hardware topology map according to the hardware indicator characteristics, and calculation of the optimal computing power combination suitable for the hardware topology map using a proximal strategy optimization algorithm; Step two, obtaining the network topology structure and real-time flow matrix, selecting the optimal communication path using an adaptive routing algorithm, and encoding and decoding the transmission data using a hybrid precision compression technique to achieve efficient transmission of the transmission data; Step three, obtaining historical data and software exception stack information from the hardware monitoring log, and constructing a fault prediction model based on a deep learning algorithm; Step four, inputting real-time data in the hardware monitoring log into the fault prediction model to obtain a prediction result, classifying and storing the prediction result according to a pre-set result classification standard, and training a checkpoint based on the prediction result to achieve rapid recovery.
[0005] Further, the specific process of constructing the hardware topology map is as follows: S101, real-time acquisition of the power utilization rate, multi-processor activity, and high-bandwidth memory bandwidth of each heterogeneous computing node in the intelligent computing cluster through a lossless network, and filling in the missing data collected using a multiple imputation algorithm; S102, using a hierarchical clustering algorithm to group hardware nodes according to feature similarity, and setting the edge weight based on inter-node communication efficiency; S103, using a proximal strategy optimization algorithm to minimize task completion time as the target, and outputting a computing power allocation scheme; S104, using Gephi to generate a visual hardware topology map, which includes a three-dimensional feature vector of inter-node interconnection bandwidth, delay, and computing power scale.
[0006] Further, the specific process of selecting the optimal communication path is as follows: S201, obtaining the network topology structure and real-time flow matrix, and monitoring the real-time bandwidth utilization rate, delay data, and packet loss rate, specifically using iperf3 to measure the bandwidth, ping to monitor the delay, and tcpdump to count the packet loss rate, and constructing a global network view based on the network topology structure; S202, using a distributed adaptive routing algorithm, combining time and distance marking variables, and each node independently calculating a routing table and synchronously updating it through the BGP protocol; S203, preferentially selecting a path with sufficient remaining bandwidth and the lowest delay as the optimal path, and the next one as a backup path; S204, during the data transmission process through the optimal path, real-time monitoring is performed, and when link failure or congestion is detected, the system automatically switches from the optimal path to the backup path.
[0007] Further, the specific method for constructing the fault prediction model is as follows: S301, obtaining historical data and software abnormal stack information in the hardware monitoring log of the computing node in the intelligent computing cluster, the historical data including GPU temperature, video memory error rate, network packet loss rate, constructing a data set based on time sequence characteristics, and dividing the data set into a training set, a validation set and a test set according to a ratio of 7:2:1; S302, constructing a fault prediction model using a bidirectional LSTM network structure, setting the input layer dimension as the number of time sequence characteristics, and setting the output layer dimension as the fault probability; S303, using a cross-entropy loss function and an Adam optimizer to iteratively train the model until the F1-score of the validation set reaches a preset standard value, obtaining a trained fault prediction model.
[0008] Further, the specific process of grading storage of the prediction result is as follows: Obtain the prediction result output by the fault prediction model, that is, the fault probability P, and obtain the probability judgment range (Pmin, Pmax), and make the following judgments: If the fault probability is less than or equal to Pmin, it is classified as a first-level fault prediction result, and the fault point is located in the cold storage layer; If the fault probability is less than Pmin and greater than Pmax, it is classified as a second-level fault prediction result, and the fault point is located in the warm storage layer; If the fault probability is greater than or equal to Pmax, it is classified as a third-level fault prediction result, and the fault point is located in the hot storage layer, an emergency warning signal is generated, and at the same time the fault data is retained to start the warning defense measures.
[0009] Further, the specific process of grading storage of the prediction result is as follows: S401, constructing a data set containing historical scheduling schemes and corresponding task completion times; S402, defining the state space as a feature vector of the hardware topology map, the action space as a selectable computing power allocation scheme, and the reward function as the negative task completion time; S403, based on the prediction result output by the fault prediction model, dynamically allocating the storage level, and setting the checkpoint in turn S404, constructing a fault troubleshooting model based on XGBoost, the input features including training loss, accuracy and learning rate, using experience replay and target network technology to iteratively update the policy network parameters until the fault troubleshooting model converges; S405, after the checkpoint is generated, inputting the fault troubleshooting model in turn to calculate the score, and the checkpoint with a score reaching a preset score threshold is the fault point.
[0010] Further, the specific process of determining the checkpoint is as follows: by setting the checkpoint hierarchical storage mechanism, the recent several checkpoints are stored in the hot data layer using the NVMe solid state disk, and the historical version is reserved in the cold data layer using the distributed object storage, and the hot data layer is preferentially loaded during fault recovery.
[0011] To sum up, due to the adoption of the technical scheme, the present application has the following advantages: The large model training oriented intelligent computing cluster parallel training performance optimization method realizes efficient transmission of transmission data by collecting hardware feature indexes of heterogeneous computing nodes in the intelligent computing cluster in real time, constructing a hardware topology map, calculating an optimal computing power combination suitable for the hardware topology map, selecting an optimal communication path using an adaptive routing algorithm, and encoding and decoding transmission data in combination with a mixed precision compression technology. Based on a deep learning algorithm, a fault prediction model is constructed, real-time data in the hardware monitoring log is input into the fault prediction model to obtain a prediction result, the prediction result is stored according to a preset result grading standard, and a checkpoint is trained based on the prediction result to realize rapid recovery, improve the super node computing power utilization rate, and dynamically adjust the communication path according to the network congestion state to improve the fault recovery speed. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 The overall method flowchart of the present application is shown. DETAILED DESCRIPTION
[0013] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application. EMBODIMENT
[0014] As Figure 1 shown, the large model training oriented intelligent computing cluster parallel training performance optimization method comprises the following steps: Step one, real-time collection of hardware feature indexes of heterogeneous computing nodes in the intelligent computing cluster, and construction of a hardware topology map according to the hardware index features, calculation of an optimal computing power combination suitable for the hardware topology map using a near-end strategy optimization algorithm; The specific process of constructing the hardware topology map is as follows: S101, the power utilization rate, multi-processor activity and high-bandwidth memory bandwidth of each heterogeneous computing node in the intelligent computing cluster are obtained in real time through a lossless network, and the missing data collected is filled in using a multiple imputation algorithm; S102, using hierarchical clustering algorithm, grouping hardware nodes according to feature similarity, and setting the weight of edge based on the communication efficiency between nodes; S103, using the proximal policy optimization algorithm, minimizing the task completion time as the goal, outputting the computing power allocation scheme; S104, using Gephi to generate a visual hardware topology map, which contains the three-dimensional feature vector of the interconnection bandwidth, delay and computing power scale between nodes.
[0015] Step two, obtain the network topology structure and real-time traffic matrix, use adaptive routing algorithm to select the optimal communication path, and combine the mixed precision compression technology to encode and decode the transmission data, realize the efficient transmission of transmission data; The specific process of selecting the optimal communication path is as follows: S201, obtain the network topology structure and real-time traffic matrix, and monitor the bandwidth utilization rate, delay data and packet loss rate in real time, specifically use iperf3 to measure bandwidth, ping to monitor delay, and tcpdump to count packet loss rate, and build a global network view based on the network topology structure; S202, using distributed adaptive routing algorithm, combining time label variable and distance label variable, each node independently calculates routing table, and updates through BGP protocol synchronization; S203, preferentially select the path with sufficient remaining bandwidth and the lowest delay as the optimal path, and the second as the backup path; S204, during the data transmission process through the optimal path, real-time monitoring is carried out, and when link failure and congestion are monitored, the optimal path is automatically switched to the backup path.
[0016] Step three, obtain the historical data in the hardware monitoring log and the software abnormal stack information, and build a fault prediction model based on deep learning algorithm; Step four, input the real-time data in the hardware monitoring log into the fault prediction model to obtain the prediction result, and store the prediction result according to the preset result grading standard, and train the checkpoint based on the prediction result to realize fast recovery.
[0017] The specific method of building the fault prediction model is as follows: S301, obtain the historical data and software abnormal stack information in the hardware monitoring log of the computing node in the intelligent algorithm cluster, the historical data includes GPU temperature, video memory error rate, network packet loss rate, build a data set based on time series characteristics, and divide the data set into training set, validation set and test set according to the ratio of 7:2:1; S302, using bidirectional LSTM network structure to build fault prediction model, setting input layer dimension as time series feature quantity, and output layer dimension as fault probability; S303, adopt cross-entropy loss function and Adam optimizer, iteratively train the model until the validation set F1-score reaches the preset standard value, and get the trained fault prediction model.
[0018] Further, the specific process of hierarchical storage of the prediction result is as follows: Get the prediction result output by the fault prediction model, i.e. the fault probability P, and get the probability judgment range (Pmin, Pmax), and make the following judgments: If the fault probability is less than or equal to Pmin, it is classified as a first-level fault prediction result, and the fault point is located in the cold storage layer; If the fault probability is less than Pmin and greater than Pmax, it is classified as a second-level fault prediction result, and the fault point is located in the warm storage layer; If the fault probability is greater than or equal to Pmax, it is classified as a third-level fault prediction result, and the fault point is located in the hot storage layer, an emergency warning signal is generated and the fault data is retained to start the warning defense measures.
[0019] The specific process of hierarchical storage of the training checkpoint based on the prediction result is as follows: S401, construct a data set containing historical scheduling schemes and corresponding task completion times; S402, define the state space as the feature vector of the hardware topology map, the action space as the optional computing power allocation scheme, and the reward function as the negative task completion time; S403, based on the prediction result output by the fault prediction model, dynamically allocate the storage hierarchy, and set the checkpoints in turn S404, build a fault troubleshooting model based on XGBoost, input features include training loss, accuracy and learning rate, use experience replay and target network technology, iteratively update policy network parameters until the fault troubleshooting model converges; S405, after the checkpoint is generated, input the fault troubleshooting model in turn to calculate the score, and the checkpoint with a score reaching the preset score threshold is the fault point.
[0020] The specific process of determining the checkpoint is as follows: by setting the checkpoint hierarchical storage mechanism, the recent several checkpoints are stored in the hot data layer using NVMe solid state disk, and the historical version is preserved in the cold data layer using distributed object storage, and the hot data layer is preferentially loaded during fault recovery.
[0021] The application realizes efficient transmission of transmission data by collecting hardware feature indexes of heterogeneous computing nodes in the intelligent computing cluster in real time, constructing a hardware topology map, calculating the optimal computing power combination suitable for the hardware topology map, selecting the optimal communication path by using an adaptive routing algorithm, and encoding and decoding transmission data by using a mixed precision compression technology, constructing a fault prediction model based on a deep learning algorithm, inputting real-time data in the hardware monitoring log into the fault prediction model to obtain a prediction result, classifying and storing the prediction result according to a preset result classification standard, training a checkpoint based on the prediction result to realize rapid recovery, improving the super node computing power utilization rate, and dynamically adjusting the communication path according to the network congestion state to improve the fault recovery speed.
[0022] The size of the threshold is set for the purpose of comparison, and the size of the threshold depends on how much sample data and the base number set by the person skilled in the art for each group of sample data; as long as the proportional relationship between the parameters and the quantized values is not affected.
[0023] The above is only a preferred specific embodiment of the application, but the protection scope of the application is not limited thereto, any person skilled in the art can make equivalent replacement or change according to the technical solution and the inventive concept of the application within the technical range disclosed by the application, which should be covered within the protection scope of the application.
Claims
1. A method for optimizing the performance of parallel training of intelligent computing clusters for large model training, characterized by: The following steps are involved: Step 1: Collect hardware characteristic indicators of heterogeneous computing nodes in the intelligent computing cluster in real time, build a hardware topology map based on the hardware indicator characteristics, and use the proximal strategy optimization algorithm to calculate the optimal computing power combination suitable for the hardware topology map; Step 2: Obtain the network topology and real-time traffic matrix, use the adaptive routing algorithm to select the optimal communication path, and combine the mixed precision compression technology to encode and decode the transmission data to achieve efficient transmission of the transmission data; Step 3: Obtain historical data and software exception stack information from hardware monitoring logs, and build a fault prediction model based on deep learning algorithms; Step 4: Input the real-time data in the hardware monitoring log into the fault prediction model to obtain the prediction results, and store the prediction results in a graded manner according to the preset result classification standards, and train checkpoints based on the prediction results to achieve rapid recovery.
2. The method for optimizing the performance of parallel training of intelligent computing clusters for large model training according to claim 1 is characterized in that: The specific process of building a hardware topology map is as follows: S101. Obtain the power utilization, multi-processor activity, and high-bandwidth memory bandwidth of each heterogeneous computing node in the intelligent computing cluster in real time through a lossless network. Simultaneously, use a multiple interpolation algorithm to fill in the missing data collected. S102, using a hierarchical clustering algorithm to group hardware nodes according to feature similarity, and setting edge weights based on inter-node communication efficiency; S103: Using a proximal strategy optimization algorithm, with the goal of minimizing task completion time, output a computing power allocation plan; S104. Generate a visualized hardware topology map using Gephi, where the hardware topology map includes three-dimensional feature vectors of interconnection bandwidth, latency, and computing power scale between nodes.
3. The method for optimizing the performance of parallel training of intelligent computing clusters for large model training according to claim 1 is characterized in that: The specific process of selecting the optimal communication path is as follows: S201. Obtain the network topology and real-time traffic matrix, and monitor bandwidth utilization, latency data, and packet loss rate in real time. Specifically, use iperf3 to measure bandwidth, ping to monitor latency, and tcpdump to calculate packet loss rate, and build a global network view based on the network topology. S202, using a distributed adaptive routing algorithm, combined with time tag variables and distance tag variables, each node independently calculates the routing table and updates it synchronously through the BGP protocol; S203: Prioritize the path with sufficient remaining bandwidth and the lowest delay as the optimal path, and then select the backup path; S204 , during data transmission via the optimal path, real-time monitoring is performed, and when link failure or congestion is detected, the optimal path is automatically switched to a backup path.
4. The method for optimizing the performance of parallel training of intelligent computing clusters for large model training according to claim 1 is characterized in that: The specific method of building a fault prediction model is as follows: S301. Obtain historical data and software exception stack information from the hardware monitoring logs of computing nodes in the intelligent computing cluster, wherein the historical data includes GPU temperature, video memory error rate, and network packet loss rate. A data set is constructed based on time series features, and the data set is divided into a training set, a validation set, and a test set in a ratio of 7:2:
1. S302. Use a bidirectional LSTM network structure to build a fault prediction model, set the input layer dimension to the number of time series features, and the output layer dimension to the fault probability; S303: Using the cross entropy loss function and the Adam optimizer, the model is iteratively trained until the F1-score of the validation set reaches a preset standard value, thereby obtaining a trained fault prediction model.
5. The method for optimizing the performance of parallel training of intelligent computing clusters for large model training according to claim 1 is characterized in that: The specific process of hierarchical storage of prediction results is as follows: Obtain the prediction result output by the fault prediction model, that is, the fault probability P, and obtain the probability judgment range (Pmin, Pmax) at the same time, and make the following judgments: If the failure probability is less than or equal to Pmin, it is classified as a level 1 failure prediction result, and the failure point is located in the cold storage layer; If the failure probability is less than Pmin and greater than Pmax, it is classified as a secondary failure prediction result, and the failure point is located in the warm storage layer; If the failure probability is greater than or equal to Pmax, it is classified into three levels of fault prediction results, and the fault point is located in the thermal storage layer, generating an emergency warning signal and retaining the fault data to start the warning defense measures.
6. The method for optimizing the performance of parallel training of intelligent computing clusters for large model training according to claim 1, characterized in that: The specific process of storing training checkpoints in a hierarchical manner based on prediction results is as follows: S401, constructing a data set including historical scheduling plans and corresponding task completion times; S402: Define the state space as the eigenvectors of the hardware topology map, the action space as the optional computing power allocation scheme, and the reward function as the negative task completion time; S403: Based on the prediction results output by the fault prediction model, dynamically allocate storage tiers and set checkpoints in sequence. S404. Build a troubleshooting model based on XGBoost, input features including training loss, accuracy, and learning rate, use experience replay and target network technology, and iteratively update the policy network parameters until the troubleshooting model converges. S405: After the checkpoints are generated, they are sequentially input into the fault troubleshooting model to calculate scores. The checkpoint defects that reach the preset score threshold are considered fault points.
7. The method for optimizing the performance of parallel training of intelligent computing clusters for large model training according to claim 6 is characterized in that: The specific process of determining checkpoints is as follows: By setting up a checkpoint tiered storage mechanism, NVMe solid-state drives are used in the hot data layer to store the most recent checkpoints, and distributed object storage is used in the cold data layer to retain historical versions. The hot data layer is loaded first during fault recovery.
Citation Information
Cited By
Ensemble communication method of core particle equipment cluster suitable for unified bus interconnection
CN121771104A
A collection communication method suitable for a core grain device cluster of a unified bus interconnection
CN121771104B