Online automatic tuning and compiling system and method for large model training
By designing an online automatic tuning and compilation system for large model training, the performance indicator data is collected in real time and hyperparameter configuration is dynamically adjusted, the problems of inefficiency and inequality of traditional tuning methods are solved, and efficient hyperparameter tuning and computing performance improvement are achieved.
Patent Information
- Application Number
- CN202510571510.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
In large model training, traditional hyperparameter tuning methods are inefficient and not universal, making them difficult to adapt to the dynamic requirements in multi-operator fusion and distributed computing environments.
An online automatic tuning and compilation system for large model training is designed, including automatic tuning client, server side, configuration manager, Triton operator registrant, communication manager, automatic tuning scheduler and data collector. By collecting performance indicator data in real time, dynamically adjusting hyperparameter configuration, and realizing automatic tuning.
It significantly improves the computing performance of the fusion operator, can handle large-scale hyperparameter space, adapt to heterogeneous hardware environments, maximizes the utilization of GPU resources, reduces manual intervention, shortens the tuning cycle, and improves the efficiency and accuracy of large-scale model training.
Smart Images

Figure CN120087455A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computational framework programming and compiler for large model training in artificial intelligence, and particularly relates to an online automatic tuning compilation system and method for large model training. Background Art
[0002] In large model training, as the model scale increases, the number of operators and hyperparameters involved in the training process increases sharply. Traditional hyperparameter tuning methods often require a large amount of manual intervention and repeated experiments, with low efficiency and lack of universality. Existing automatic tuning methods mostly focus on the tuning of single operators and are difficult to adapt to the dynamic requirements in the multi-operator fusion and distributed computing environment. Summary of the Invention
[0003] The purpose of the present invention is to provide an online automatic tuning compilation system and method for large model training in view of the deficiencies of the prior art.
[0004] The purpose of the present invention is achieved through the following technical solutions: An online automatic tuning compilation system for large model training, including an automatic tuning client, an automatic tuning server, a configuration manager, a Triton operator register, a communication manager, an automatic tuning scheduler, and a data collector; The data collector is used to collect performance metric data during the large model training process as the optimization target; The automatic tuning client is responsible for, after the large model training is started, sending the performance metric data collected by the data collector and calling the communication manager to send it to the automatic tuning server; The automatic tuning server is used to, after receiving the automatic tuning request from the automatic tuning client, initialize the hyperparameter space according to different tuning strategies, and adjust the hyperparameter configuration required for the next round of iterative training as the updated hyperparameter configuration and call the communication manager to send it to the automatic tuning client according to the performance metric data fed back by the automatic tuning client; The configuration manager is used to provide a switch for starting online automatic tuning to the user, and configure the updated hyperparameter configuration received by the automatic tuning client into the Triton fusion operator in the large model framework; The Triton operator register is used to register the Triton fusion operator into the large model training framework and receive the updated hyperparameter configuration; The communication manager is responsible for implementing data transmission between the automatic tuning client and the automatic tuning server using a communication protocol; The automatic tuning scheduler is used to pass the hyperparameter configuration to multiple GPU servers participating in the large model training in the cluster.
[0005] Further, the data collector includes a performance metric collection module, a data preprocessing module, and a metric aggregation module; The performance metric collection module is used to collect the performance metric data of each training node in the local training process in real time. The performance metric data includes, but is not limited to, computing throughput, memory usage, GPU utilization, and communication overhead; The data preprocessing module is used to clean and standardize the collected performance metric data to obtain the processed performance metric data; The metric aggregation module is used to summarize and statistically analyze the processed performance metric data of different training nodes and send it to the automatic tuning client.
[0006] Further, the automatic tuning client includes a data collection interface module, a metric preprocessing module, a configuration update module, and a communication interface module; The data collection interface module is used to receive the performance metric data from the data collector; The metric preprocessing module is used to detect outliers in the performance metric data from the data collector, remove the noise data, then standardize the data in different dimensions to obtain the standardized performance metric data; finally, perform time series aggregation on the standardized performance metric data to obtain the time series aggregated performance metric data; The configuration update module is responsible for receiving the updated hyperparameter configuration sent by the automatic tuning server side and applying the updated hyperparameter configuration to the local training process; The communication interface module is used to send the time series aggregated performance metric data to the automatic tuning server side through the communication manager.
[0007] Further, the automatic tuning server side includes a hyperparameter space initialization module, a metric analysis module, an optimization strategy selection module, and a configuration optimization module; The hyperparameter space initialization module is used to initialize the hyperparameter space according to different tuning strategies after receiving the automatic tuning request from the automatic tuning client; The metric analysis module is used to analyze the performance metric data from the automatic tuning client and construct a performance evaluation model; the performance evaluation model comprehensively considers computing efficiency, memory efficiency, and communication efficiency and generates a comprehensive performance score for the current configuration; The optimization strategy selection module is used to dynamically select a suitable optimization strategy according to the characteristics and optimization objectives of the training scenario during the local training process of the large model and the generated comprehensive performance score; the optimization strategies include Bayesian optimization strategy, reinforcement learning strategy, or genetic algorithm strategy; The configuration optimization module adjusts the hyperparameter configuration required for the next round of iterative training through a selected suitable optimization strategy as the updated hyperparameter configuration and calls the communication manager to send it to the auto-tuning client.
[0008] Furthermore, the configuration manager includes a configuration storage module, a configuration update module, a version control layer, and a monitoring layer; The configuration storage module uses a distributed database, which not only stores the optimal solution exceeding the configuration in the previous round of iterative training, but also stores the correspondence between historical configuration records and performance data; The configuration update module, after the user turns on the switch to start online auto-tuning, configures the updated hyperparameter configuration received by the auto-tuning client into the Triton fusion operator in the large model framework; The version control layer manages the version evolution process of the configuration, implements a fast rollback mechanism for the configuration, and can restore to a stable version in a timely manner when performance anomalies are found; The monitoring layer continuously monitors the effective status and performance impact of the configuration, evaluates the effect of configuration updates in real time by collecting feedback information from each node, and can notify the system administrator in a timely manner when configuration anomalies are found.
[0009] Furthermore, the communication manager uses the HTTP or gRPC protocol to implement data transmission.
[0010] Furthermore, the Triton operator register includes a compilation module, a registration module, an optimization module, and a verification module; The compilation module compiles the Triton fusion operator into a binary file; The registration module registers the compiled Triton fusion operator into the large model training framework; The optimization module optimizes the compiled Triton fusion operator; The verification module verifies the correctness of the compiled Triton fusion operator registered in the large model training framework.
[0011] Furthermore, the auto-tuning scheduler includes a load balancing module, a configuration synchronization module, and a status monitoring module; The load balancing module dynamically allocates tuning tasks according to the load status of each GPU node; The configuration synchronization module ensures the consistency of hyperparameter configurations of all nodes in the cluster; The status monitoring module monitors the running status and tuning progress of each node in real time.
[0012] Further, the automatic tuning client further includes an anomaly detection module and a logging module; The anomaly detection module is used to monitor the anomaly situation during the large model training process and notify the user when an anomaly is detected; The logging module is used to record the key operations and performance changes during the tuning process for user auditing and system maintenance.
[0013] The present invention also provides an online automatic tuning compilation method for large model training. This method is implemented based on the above-mentioned online automatic tuning compilation system for large model training. Specifically, this method includes: Collect performance metric data during the large model training process as the optimization target through the data collector; Through the automatic tuning client, after the large model training starts, it is responsible for sending the performance metric data collected by the data collector and calling the communication manager to send it to the automatic tuning server side; Through the automatic tuning server side, after receiving the automatic tuning request from the automatic tuning client, initialize the hyperparameter space according to different tuning strategies, and adjust the hyperparameter configuration required for the next round of iterative training as the updated hyperparameter configuration and call the communication manager to send it to the automatic tuning client according to the performance metric data fed back by the automatic tuning client; Through the configuration manager, provide a switch to start online automatic tuning to the user, and configure the updated hyperparameter configuration received by the automatic tuning client into the Triton fusion operator in the large model framework; Through the Triton operator registrar, register the Triton fusion operator into the large model training framework and receive the updated hyperparameter configuration; Through the communication manager, be responsible for implementing data transmission between the automatic tuning client and the automatic tuning server side using the communication protocol; Through the automatic tuning scheduler, pass the hyperparameter configuration to multiple GPU servers participating in the large model training in the cluster.
[0014] The beneficial effects of the present invention are: Through the online automatic tuning compilation system provided by the present invention, it is possible to automatically search and adjust the hyperparameter configuration during the large model training process, significantly improving the computing performance of the fusion operator. Compared with traditional methods, the present invention can not only handle large-scale hyperparameter spaces but also adapt to heterogeneous hardware environments, maximizing the utilization of GPU resources. In addition, the system can also reduce manual intervention through the automated tuning process, shorten the tuning cycle, thereby improving the efficiency and accuracy of large model training. Description of the Drawings
[0015] Figure 1It is a structural diagram of an online automatic tuning compilation system for large model training; Figure 2 It is a structural diagram of a data collector; Figure 3 It is a structural diagram of an automatic tuning client; Figure 4 It is a structural diagram of an automatic tuning server; Figure 5 It is a structural diagram of a configuration manager; Figure 6 It is a structural diagram of a Triton operator register; Figure 7 It is a structural diagram of an automatic tuning scheduler. Detailed implementation
[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts are within the protection scope of the present invention.
[0017] Technical term explanations: Triton, a language proposed by OpenAI for writing efficient custom deep learning primitives.
[0018] HTTP, Hypertext Transfer Protocol, is an Internet protocol for distributed, collaborative, hypermedia information systems. It is the basis for data communication on the World Wide Web, allowing hypertext data to be transferred over the Internet from one system to another. HTTP defines the format of requests and responses between clients and servers, as well as the request methods that a client may send, such as GET, POST, PUT, DELETE, etc.
[0019] gRPC, gRPC Remote Procedure Call, is a high-performance, open-source, and general-purpose RPC (Remote Procedure Call) framework, mainly developed by Google. It allows transparent communication between client and server applications and supports multiple programming languages.
[0020] I. Overall architecture of an online automatic tuning compilation system for large model training As Figure 1As shown in the figure, an online automatic tuning compilation system for large model training provided by the present invention adopts a distributed architecture design, mainly including seven core functional modules: an automatic tuning client, an automatic tuning server, a configuration manager, a Triton operator register, a communication manager, an automatic tuning scheduler, and a data collector. Each module communicates through a standardized interface to form a complete closed-loop optimization system.
[0021] The data collector is used to collect performance metric data during the large model training as the optimization target.
[0022] The automatic tuning client is deployed on each training node. After the large model training starts, it is responsible for collecting the performance metric data collected by the data collector during the large model training and calling the communication manager to send it to the automatic tuning server.
[0023] The automatic tuning server, first after receiving the automatic tuning request from the automatic tuning client, initializes the hyperparameter space according to different tuning strategies, and based on the performance metric data fed back by the automatic tuning client, adjusts the hyperparameter configuration required for the next round of iterative training as the updated hyperparameter configuration and calls the communication manager to send it to the automatic tuning client.
[0024] The configuration manager saves the optimal solution exceeding configuration in the previous round of iterative training and provides a switch for starting online automatic tuning to the user; after the user turns on the switch for starting online automatic tuning, it configures the updated hyperparameter configuration received by the automatic tuning client into the Triton fusion operator in the large model framework.
[0025] The Triton operator register registers the Triton fusion operator into the large model training framework and receives the updated hyperparameter configuration.
[0026] The communication manager is responsible for implementing data transmission between the automatic tuning client and the automatic tuning server using a reliable communication protocol; the communication manager uses the HTTP or gRPC protocol to implement data transmission.
[0027] The automatic tuning scheduler passes the hyperparameter configuration to multiple GPU servers participating in the large model training in the cluster.
[0028] II. Detailed Design of Core Modules 2.1 Data Collector As Figure 2 shown, the data collector includes a performance metric collection module, a data preprocessing module, and a metric aggregation module.
[0029] The performance metric collection module is used to collect the performance metric data of each training node in the local training process in real time. The performance metric data includes but is not limited to computing throughput, memory usage, GPU utilization, and communication overhead.
[0030] The data preprocessing module is used to clean and standardize the collected performance metric data to obtain the processed performance metric data.
[0031] The metric aggregation module is used to summarize and statistically analyze the processed performance metric data of different training nodes and send them to the auto-tuning client.
[0032] 2.2 Auto-Tuning Client As Figure 3 shown, the auto-tuning client includes a data collection interface module, a metric preprocessing module, a configuration update module, and a communication interface module.
[0033] The data collection interface module is used to receive the performance metric data from the data collector; through the standardized interface definition, the compatibility with different training frameworks is ensured.
[0034] The metric preprocessing module is used to detect outliers in the performance metric data from the data collector, remove the noise data, then standardize the data in different dimensions to obtain the standardized performance metric data, making the standardized performance metric data comparable; finally, perform time-series aggregation on the standardized performance metric data to obtain the time-series aggregated performance metric data, and the time-series aggregated performance metric data is used as a statistical metric reflecting the performance trend.
[0035] The configuration update module is responsible for receiving the updated hyperparameter configuration sent by the auto-tuning server side and applying the updated hyperparameter configuration to the local training process; the configuration update module realizes the atomic update of the configuration to ensure the consistency during the configuration update process. At the same time, it also maintains a local configuration cache, which can quickly roll back to the previous stable configuration in case of network anomalies.
[0036] The communication interface module is used to send the time-series aggregated performance metric data to the auto-tuning server side through the communication manager. The communication interface module realizes functions such as data compression, encrypted transmission, and resume from breakpoint, ensuring the reliability of data transmission in a complex network environment.
[0037] The auto-tuning client also includes an anomaly detection module and a log recording module. The anomaly detection module is used to monitor the anomalies during the large model training process and notify the user when an anomaly is detected. The log recording module is used to record the key operations and performance changes during the tuning process for user auditing and system maintenance.
[0038] The automatic tuning client further includes an anomaly detection module and a logging module. The anomaly detection module is used to monitor anomalies during the large model training process and notify the user when an anomaly is detected. The logging module is used to record key operations and performance changes during the tuning process for user auditing and system maintenance.
[0039] 2.3 Automatic Tuning Server As Figure 4 shown, the automatic tuning server is the core decision-making module of an online automatic tuning compilation system for large model training, including four main components: a hyperparameter space initialization module, a metric analysis module, an optimization strategy selection module, and a configuration optimization module.
[0040] The hyperparameter space initialization module is used to initialize the hyperparameter space according to different tuning strategies after receiving a request for automatic tuning from the automatic tuning client. The metric analysis module is used to analyze performance metric data from the automatic tuning client and construct a performance evaluation model; the performance evaluation model comprehensively considers computational efficiency, memory efficiency, and communication efficiency to generate a comprehensive performance score for the current configuration. The performance analyzer also implements a trend analysis function to predict the performance change trend under the current configuration through the comprehensive performance score, providing a basis for optimization decisions.
[0041] The optimization strategy selection module is used to dynamically select a suitable optimization strategy according to the characteristics and optimization objectives of the training scenario during the local training of the large model and the generated comprehensive performance score; the optimization strategies include Bayesian optimization strategy, reinforcement learning strategy, or genetic algorithm strategy; the optimization strategy selection module also maintains a strategy evaluation mechanism to evaluate and adjust the effects of different optimization strategies through historical optimization effects.
[0042] The configuration optimization module adjusts the hyperparameter configuration required for the next round of iterative training as the updated hyperparameter configuration by the selected suitable optimization strategy and calls the communication manager to send it to the automatic tuning client; during the adjustment process, the configuration optimization module first verifies whether the new hyperparameter configuration meets the hardware constraints and safety boundaries, and then encodes the hyperparameter configuration that meets the hardware constraints and safety boundaries to generate a standardized configuration description as the updated hyperparameter configuration. The configuration optimization module also maintains a configuration version control system to support configuration rollback and tracking.
[0043] 2.4 Configuration Manager As Figure 5 shown, the configuration manager includes a configuration storage module, a configuration update module, a version control layer, and a monitoring layer.
[0044] The configuration storage module is used to adopt a distributed database, which stores not only the optimal solution exceeding the configuration in the previous round of iterative training, but also the corresponding relationship between historical configuration records and performance data.
[0045] The configuration update module is used to configure the updated hyperparameter configuration received by the auto-tuning client into the Triton fusion operator in the large model framework after the user turns on the switch to start online auto-tuning.
[0046] The version control layer is used to manage the version evolution process of the configuration, implement a fast configuration rollback mechanism, and be able to quickly restore to a stable version when performance anomalies are found.
[0047] The monitoring layer is used to continuously monitor the effective status of the configuration and its performance impact. By collecting feedback information from each node, it can evaluate the effect of configuration updates in real time and notify the system administrator in a timely manner when configuration anomalies are found.
[0048] 2.5 Triton Operator Register As Figure 6 shown, the Triton operator register includes a compilation module, a registration module, an optimization module, and a verification module.
[0049] The compilation module is used to compile the Triton fusion operator into a binary file. The compilation process is divided into multiple stages: first, syntax analysis and semantic checking are performed, then intermediate code generation and optimization are carried out, and finally machine code for the target platform is generated. The compilation module supports multiple backends and can generate optimized code for different hardware platforms.
[0050] The registration module is used to register the compiled Triton fusion operator into the large model training framework; the registration process includes steps such as metadata registration, symbol table update, and memory allocation. The registration module implements a hot-plug mechanism and supports dynamic update of operator implementation during training.
[0051] The optimization module is used to optimize the compiled Triton fusion operator. The optimization module implements multi-level code optimization. At the IR level, basic optimizations such as constant propagation and dead code elimination are performed. At the platform-related level, memory access optimization, instruction scheduling optimization, etc. are implemented. The optimizer also includes an adaptive optimization mechanism that can dynamically adjust the optimization strategy according to runtime information.
[0052] The verification module is used to verify the correctness of the compiled Triton fusion operator registered in the large model training framework. The verification process includes interface consistency check, numerical accuracy verification, performance benchmark testing, etc. The verifier also maintains a test case library for regression testing and performance comparison.
[0053] 2.6 Automatic Tuning Scheduler As shown Figure 7 The automatic tuning scheduler further includes a load balancing module, a configuration synchronization module, and a status monitoring module. The load balancing module is used to dynamically allocate tuning tasks according to the load conditions of each GPU node. The configuration synchronization module is used to ensure the consistency of hyperparameter configurations of all nodes in the cluster. The status monitoring module is used to monitor the running status and tuning progress of each node in real time.
[0054] III. System Working Process The working process of an online automatic tuning compilation system for large model training can be divided into the following main stages: Initialization stage: When the online automatic tuning compilation system for large model training starts, each functional module completes the initialization configuration; establishes a communication connection between the automatic tuning client and the automatic tuning server; loads the initial hyperparameter configuration and optimization strategy; initializes the performance monitoring and data acquisition module.
[0055] Data acquisition stage: The data collector continuously monitors and acquires the performance metrics during the local training process; cleans and standardizes the acquired raw performance metric data; packages and sends the processed performance metric data to the automatic tuning server; maintains the local performance metric cache.
[0056] Optimization decision stage: The automatic tuning server receives and aggregates the performance metric data from each node; the performance analyzer evaluates the optimization effect of the current configuration; the optimization strategy engine selects the optimization direction according to the analysis result; generates new hyperparameter configuration suggestions.
[0057] Configuration update stage: The configuration manager verifies the validity of the new configuration; distributes the configuration change notifications to each node; each node completes the configuration update and returns the confirmation information; monitors the system status after the configuration update.
[0058] Feedback adjustment stage: Collects the performance feedback after the configuration update; evaluates the optimization effect, updates the optimization strategy; records the optimization history, updates the experience database; adjusts the direction and intensity of the next round of optimization.
[0059] IV. Key Technology Implementation 4.1 Adaptive Optimization Strategy The online automatic tuning compilation system for large model training implements an adaptive optimization strategy selection mechanism. First, a strategy evaluation model is established, which considers the following factors: historical optimization effect, computing resource consumption, convergence speed, etc. Then, according to the current training stage and performance bottleneck, the most suitable optimization strategy is dynamically selected.
[0060] During the optimization process, the online automatic tuning and compilation system for large model training will continuously evaluate the effectiveness of the strategy and adjust the strategy parameters or switch to other strategies based on the evaluation results. For example, a genetic algorithm with strong exploratory nature may be used in the early stages of training, while a refined Bayesian optimization may be used in the later stages.
[0061] 4.2 Distributed Collaboration Mechanism In order to handle the collaborative optimization problem in a cluster environment, the online automatic tuning and compilation system for large model training implements a distributed collaborative mechanism. This mechanism includes: Global state synchronization: maintain the global optimization state within the cluster; Configuration consistency guarantee: Distributed transactions are used to ensure the consistency of configuration updates; Load balancing: dynamically adjust task allocation based on node performance; Fault-tolerance processing: Implement automatic recovery mechanism when node fails.
[0062] 4.3 Real-time Performance Modeling The online automatic tuning and compilation system for large model training builds a real-time performance evaluation model that can: Rapid response to performance changes: Real-time performance evaluation through streaming data processing; Predict performance trends: predict performance change trends based on historical data; Identify performance bottlenecks: automatically locate key factors affecting performance; Generate optimization suggestions: Output optimization directions based on the performance model.
[0063] On the other hand, the present invention also provides an online automatic tuning and compiling method for large model training, which is implemented by the above-mentioned online automatic tuning and compiling system for large model training, and specifically comprises: Through the data collector, the performance index data in the large model training process is collected as an optimization target; After the large model training is started, the automatic tuning client is responsible for sending the performance indicator data collected by the data collector to the automatic tuning server by calling the communication manager; After receiving the automatic tuning request from the automatic tuning client, the automatic tuning server initializes the hyper-parameter space according to different tuning strategies, and adjusts the hyper-parameter configuration required for the next round of iterative training as the updated hyper-parameter configuration according to the performance indicator data fed back by the automatic tuning client, and calls the communication manager to send it to the automatic tuning client; Through the configuration manager, a switch for starting online automatic tuning is provided to the user, and the updated hyperparameter configuration received by the automatic tuning client is configured into the Triton fusion operator in the large model framework; Through the Triton operator register, register the Triton fusion operator into the large model training framework and receive the updated hyperparameter configuration; Through the communication manager, be responsible for implementing data transmission between the auto-tuning client and the auto-tuning server using the communication protocol; Through the auto-tuning scheduler, pass the hyperparameter configuration to multiple GPU servers participating in large model training in the cluster.
[0064] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.
Claims
1. An online automatic tuning and compilation system for large model training, characterized in that: Includes automatic tuning client, automatic tuning server, configuration manager, Triton operator registrar, communication manager, automatic tuning scheduler and data collector; The data collector is used to collect performance indicator data during the large model training process as an optimization target; The automatic tuning client is used to call the communication manager to send the performance indicator data collected by the data collector to the automatic tuning server after the large model training is started; The automatic tuning server is used to initialize the hyper-parameter space according to different tuning strategies after receiving the automatic tuning request from the automatic tuning client, and adjust the hyper-parameter configuration required for the next round of iterative training as the updated hyper-parameter configuration according to the performance indicator data fed back by the automatic tuning client, and call the communication manager to send it to the automatic tuning client; The configuration manager is used to provide the user with a switch for starting online automatic tuning, and configure the updated hyperparameter configuration received by the automatic tuning client into the Triton fusion operator in the large model framework; The Triton operator register is used to register the Triton fusion operator to the large model training framework and receive updated hyperparameter configuration; The communication manager is responsible for implementing data transmission between the automatic tuning client and the automatic tuning server by using a communication protocol; The automatic tuning scheduler is used to pass the hyperparameter configuration to multiple GPU servers in the cluster that participate in large model training.
2. According to claim 1, an online automatic tuning and compiling system for large model training is characterized in that: The data collector includes a performance index collection module, a data preprocessing module and an index aggregation module; The performance indicator collection module is used to collect performance indicator data of each training node in the local training process in real time, and the performance indicator data includes but is not limited to computing throughput, memory usage, GPU utilization and communication overhead; The data preprocessing module is used to clean and standardize the collected performance indicator data to obtain processed performance indicator data; The indicator aggregation module is used to aggregate and count the processed performance indicator data of different training nodes and send them to the automatic tuning client.
3. According to claim 1, an online automatic tuning and compiling system for large model training is characterized in that: The automatic tuning client includes a data acquisition interface module, an indicator preprocessing module, a configuration update module and a communication interface module; The data acquisition interface module is used to receive performance indicator data from the data collector; The indicator preprocessing module is used to perform outlier detection on the performance indicator data from the data collector, remove noise data, and then standardize the data of different dimensions to obtain standardized performance indicator data; finally, perform time series aggregation on the standardized performance indicator data to obtain time series aggregated performance indicator data; The configuration update module is responsible for receiving the updated hyper-parameter configuration sent by the automatic tuning server and applying the updated hyper-parameter configuration to the local training process; The communication interface module is used to send the performance indicator data after time series aggregation to the automatic tuning server through the communication manager.
4. According to claim 1, an online automatic tuning and compiling system for large model training is characterized in that: The automatic tuning server includes a hyperparameter space initialization module, an indicator analysis module, an optimization strategy selection module and a configuration optimization module; The hyper-parameter space initialization module is used to initialize the hyper-parameter space according to different tuning strategies after receiving the automatic tuning request from the automatic tuning client; The indicator analysis module is used to analyze the performance indicator data from the automatic tuning client and build a performance evaluation model; the performance evaluation model comprehensively considers computing efficiency, memory efficiency and communication efficiency to generate a comprehensive performance score for the current configuration; The optimization strategy selection module is used to dynamically select a suitable optimization strategy according to the characteristics of the training scenario and the optimization target and the generated comprehensive performance score during the local training of the large model; the optimization strategy includes a Bayesian optimization strategy, a reinforcement learning strategy or a genetic algorithm strategy; The configuration optimization module adjusts the hyper-parameter configuration required for the next round of iterative training by selecting a suitable optimization strategy as the updated hyper-parameter configuration and calls the communication manager to send it to the automatic tuning client.
5. According to claim 1, an online automatic tuning and compiling system for large model training is characterized in that: The configuration manager includes a configuration storage module, a configuration update module, a version control layer and a monitoring layer; The configuration storage module is used to use a distributed database, which not only stores the optimal solution exceeding the configuration in the previous round of iterative training, but also stores the corresponding relationship between historical configuration records and performance data; The configuration update module is used to configure the updated hyper-parameter configuration received by the automatic tuning client into the Triton fusion operator in the large model framework after the user turns on the switch to start the online automatic tuning; The version control layer is used to manage the version evolution process of the configuration, and implements a fast rollback mechanism for the configuration, so that it can be restored to a stable version in a timely manner when performance anomalies are found; The monitoring layer is used to continuously monitor the effectiveness status and performance impact of the configuration, collect feedback information from each node, evaluate the effect of the configuration update in real time, and notify the system administrator in time when a configuration anomaly is found.
6. The online automatic tuning and compiling system for large model training according to claim 1, characterized in that: The communication manager uses HTTP or gRPC protocol to realize data transmission.
7. The online automatic tuning and compiling system for large model training according to claim 1, characterized in that: The Triton operator registrar includes a compilation module, a registration module, an optimization module and a verification module; The compiling module is used to compile the Triton fusion operator into a binary file; The registration module is used to register the compiled Triton fusion operator into the large model training framework; The optimization module is used to optimize the compiled Triton fusion operator; The verification module is used to verify the correctness of the compiled Triton fusion operator registered to the large model training framework.
8. The online automatic tuning and compiling system for large model training according to claim 1, characterized in that: The automatic tuning scheduler includes a load balancing module, a configuration synchronization module and a status monitoring module; The load balancing module is used to dynamically allocate tuning tasks according to the load status of each GPU node; The configuration synchronization module is used to ensure the consistency of hyperparameter configuration of all nodes in the cluster; The status monitoring module is used to monitor the operating status and tuning progress of each node in real time.
9. The online automatic tuning and compiling system for large model training according to claim 3, characterized in that: The automatic tuning client also includes an anomaly detection module and a log recording module; The anomaly detection module is used to monitor anomalies during the large model training process and notify the user when an anomaly is detected; The logging module is used to record key operations and performance changes during the tuning process to facilitate user auditing and system maintenance.
10. An online automatic tuning and compilation method for large model training, characterized in that: The method is implemented based on the online automatic tuning and compilation system for large model training described in any one of claims 1 to 9, and the method specifically includes: Through the data collector, the performance index data in the large model training process is collected as an optimization target; After the large model training is started, the automatic tuning client is responsible for sending the performance indicator data collected by the data collector to the automatic tuning server by calling the communication manager; After receiving the automatic tuning request from the automatic tuning client, the automatic tuning server initializes the hyper-parameter space according to different tuning strategies, and adjusts the hyper-parameter configuration required for the next round of iterative training as the updated hyper-parameter configuration according to the performance indicator data fed back by the automatic tuning client, and calls the communication manager to send it to the automatic tuning client; Through the configuration manager, a switch for starting online automatic tuning is provided to the user, and the updated hyperparameter configuration received by the automatic tuning client is configured into the Triton fusion operator in the large model framework; Register the Triton fusion operator to the large model training framework through the Triton operator register and receive the updated hyperparameter configuration; The communication manager is responsible for realizing data transmission between the automatic tuning client and the automatic tuning server by using the communication protocol; Through the automatic tuning scheduler, the hyperparameter configuration is passed to multiple GPU servers in the cluster that participate in large model training.
Citation Information
Patent Citations
Large-scale distributed inference engine and system based on deep learning
CN113568757A
Automatic deployment method based on TensorRT-LLM model reasoning acceleration service
CN117992078A
Triton compiler assembly line-oriented optimization system and optimization method
CN118605850A
KR20240052512A