Intelligent elasticity test system based on HPC multi-dimensional performance indexes
Through the intelligent elastic testing system of multi-dimensional performance indicators, combined with multi-source data acquisition, intelligent fault prediction and dynamic checkpoint management, the problem of high failure rate in high-performance computing systems is solved, rapid failure recovery and task continuity are achieved, and system efficiency is improved.
Patent Information
- Application Number
- CN202510379809.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
The failure rate in high-performance computing systems is high, the traditional checkpoint/restart strategy has large storage overhead, long recovery delay, and insufficient dynamic adaptability, which affects task continuity and resource utilization.
Multi-source performance data acquisition and preprocessing, intelligent fault prediction, dynamic checkpoint management and task migration mechanisms are adopted, combined with multi-dimensional performance indicators and machine learning, fault warning and rapid recovery are achieved, and task status is frozen using CRIU technology and migrated through Kubernetes scheduling.
It realizes low overhead and efficient failure recovery for HPC systems, improves the continuity of computing tasks and system efficiency, and reduces system interruption losses.
Smart Images

Figure CN120256181A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to high-performance computing systems, and particularly to an intelligent elastic testing system based on HPC multi-dimensional performance metrics. Background Art
[0002] With the rapid development of high-performance computing (HPC) systems and large-scale distributed systems, the continuous expansion of system scale has led to a significant increase in the failure rate. The mean time between failures (MTBF) of petaflop-scale systems is extremely low, seriously affecting task continuity and resource utilization. Traditional checkpoint / restart, replication, and static fault tolerance strategies have problems such as high checkpoint storage overhead, long recovery latency, and insufficient dynamic adaptability during large-scale parallel tasks or large model training processes. Summary of the Invention
[0003] To solve the above technical problems, the present invention proposes an intelligent elastic testing system based on HPC multi-dimensional performance metrics, which uses multi-source data acquisition, machine learning-driven fault prediction, dynamic checkpoint management, and task migration mechanisms to achieve early warning and rapid recovery of faults, thereby reducing the interruption losses caused by system faults and improving the continuity of computing tasks and the overall system efficiency.
[0004] The technical solution of the present invention includes:
[0005] An intelligent elastic testing system based on HPC multi-dimensional performance metrics, comprising:
[0006] A multi-source performance data acquisition and preprocessing module, configured to acquire multi-dimensional performance data and perform preprocessing;
[0007] An intelligent fault prediction module, configured to construct a fault feature data set based on multi-dimensional performance data and fault logs, and according to the fault feature data set, use a model to achieve real-time prediction of node anomalies, and when the risk value obtained according to the prediction result exceeds the risk threshold, trigger subsequent checkpoint adjustment and task migration processes;
[0008] A dynamic checkpoint management module, configured to perform dynamic checkpoint management by combining full-volume and incremental checkpoints: in the stable state of the system, save full-volume checkpoints regularly; in high-risk situations, only record state changes through incremental checkpoint technology;
[0009] A task migration management module, configured to quickly freeze the current task state through the CRIU technology and migrate the frozen task state to a healthy node when detecting node anomalies or fault risks;
[0010] A closed-loop feedback and system optimization module, configured to centrally record the operation data of each module in a unified log system, monitor through Prometheus, and dynamically adjust the parameters of each module according to the monitoring feedback.
[0011] Beneficial effects:
[0012] The present invention provides an intelligent elastic fault tolerance system and method based on multi-dimensional performance monitoring. Through the collaborative work of four major modules, namely multi-source data collection, intelligent fault prediction, dynamic checkpoint management, and task migration, it realizes the early warning and rapid recovery of common faults in HPC systems and large-scale distributed training tasks. Specifically, the solution achieves:
[0013] Comprehensive monitoring and fusion at the data level, providing refined inputs for fault prediction through multi-tool collection and unified preprocessing;
[0014] An intelligent fault prediction mechanism, realizing real-time warning by combining random forest, LSTM, and XGBoost models;
[0015] Dynamic and adaptive checkpoint management, significantly reducing storage and I / O overhead while retaining low-latency recovery capabilities;
[0016] A seamless task migration solution, realizing task state freezing, migration, and recovery by using CRIU and the cluster scheduling system;
[0017] A closed-loop feedback system, continuously optimizing the performance of each module through monitoring data to enhance the overall system robustness and fault tolerance capabilities.
[0018] This solution draws on the mature experience of Prometheus monitoring, CRIU freezing, Kubernetes scheduling, and existing fault prediction technologies. At the same time, it realizes system self-adaptive optimization through a multi-level checkpoint strategy and a closed-loop feedback mechanism, providing a low-overhead and high-efficiency fault tolerance solution for high-performance computing environments. Brief description of the drawings
[0019] Figure 1 A block diagram showing the intelligent elastic test system based on HPC multi-dimensional performance indicators of the present invention.
[0020] Figure 2 A flowchart showing the working process of the intelligent elastic test system based on HPC multi-dimensional performance indicators of the present invention.
[0021] Reference numerals: multi-source performance data collection and preprocessing module 110, intelligent fault prediction module 120, dynamic checkpoint management module 130, task migration management module 140, closed-loop feedback and system optimization module 150. Detailed implementation manners
[0022] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0023] The following describes the existing technical support in the art to facilitate understanding of the concept of the present invention, and the innovative points of the present invention are introduced therein.
[0024] 1. Multi-source performance data collection and preprocessing
[0025] 1.1 Data collection techniques and tools
[0026] Hardware-level data collection
[0027] Use PAPI (Performance Application Programming Interface) to collect CPU hardware performance counter data, such as cache hit rate, instruction execution rate, etc.; at the same time, monitor GPU utilization, video memory bandwidth, and power consumption through NVPROF or ROCm Profiler (ROCm performance analysis tool).
[0028] System-level metric collection
[0029] Use tools such as node_exporter and cAdvisor to obtain environmental metrics such as memory occupancy, network I / O, disk read and write, temperature, and voltage of each node in real time.
[0030] Data aggregation and storage
[0031] Use Prometheus to scrape data from each collection endpoint in pull mode and store it using a time series database; combine with Grafana to build a data visualization dashboard to display the real-time changes of key metrics.
[0032] 1.2 Data preprocessing and fusion
[0033] Data cleaning
[0034] Adopt techniques based on statistical methods and anomaly detection to eliminate outliers and ensure the quality of training data for subsequent prediction models.
[0035] Multi-dimensional data fusion
[0036] Use feature engineering to uniformly normalize data from different collection tools, construct a multi-dimensional performance index feature vector, and provide input for subsequent machine learning models. For example, for the same node, integrate metrics such as CPU utilization, memory bandwidth, temperature, and network latency to form a comprehensive feature.
[0037] In terms of real-time performance data processing, use the above data collection tools to capture and store metrics through Prometheus. Prometheus obtains multi-dimensional performance data from each performance monitoring endpoint in a pull mode, and stores the data in a time series. Use Grafana as a data visualization platform, combined with Prometheus' time series database, to build a performance monitoring dashboard, realize the visualization of multi-dimensional performance data, display key metrics such as CPU / GPU load, memory usage, and network traffic of the node, and at the same time provide trend analysis of historical performance data to facilitate viewing the system state before and after faults.
[0038] Existing technical support:
[0039] Existing systems (such as Prometheus + Grafana) have verified efficient data collection and visualization capabilities in large-scale systems.
[0040] 2. Intelligent fault prediction module
[0041] 2.1 Model construction and training
[0042] Model selection and integration
[0043] According to the fault characteristics, use the random forest model to capture the relationships between non-linear features; use the long short-term memory network (LSTM) to capture the trend changes in time series data; at the same time, introduce the XGBoost model to improve the classification accuracy of sudden faults (such as network interruptions, hardware crashes).
[0044] Dataset construction
[0045] Use historical performance data and fault logs to construct a fault feature dataset, and perform data annotation and cleaning to ensure that the training data covers various fault modes (node overheating, communication anomalies, hardware failures, etc.).
[0046] Hyperparameter tuning and cross-validation
[0047] Use grid search and cross-validation methods to tune the hyperparameters of each model to ensure that the model reaches the optimal state in terms of metrics such as prediction accuracy, recall rate, and F1 score.
[0048] 2.2 Model real-time prediction
[0049] Real-time data stream processing
[0050] The preprocessed multi-dimensional data is fed into the prediction module through a streaming computing framework (such as Spark Streaming or Kafka Streams) to achieve online fault risk assessment.
[0051] Prediction result fusion and decision-making
[0052] The prediction results of each model are weighted and fused, and a threshold is set to judge the fault risk level. When the risk value exceeds the predetermined threshold, subsequent dynamic adjustment and task migration operations are triggered.
[0053] Existing technical support:
[0054] Similar systems such as CheckFreq and ByteCheckpoint (a system for optimizing the training checkpoints of native large Pytorch models) have used some machine learning methods in dynamic detection and prediction. Here, deep models such as LSTM are further combined to improve the real-time performance and accuracy of prediction; in addition, existing research (such as the fault prediction method based on LSTM) provides effective algorithm verification and tuning solutions.
[0055] 3. Dynamic checkpoint management module
[0056] 3.1 Refinement of checkpoint technology
[0057] Dynamic adjustment of checkpoint saving strategy
[0058] Low-frequency full checkpoint saving is adopted when the system is stable; when the fault prediction module detects an increase in risk, it switches to high-frequency incremental checkpoint saving.
[0059] Incremental and multi-level checkpoints
[0060] The incremental checkpoint technology is used to only save the data blocks that have changed since the last checkpoint, reducing storage and I / O overhead; combined with multi-level checkpoints, different recovery strategies are designed for local and global faults respectively.
[0061] Data consistency and recovery verification
[0062] Non-volatile storage (such as SSD or distributed file system) is used to save checkpoint data, and data consistency is ensured through checksum and / or redundant coding technology. During the fault recovery process, the full checkpoint is first loaded, and then the incremental checkpoints are applied in sequence to restore the state.
[0063] 3.2 Technical implementation
[0064] API interface and framework integration
[0065] Develop standardized APIs for checkpoint saving and loading, and support integration into the distributed training environment through the PyTorch Checkpointing API or a custom library (such as DMTCP).
[0066] Scheduling interface call
[0067] The dynamically adjusted policy is triggered by the fault prediction module, and by calling the storage system and network interfaces, it realizes the asynchronous writing and synchronous verification of checkpoint data.
[0068] Existing technical support:
[0069] CheckFreq and ByteCheckpoint provide mature implementation experiences of incremental and multi-level checkpoint technologies. On this basis, this solution further implements a dynamic adjustment mechanism, drawing on the solutions of existing systems in I / O optimization and data consistency guarantee.
[0070] 4. Task migration management module
[0071] 4.1 Task status freezing and migration
[0072] Status freezing
[0073] When a fault risk is predicted or a node anomaly is detected, use CRIU (Checkpoint / Restore In Userspace) technology to quickly freeze the task status, including memory images, CPU registers, and process contexts.
[0074] Task migration strategy
[0075] Combined with the Kubernetes Pod scheduling API or the Slurm task scheduling interface, after the system detects an abnormal node health status, transfer the frozen task status to a healthy node.
[0076] Recovery and rescheduling
[0077] On the target node, first load the latest checkpoint data and use the frozen status information to restore the task, while dynamically adjusting the resource allocation to meet the computing requirements of the restored task.
[0078] 4.2 Implementation details of task migration
[0079] Communication and data transmission
[0080] During the task migration process, use an efficient distributed file transfer protocol or message queue to ensure that the task status data is not lost and has a low latency during network transmission.
[0081] Automated monitoring and feedback
[0082] The task migration module is deeply integrated with the monitoring system, which can provide real-time feedback on the migration progress and task recovery situation, and adjust the scheduling strategy according to the feedback, such as rescheduling the remaining tasks or updating the node health status information.
[0083] The innovation of the present invention
[0084] On this basis, the present invention introduces the CRIU technology to achieve finer-grained task state freezing and seamless recovery. When the fault prediction module determines that a node is abnormal or a hardware crash occurs, the task migration strategy is triggered. Task migration is achieved through Kubernetes Pod scheduling or Slurm job scheduling. The task process state is frozen on the faulty node, including the memory snapshot, CPU registers, and task context, and the frozen task state is transferred to the target node for recovery. When recovering the task on the target node, the most recent checkpoint data is preferentially loaded, and computing resources are dynamically allocated.
[0085] Through the above technologies and methods, a dynamic checkpoint management module is developed to provide a programming interface based on PyTorch to support efficient checkpoint saving and recovery in distributed training scenarios; the module is integrated with Kubernetes to automatically sense the task running state and node health status. At the same time, a task migration management module is developed to achieve deep integration with Kubernetes, automatically trigger task migration and resource reallocation, support logging and monitoring of the task migration process, and facilitate fault analysis and debugging.
[0086] Existing technical support:
[0087] The reliability of task automatic migration and resource scheduling has been verified in large-scale cluster environments by the Kubernetes and Slurm systems.
[0088] 5. Closed-loop feedback and system optimization
[0089] 5.1 Feedback mechanism design
[0090] Monitoring and logging
[0091] The operation results of each module (including the accuracy of fault prediction, checkpoint saving time, task migration success rate, etc.) are recorded in a unified log system in real time, and visualized using Prometheus monitoring metrics.
[0092] Automatic feedback adjustment
[0093] According to the actual operation situation of the system, the weights of the fault prediction model, the checkpoint saving frequency, and the task migration strategy are automatically adjusted to form an adaptive closed-loop optimization system.
[0094] 5.2 Continuous optimization strategy
[0095] Online model update
[0096] Use newly collected fault data and operation logs to perform online incremental training on the prediction model to improve the model's robustness.
[0097] Policy dynamic adjustment
[0098] Through experiments, verify and dynamically adjust the checkpoint saving and task migration thresholds corresponding to different fault types to adapt to different application scenarios and hardware environments.
[0099] Existing technology support:
[0100] Existing AIOps systems and intelligent monitoring platforms have accumulated a lot of experience in real-time data feedback and online model updates. This application draws on these mature technologies to achieve dynamic adaptive optimization of the overall system.
[0101] After describing the existing technology on which the present invention is based and the innovative points of the present invention, the following refers to Figure 1 Describe the intelligent elastic test system of the present invention based on HPC multi-dimensional performance indicators. As Figure 1 shown, the system includes:
[0102] The multi-source performance data acquisition and preprocessing module 110 is used to collect multi-dimensional performance data and perform preprocessing. In one embodiment, the multi-source performance data acquisition and preprocessing module 110 uses hardware-level (PAPI, NVPROF / ROCmProfiler) and system-level (node_exporter, cAdvisor) tools to collect multi-dimensional performance data such as CPU, GPU, memory, network, temperature, and power consumption.
[0103] The multi-source performance data acquisition and preprocessing module 110 also realizes data pulling and time series storage through Prometheus, and constructs a real-time visualization dashboard with the help of Grafana to display the collected multi-dimensional performance data.
[0104] Preprocessing the multi-dimensional performance data includes: cleaning, normalizing, and multi-dimensional feature fusion of the collected multi-dimensional performance data to construct a unified feature vector for subsequent analysis.
[0105] The intelligent fault prediction module 120 is used to construct a fault feature data set based on multi-dimensional performance data and fault logs, and according to the fault feature data set, use a model to realize real-time prediction of node anomalies, and when the risk value obtained according to the prediction result exceeds the risk threshold, trigger subsequent checkpoint adjustment and task migration processes.
[0106] In one embodiment, the model includes one of models such as random forest, long short-term memory network (LSTM), and XGBoost. Anomalies can include overheating, communication anomalies, hardware failures, etc.
[0107] The intelligent fault prediction module 120 also uses hyperparameter tuning and cross-validation to ensure that the model reaches the optimal level in terms of accuracy, recall rate, and F1 metric, and performs online prediction on real-time data through a streaming computing framework (such as Apache Spark streaming processing engine or Kafka streaming interface). Real-time data refers to multi-dimensional performance metric data continuously generated during the operation of HPC nodes. The data strongly associated with node anomalies include: CPU utilization, memory occupancy, network latency, I / O throughput, and hardware temperature. When the hardware is operating normally, the above performance data are all within a specific threshold range. Therefore, when the above indicators exceed the threshold range and do not recover within a short period of time, the system will determine that the HPC node has an anomaly and trigger anomaly detection.
[0108] The process of determining that the risk value exceeds the risk threshold includes: the intelligent fault prediction module 120 performs weighted fusion on the prediction results to obtain the risk value, and sets the risk threshold. When the risk value exceeds the risk threshold, it triggers the subsequent checkpoint adjustment and task migration processes.
[0109] The dynamic checkpoint management module 130 is used to perform dynamic checkpoint management by combining full-scale and incremental checkpoints: in the stable state of the system, full-scale checkpoints are saved regularly; in high-risk situations, only state changes are recorded through incremental checkpoint technology to reduce storage and I / O burdens.
[0110] The dynamic checkpoint management module 130 also uses a hierarchical checkpoint design to implement local and global recovery strategies, and ensures data consistency through non-volatile storage (SSD or distributed file system), and uses checksum or redundant coding to ensure the correctness of checkpoint data.
[0111] The dynamic checkpoint management module 130 is also deeply integrated with the distributed training environment through a standardized API interface (such as based on PyTorch CheckpointingAPI or DMTCP interface) to achieve fast and low-overhead saving and loading of checkpoint data.
[0112] The task migration management module 140 is used to quickly freeze the current task state through CRIU (Checkpoint / Restore In Userspace) technology and migrate the frozen task state to a healthy node when a node anomaly or fault risk is detected.
[0113] Among them, the task state includes memory image, CPU registers, and process context.
[0114] Migrating the frozen task status to a healthy node includes: invoking the Kubernetes Pod scheduling API or the Slurm task scheduling interface to migrate the frozen task status to a healthy node.
[0115] On the healthy node, first load the latest checkpoint data, and then use the frozen task status information to complete task recovery. At the same time, ensure low latency and integrity of data transmission through efficient distributed file transfer and message queues.
[0116] Efficient distributed file transfer refers to achieving file transfer in an efficient and reliable manner in a distributed environment, which usually involves various technical architectures and protocols. The specific technical architecture and protocol used depend on the actual scenario of the system.
[0117] The closed-loop feedback and system optimization module 150 is used to centrally record the operation data of each module in a unified log system, monitor it through Prometheus, and dynamically adjust the parameters of each module according to the monitoring feedback.
[0118] Among them, the operation data includes the accuracy of fault prediction, the latency of checkpoint saving, the success rate of task migration, etc.
[0119] Dynamically adjusting the parameters of each module according to the monitoring feedback includes: dynamically adjusting the weights of the models adopted by the intelligent fault prediction module 120, the checkpoint saving frequency, and the task migration threshold according to the monitoring feedback, realizing continuous online optimization of the system, and forming an adaptive closed-loop regulation mechanism.
[0120] The task migration threshold is a critical indicator set by the system during monitoring and dynamic regulation. When the state (such as performance indicators, response time, error rate, etc.) of a certain module or task in the system reaches or exceeds this threshold, the system will automatically trigger a task migration operation to migrate the task from the currently highly loaded or faulty node to a node with better state.
[0121] Next, refer to Figure 2 to describe an implementation process of the working process of the system of the present invention. As Figure 2 shown, the working process includes:
[0122] Data collection and preprocessing: Each node collects performance indicators in real time through tools such as PAPI, NVPROF, and node_exporter; the data is pulled and stored by Prometheus, and after data cleaning and feature fusion, a unified input vector is formed.
[0123] Real-time fault prediction: The preprocessed data stream is input into the intelligent fault prediction module, and the fault risk probability is calculated using random forest, LSTM, and XGBoost models; the weighted fusion prediction result is compared with a preset threshold, and if the risk increases, a warning signal is issued.
[0124] Dynamic checkpoint strategy adjustment: When the fault risk signal is triggered, the checkpoint management module automatically adjusts the checkpoint saving strategy, switches to the high-frequency and low-overhead incremental checkpoint saving mode, and at the same time ensures the regular update of the full checkpoint to support global recovery.
[0125] Task migration and recovery: The task migration module responds to the warning signal, calls CRIU to freeze the task status of the abnormal node, and triggers task migration through the Kubernetes or Slurm interface, quickly transmits the task status data to the healthy node and resumes operation.
[0126] Closed-loop feedback and adaptive optimization: The monitoring module records the execution of each link in real time, feeds back to the fault prediction and strategy adjustment module, supports online model update and dynamic threshold adjustment, and realizes continuous adaptive optimization of the system.
[0127] Although the present invention has been described in terms of a limited number of embodiments, those skilled in the art in this technical field will appreciate that other embodiments can be envisioned within the scope of the invention as thus described. Additionally, it should be noted that the language used in this specification has been principally selected for readability and teaching purposes rather than for the purpose of explaining or delimiting the subject matter of the invention.
Claims
1. An intelligent elastic test system based on HPC multi-dimensional performance indicators, characterized in that, It includes: A multi-source performance data acquisition and preprocessing module (110) for acquiring multi-dimensional performance data and performing preprocessing; An intelligent fault prediction module (120) for constructing a fault feature data set based on multi-dimensional performance data and fault logs, and using a model to realize real-time prediction of node anomalies according to the fault feature data set. When the risk value obtained according to the prediction result exceeds the risk threshold, it triggers the subsequent checkpoint adjustment and task migration processes; A dynamic checkpoint management module (130) for performing dynamic checkpoint management by combining full-scale and incremental checkpoints: in the stable state of the system, regularly save full-scale checkpoints; in high-risk situations, only record state changes through incremental checkpoint technology; A task migration management module (140) for quickly freezing the current task state through the CRIU technology and migrating the frozen task state to a healthy node when node anomalies or fault risks are detected; A closed-loop feedback and system optimization module (150) for centrally recording the operation data of each module in a unified log system, monitoring through Prometheus, and dynamically adjusting the parameters of each module according to the monitoring feedback.
2. The intelligent elastic test system based on HPC multi-dimensional performance indicators according to claim 1, wherein The multi-source performance data acquisition and preprocessing module (110) uses hardware-level and system-level tools to acquire CPU, GPU, memory, network, temperature, and power consumption data. The hardware level includes PAPI, NVPROF, and / or ROCm Profiler, and the system level includes node_exporter and cAdvisor.
3. The intelligent elastic test system based on HPC multi-dimensional performance indicators according to claim 1, characterized in that, The multi-source performance data acquisition and preprocessing module (110) also realizes data pulling and time series storage through Prometheus, and constructs a real-time visualization dashboard with the help of Grafana to display the acquired multi-dimensional performance data.
4. The intelligent elastic test system based on HPC multi-dimensional performance indicators according to claim 1, characterized in that, Preprocessing the multi-dimensional performance data includes: cleaning, normalizing, and multi-dimensional feature fusion of the acquired multi-dimensional performance data, and constructing a unified feature vector for subsequent analysis.
5. The intelligent elastic test system based on HPC multi-dimensional performance indicators according to claim 1, wherein, The model adopted by the intelligent fault prediction module (120) includes one of models such as random forest, long short-term memory network (LSTM), and XGBoost. Node anomalies include overheating, communication anomalies, and hardware failures.
6. The intelligent elastic test system based on HPC multi-dimensional performance indicators according to claim 5, wherein The intelligent fault prediction module (120) also uses hyperparameter tuning and cross-validation to ensure that the model reaches the optimal level in terms of accuracy, recall rate, and F1 metric, and performs online prediction on real-time data through a streaming computing framework.
7. The intelligent elastic test system based on HPC multi-dimensional performance indicators according to claim 5, characterized in that The process of judging that the risk value exceeds the risk threshold includes: the intelligent fault prediction module (120) performs weighted fusion on the prediction results to obtain the risk value, sets the risk threshold, and when the risk value exceeds the risk threshold, triggers the subsequent checkpoint adjustment and task migration processes.
8. The intelligent elastic test system based on HPC multi-dimensional performance indicators according to claim 1, characterized in that The dynamic checkpoint management module (130) is deeply integrated with the distributed training environment through a standardized API interface to realize the saving and loading of checkpoint data.
9. The intelligent elastic test system based on HPC multi-dimensional performance indicators according to claim 1, characterized in that The task status frozen by the task migration management module (140) includes memory images, CPU registers, and process contexts; migrating the frozen task status to a healthy node includes: calling the Kubernetes Pod scheduling API or the Slurm task scheduling interface to migrate the frozen task status to a healthy node.
10. The intelligent elastic test system based on HPC multi-dimensional performance indicators according to claim 1, wherein, The operation data recorded by the closed-loop feedback and system optimization module (150) includes the failure prediction accuracy rate, checkpoint saving latency, and task migration success rate; Dynamically adjusting the parameters of each module according to the monitoring feedback includes: dynamically adjusting the weights of the models adopted by the intelligent fault prediction module (120), the checkpoint saving frequency, and the task migration threshold according to the monitoring feedback.