Cluster reliability test method and system based on fault simulation
Through the cluster reliability testing method based on fault simulation, combined with deep learning and machine learning algorithms, the problems of low efficiency and inaccurate results in the existing technology are solved, and efficient and accurate reliability detection and evaluation are achieved.
Patent Information
- Application Number
- CN202510182427.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-20
AI Technical Summary
The existing cluster reliability testing methods have insufficient subjectivity of static analysis, low efficiency of manual fault injection, large differences between the simulation system and the real environment, and lack of a unified model or index system, making it difficult to accurately quantify the impact of faults on cluster reliability.
The cluster reliability testing method based on fault simulation is adopted, and fault simulation is obtained by obtaining injection parameter information and failover models, key performance indicators are collected in real time, and data analysis is performed by combining algorithms such as long and short-term memory networks, vector machines and attention mechanisms to generate reliability scores, impact analysis results and risk prediction results.
It realizes efficient, comprehensive and accurate simulation of complex fault scenarios, significantly improves testing efficiency and reliability detection levels, can intuitively reflect the reliability level of the cluster system, and facilitates comparison and analysis between different cluster systems.
Smart Images

Figure CN120179435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software reliability, and particularly to a method and system for cluster reliability testing based on fault simulation. Background Art
[0002] Modern distributed systems are widely used in fields such as cloud computing, big data processing, and high-performance computing. These systems are usually supported by clusters composed of multiple machines, and the nodes in the cluster jointly complete tasks through complex interaction mechanisms. With the expansion of the cluster scale, its reliability has become the core concern in system design and operation and maintenance. Once some nodes in the cluster fail, it may have a serious impact on the performance and availability of the entire system. Therefore, how to effectively evaluate the reliability of the cluster in the fault scenario has become the focus of the industry.
[0003] The current cluster reliability testing methods mainly focus on the following technical paths: Static analysis method: By analyzing code logic, architecture design, and configuration files, identify weak links that may cause faults. This method usually relies on expert experience and rule base support; Manual fault injection method: The tester manually creates faults (such as network disconnection, service crash, etc.) on some nodes of the cluster, and then observes the system behavior and evaluates the reliability; Simulation system method: Use virtualization technology to simulate the actual cluster operation scenario in a simulation environment, and test by simulating faults; Monitoring log analysis: By analyzing system log data for a long time, locate potential reliability problems in the system and evaluate their impacts. Although the above methods can help evaluate the reliability of the cluster to a certain extent, there are still the following significant problems: Static analysis cannot cover the complex dynamic interactions in actual operation, and the results are usually subjective, lacking comprehensiveness and accuracy; The manual fault injection process consumes a large amount of time and human resources, and it is difficult to systematically simulate various complex fault scenarios, with low efficiency and limited coverage; There are differences between the test environment of the simulation system and the real cluster environment, which is out of touch with reality, resulting in doubts about the reliability of the test results; The existing technologies lack a unified model or index system, and it is difficult to accurately quantify the impact of faults on the cluster reliability. Summary of the Invention
[0004] The present invention aims to provide a method and system for cluster reliability testing based on fault simulation to solve the above technical problems, improve the test efficiency and reliability detection level, and achieve accurate quantification of the reliability detection results.
[0005] To solve the above technical problems, the present invention provides a method for cluster reliability testing based on fault simulation, including the following steps:
[0006] Obtain injection parameter information, and perform fault simulation in the cluster based on the injection parameter information and the fault transfer model to obtain complex fault scenarios;
[0007] Obtain key performance indicator records in real time based on complex fault scenarios;
[0008] Execute reliability detection steps based on a preset reliability detection model and key performance indicator records: obtain the time series feature change results based on the key performance indicator records and the long short-term memory network algorithm; perform classification and fitting regression prediction based on the key performance indicator records, support vector machine, and attention mechanism to obtain classification prediction results; obtain the index mutation situation based on the key performance indicator records and the anomaly detection algorithm; obtain a data analysis report based on the time series feature change results, classification prediction results, and index mutation situation;
[0009] Obtain reliability scores, impact analysis results, and risk prediction results based on the data analysis report, weighted scoring algorithm, anomaly detection algorithm, and linear regression analysis to complete the reliability detection of the cluster.
[0010] The above solution first conducts fault simulation in the real cluster operating environment through corresponding fault simulation technologies to ensure that the test results are closer to the actual scenario, achieving efficient, comprehensive, and accurate fault simulation of various complex fault scenarios in the cluster; at the same time, it collects corresponding key performance indicator record information in real time, conducts data detection and analysis of the key performance indicator records through a reliability detection model, and finally generates reliability scores, impact analysis results, and risk prediction results in combination with the data analysis report to achieve accurate quantification of the reliability detection results; the reliability detection results obtained by this method can intuitively reflect the reliability level of the cluster system, facilitate comparative analysis between different cluster systems, significantly reduce manual intervention, and improve the test efficiency and detection ability.
[0011] Furthermore, obtain injection parameter information, and conduct fault simulation in the cluster based on the injection parameter information and the fault transfer model to obtain complex fault scenarios, including: constructing a fault transfer model based on Markov chains; obtaining injection parameter information, and using parallel injection technology to conduct fault simulation in the cluster based on the injection parameter information and the fault transfer model to obtain complex fault scenarios.
[0012] In the above solution, the fault transfer model constructed based on Markov chains can be efficiently and flexibly applied to various complex system fault transfer scenarios, and then different fault simulation data information, including multiple simulated faults, can be accurately injected into the specified nodes of the cluster through the injection parameter information and the fault transfer model, so as to fully realize the simulation of complex fault scenarios in the real cluster, provide detailed data support for subsequent analysis and performance evaluation, ensure that the test results are closer to the actual scenario, can provide direct reference for operation and maintenance decisions, and help in the formulation of fault location and optimization recovery strategies.
[0013] Further, when obtaining injection parameter information and performing fault simulation in the cluster based on the injection parameter information and the failover model to obtain a complex fault scenario, it further includes: obtaining fault data information based on real-time monitoring and feedback mechanism, and when the fault data information does not meet the preset controllable range, adjusting the injection parameter information based on the fault data information to obtain the final complex fault scenario.
[0014] In the above solution, by adopting real-time monitoring of the running state of the cluster during fault simulation and combining the feedback mechanism to obtain fault data information, the accuracy and controllability of the fault simulation process are ensured, so as to restore the complex fault scenario to the greatest extent.
[0015] Further, when obtaining injection parameter information and performing fault simulation in the cluster based on the injection parameter information and the failover model to obtain a complex fault scenario, it further includes: adjusting and obtaining fault data information based on real-time monitoring and feedback mechanism, and when the fault data information meets the preset abnormal range, adjusting the injection parameter information based on the fault data information and the preset abnormal recovery strategy to enable the cluster to resume normal operation.
[0016] In the above solution, by adopting real-time detailed monitoring of the running state of the cluster during fault simulation and combining the feedback mechanism to obtain fault data information, when an abnormality occurs in the fault simulation, the preset abnormal recovery strategy is used for adjustment and recovery in a timely manner to enable the cluster to resume normal operation, ensuring the controllability of the cluster system during the fault simulation and reliability detection process. The configuration of the cluster can also be adjusted in combination with the results obtained from the reliability detection to optimize the stability and fault tolerance.
[0017] Further, real-time obtaining of key performance indicator records based on the complex fault scenario includes: creating a custom data collection script; obtaining a fault injection signal, and obtaining source data information in real time based on the fault injection signal, the custom data collection script and the complex fault scenario; obtaining key performance indicator records based on the source data information and the preset data processing method.
[0018] In the above solution, by developing and creating a custom data collection script, the data collection process is automatically started by obtaining the fault injection signal, realizing the comprehensive real-time monitoring and recording of the running state of the cluster, and finally obtaining the key performance indicator records required for reliability detection, providing detailed data support for subsequent analysis and performance evaluation.
[0019] Further, in the step of performing reliability detection based on the preset reliability detection model and the key performance indicator records, the preset process of the reliability detection model includes: obtaining a training set and a test set based on the key performance indicator records; constructing a reliability detection model, and performing model training steps and model optimization steps based on the training set, the test set and the reliability detection model to obtain the preset reliability detection model.
[0020] In the above solution, a training set and a test set are formed by recording key performance indicators, and the reliability detection model is trained and optimized based on this to finally obtain a preset reliability detection model with higher accuracy and more reliable detection.
[0021] Furthermore, a reliability detection model is constructed, and the model training step and the model optimization step are executed based on the training set, the test set, and the reliability detection model to obtain a preset reliability detection model, including:
[0022] Construct a reliability detection model based on a deep learning framework and a machine learning algorithm;
[0023] Execute the model training step based on the training set, the test set, and the reliability detection model:
[0024] In the reliability detection model, obtain the training result of the time series feature change based on the training set and the long short-term memory network algorithm; obtain the mutation situation of the training index based on the training set and the anomaly detection algorithm; in the initial detection model, perform classification and fitting regression prediction based on the training result of the time series feature change, the mutation situation of the training index, the support vector machine, and the attention mechanism to obtain the reliability training model and the training evaluation result;
[0025] Execute the model optimization step based on the reliability training model: obtain the evaluation accuracy rate and the loss function based on the training evaluation result, and iteratively optimize the reliability training model based on the cross-validation method, the evaluation accuracy rate, and the loss function to obtain the reliability optimization model;
[0026] Obtain the test evaluation result based on the test set and the reliability optimization model, and obtain the preset reliability detection model when it is determined that the test evaluation result meets the preset requirements.
[0027] In the above solution, during the model training process, the long short-term memory network algorithm is used to analyze the change of time series features, the trend analysis and anomaly detection algorithms are used to identify the mutation of indicators in two ways according to the model and rules, that is, the anomaly detection algorithm, and the support vector machine and the attention mechanism are combined to fit the training set for classification and regression prediction to complete the construction and training process of the model. Then, the training model is iteratively optimized through cross-validation, and finally the test set is used for model testing. The model that meets the detection accuracy requirements is confirmed as the preset reliability detection model used for subsequent detection to support the subsequent generation of reliability scores and related analysis and prediction results, reports, etc., and provide data support for the system operation status evaluation and optimization suggestions.
[0028] Furthermore, based on the data analysis report, the weighted scoring algorithm, the anomaly detection algorithm, and the linear regression analysis, obtain the reliability score, the impact analysis result, and the risk prediction result to complete the reliability detection of the cluster, including:
[0029] Calculate the reliability score based on the data analysis report and the weighted scoring algorithm, and its calculation method satisfies the following formula:
[0030]
[0031] In the formula: R represents the reliability score, w1 represents the weight coefficient of the first key performance indicator, w2 represents the weight coefficient of the second key performance indicator, ΔT represents the first key performance indicator, and T base represents the initial value before the change of the first key performance indicator, ΔR represents the change value of the second key performance indicator, and R base represents the initial value before the change of the second key performance indicator;
[0032] Obtain the impact analysis result and risk prediction result based on the data analysis report, anomaly detection algorithm, and linear regression analysis;
[0033] Generate a visual reliability report based on the reliability score, impact analysis result, and risk prediction result to complete the reliability detection of the cluster.
[0034] In the above solution, by using the weighted scoring algorithm to calculate, evaluate, and further analyze the data analysis report, the reliability score is calculated, which accurately quantifies the reliability detection result when the cluster runs under the corresponding complex fault scenarios, and obtains the impact analysis result and risk prediction result, specifically analyzes the impact and its changes of the corresponding fault on the key performance indicators, and realizes the prediction of potential risks; and automatically generates a visual reliability report according to the analysis and calculation results, intuitively reflecting the reliability level of the cluster system, which is convenient for comparative analysis between different cluster systems.
[0035] Furthermore, generate a visual reliability report based on the reliability score, impact analysis result, and risk prediction result to complete the reliability detection of the cluster, including:
[0036] Generate optimization suggestions through the LLM model based on the reliability score, impact analysis result, and risk prediction result;
[0037] Generate a visual reliability report based on the reliability score, impact analysis result, risk prediction result, and optimization suggestions to complete the reliability detection of the cluster.
[0038] In the above solution, the LLM model is used to combine the existing knowledge graph related to operation, reliability score, impact analysis result, and risk prediction result to realize the automatic output of an analysis report containing optimization suggestions, providing intuitive data support and improvement methods for the system operation status evaluation and optimization adjustment, so as to improve the stability and fault tolerance of the cluster system.
[0039] The present invention first conducts fault simulation in a real cluster operating environment through corresponding fault simulation technologies, supports various types of fault injection and combined testing of complex scenarios, and significantly improves the coverage rate. At the same time, through model analysis, the interaction effects between multiple nodes can be discovered, realizing efficient, comprehensive, and accurate complex scenario fault simulation. Meanwhile, the running state of the cluster is monitored and fed back in real time, and abnormal recovery is performed when the fault simulation is abnormal, ensuring the controllability and safety of the fault simulation and subsequent testing processes. Furthermore, by collecting the corresponding key performance index record information in real time, based on the reliability detection model optimized by training, on the basis of complex fault simulation, data detection and analysis are carried out through key performance index records, and finally, reliability scores, impact analysis results, and risk prediction results are generated through data analysis reports, realizing real-time acquisition and analysis of the changes in the running state of the cluster before and after the fault and accurately quantifying the impact of the fault on reliability. The reliability detection results obtained by this method can intuitively reflect the reliability level of the cluster system, convert complex performance data into reliability indicators that are easy to understand and make decisions, facilitate comparative analysis between different cluster systems, realize the automation of the entire process of fault injection, data collection, and analysis, significantly reduce manual intervention, and improve the testing efficiency.
[0040] The present invention also provides a cluster reliability testing system based on fault simulation for implementing a cluster reliability testing method based on fault simulation, including:
[0041] A fault simulation module, configured to obtain injection parameter information and perform fault simulation in the cluster based on the injection parameter information and a fault transfer model to obtain a complex fault scenario;
[0042] A data collection module, configured to obtain key performance index records in real time based on the complex fault scenario;
[0043] A reliability detection module, configured to perform reliability detection steps based on a preset reliability detection model and key performance index records: obtain a time series feature change result based on the key performance index records and a long short-term memory network algorithm; perform classification and fitting regression prediction based on the key performance index records, a support vector machine, and an attention mechanism to obtain a classification prediction result; obtain the index mutation situation based on the key performance index records and an anomaly detection algorithm; obtain a data analysis report based on the time series feature change result, the classification prediction result, and the index mutation situation;
[0044] An intelligent analysis and generation module, configured to obtain reliability scores, impact analysis results, and risk prediction results based on the data analysis report, a weighted scoring algorithm, an anomaly detection algorithm, and a linear regression analysis to complete the reliability detection of the cluster.
[0045] The system provided by the present invention first uses corresponding fault simulation techniques through the fault simulation module to simulate faults in the real cluster running environment, ensuring that the test results are closer to the actual scenario, and realizing efficient, comprehensive and accurate fault simulation of various complex fault scenarios in the cluster. At the same time, based on the data acquisition module, the corresponding key performance index record information is collected in real time. Then, through the reliability detection module, based on the reliability detection model, the data detection and analysis of the key performance index records are carried out to generate the corresponding data analysis report. Finally, combined with the intelligent analysis generation module, the reliability score, impact analysis result and risk prediction result are generated based on the data analysis report, realizing the accurate quantification of the reliability detection result. The reliability detection result obtained by this method can intuitively reflect the reliability level of the cluster system, facilitating the comparative analysis between different cluster systems. At the same time, it significantly reduces manual intervention and improves the test efficiency and detection ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Schematic diagram of a cluster reliability test method based on fault simulation provided by an embodiment of the present invention;
[0047] Figure 2 Schematic diagram of a cluster reliability evaluation framework test method based on fault simulation provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] Embodiment 1:
[0050] This embodiment provides a cluster reliability test method based on fault simulation, as Figure 1 shown, including the following steps:
[0051] S1: Obtain injection parameter information, and simulate faults in the cluster based on the injection parameter information and the fault transfer model to obtain complex fault scenarios;
[0052] S2: Obtain key performance index records in real time based on the complex fault scenarios;
[0053] S3: Execute the reliability detection steps based on the preset reliability detection model and the key performance indicator records: Obtain the time series feature change results based on the key performance indicator records and the long short-term memory network algorithm; perform classification and fitting regression prediction based on the key performance indicator records, the support vector machine, and the attention mechanism to obtain the classification prediction results; obtain the index mutation situation based on the key performance indicator records and the anomaly detection algorithm; obtain the data analysis report based on the time series feature change results, the classification prediction results, and the index mutation situation;
[0054] S4: Obtain the reliability score, the impact analysis result, and the risk prediction result based on the data analysis report, the weighted scoring algorithm, the anomaly detection algorithm, and the linear regression analysis to complete the reliability detection of the cluster.
[0055] The above solution first conducts fault simulation in the real cluster running environment through the corresponding fault simulation technology to ensure that the test results are closer to the actual scenario, realizing efficient, comprehensive, and accurate fault simulation of multiple types of complex fault scenarios in the cluster; at the same time, it collects the corresponding key performance indicator record information in real time, conducts data detection and analysis of the key performance indicator records through the reliability detection model, and finally generates the reliability score, the impact analysis result, and the risk prediction result in combination with the data analysis report, realizing the precise quantification of the reliability detection results; the reliability detection results obtained by this method can intuitively reflect the reliability level of the cluster system, facilitating the comparative analysis between different cluster systems, and at the same time significantly reducing manual intervention and improving the test efficiency and detection ability.
[0056] Optionally, step S1 includes: constructing a fault transfer model based on the Markov chain; obtaining the injection parameter information, and using the parallel injection technology to conduct fault simulation in the cluster based on the injection parameter information and the fault transfer model to obtain complex fault scenarios.
[0057] In the specific implementation process, through the random fault injection algorithm and the fault transfer model based on Markov chain, faults are injected into specified nodes in the distributed cluster. Combining system resource monitoring signals (such as CPU, memory, disk I / O, etc.) and network status signals (such as packet loss rate, latency, etc.), accurate simulation of various fault types such as network latency, node downtime, disk failure, and resource contention is achieved. Moreover, it supports flexible configuration of various injection parameter information such as the scope, time, and frequency of fault injection, and improves the simulation efficiency through parallel injection technology. Combining real-time monitoring and feedback mechanisms ensures the accuracy and controllability of the simulation process, thereby restoring complex fault scenarios to the greatest extent; when obtaining injection parameter information, the target cluster can be intuitively and conveniently selected by the user in the tool interface, and the specific parameters of fault injection can be flexibly configured, including selecting fault types (such as network latency, node downtime, disk failure, resource contention, etc.), the scope of affected nodes, the test duration, and the preset abnormal recovery strategy; at the same time, the fault trigger condition and recovery mechanism can also be set to ensure the controllability and security of the test process.
[0058] Optionally, when performing step S1, it further includes: obtaining fault data information based on the real-time monitoring and feedback mechanism, and when the fault data information does not meet the preset controllable range, adjusting the injection parameter information based on the fault data information to obtain the final complex fault scenario.
[0059] Optionally, when performing step S1, it further includes: adjusting and obtaining fault data information based on the real-time monitoring and feedback mechanism, and when the fault data information meets the preset abnormal range, adjusting the injection parameter information based on the fault data information and the preset abnormal recovery strategy to make the cluster return to the normal operating state.
[0060] In the specific implementation process, the running state of the cluster is monitored and recorded in detail in real time through the real-time monitoring and feedback mechanism. The collected fault data information covers key performance indicators such as CPU utilization rate, memory occupancy, disk I / O rate, and network latency, and records system behavior changes through logs, providing detailed data support for subsequent analysis and performance evaluation, which helps in fault location and the formulation of optimization recovery strategies; in the actual application process, assuming a specific process parameter range is given, the sampling frequency of real-time data supports a sampling interval of 1 second to 10 minutes, and the distributed cluster with the adapted cluster scale supports 2 to 1000 nodes; the supported fault type range includes network latency (50ms to 1000ms), CPU resource occupancy (10% to 100%), memory leak (10% - 100%), disk I / O bottleneck (10MB / s - 100MB / s), etc. When the fault data information obtained by real-time sampling exceeds the supported fault type range, that is, when it is considered to meet the preset abnormal range, cluster abnormal recovery is performed to make the cluster return to the normal operating state.
[0061] Optionally, step S2 includes: creating a custom data acquisition script; obtaining a fault injection signal, and obtaining source data information in real time based on the fault injection signal, the custom data acquisition script, and the complex fault scenario; obtaining a key performance indicator record based on the source data information and a preset data processing method.
[0062] In the specific implementation process, a custom monitoring data acquisition script is developed to integrate the real-time monitoring and feedback mechanism. The data acquisition process is automatically started while the fault injection is simulated to obtain the corresponding data information. The relevant operation performance indicators of the distributed upload cluster are transmitted, and the reliability of the performance indicator transmission is ensured by combining message queues such as kafka. The key performance indicators during the operation of the distributed cluster are obtained by consuming data through data processing frameworks such as spark and mapreduce, including node resource utilization rate, network latency, task completion rate, application monitoring indicators, and service availability, etc. Finally, a key performance indicator record required for reliability detection is formed.
[0063] Optionally, in step S3, the process of presetting the reliability detection model includes: obtaining a training set and a test set based on the key performance indicator record; constructing a reliability detection model, and performing a model training step and a model optimization step based on the training set, the test set, and the reliability detection model to obtain a preset reliability detection model.
[0064] Optionally, constructing a reliability detection model, and performing a model training step and a model optimization step based on the training set, the test set, and the reliability detection model to obtain a preset reliability detection model includes:
[0065] Constructing a reliability detection model based on a deep learning framework and a machine learning algorithm;
[0066] Performing a model training step based on the training set, the test set, and the reliability detection model:
[0067] Obtaining a training result of time series feature change based on the training set and the long short-term memory network algorithm in the reliability detection model; obtaining the mutation situation of training indicators based on the training set and an anomaly detection algorithm; performing classification and fitting regression prediction based on the training result of time series feature change, the mutation situation of training indicators, a support vector machine, and an attention mechanism in the initial detection model to obtain a reliability training model and a training evaluation result;
[0068] Performing a model optimization step based on the reliability training model: obtaining an evaluation accuracy rate and a loss function based on the training evaluation result, and iteratively optimizing the reliability training model based on the cross-validation method, the evaluation accuracy rate, and the loss function to obtain a reliability optimization model;
[0069] Obtaining a test evaluation result based on the test set and the reliability optimization model, and obtaining a preset reliability detection model when it is determined that the test evaluation result meets the preset requirements.
[0070] In the specific implementation process, it is constructed by adopting a deep learning framework and combining machine learning algorithms. During the design process, temporal features are extracted based on a multi-layer neural network (such as LSTM or CNN), and classification and regression analysis are carried out in combination with decision trees, support vector machines (SVM), or attention mechanism models; in terms of data processing, after collecting the key performance indicators records and fault event logs during the system operation, data preprocessing is first carried out through steps such as data cleaning, normalization, and feature extraction; in the training stage, historical fault datasets are used for processing, the processed data is labeled to obtain training and test datasets, supervised learning is carried out using the training set, and model parameters are optimized using cross-validation, the loss function and accuracy are calculated, and the model performance is iteratively optimized through the loss function and accuracy evaluation metrics.
[0071] In the specific implementation process, when evaluating the reliability of the model, the stability of the calculation metrics is evaluated through mean, standard deviation, and deviation analysis. The mean can reflect the central tendency of the data, the standard deviation is used to measure the dispersion degree of the data, and deviation analysis helps to clarify the specific situation of the data deviating from the mean, and trend analysis and anomaly detection algorithms are used to identify index mutations in two ways based on the model and rules. For example, the CPU utilization rate continuously exceeding 80% may indicate a resource bottleneck, and the network packet loss rate exceeding 5% may cause communication anomalies, which are manifested as an extended system response time or service unavailability; during the prediction process, by comparing the system state changes before and after the fault, the deviation and abnormal distribution of key metrics are calculated in real time, supporting the subsequent generation of reliability model scores and trend analysis reports, providing data support for system operation state evaluation and optimization.
[0072] Optionally, step S4 includes:
[0073] Calculating the reliability score based on the data analysis report and the weighted scoring algorithm, and its calculation method satisfies the following formula:
[0074]
[0075] In the formula: R represents the reliability score, w1 represents the weight coefficient of the first key performance indicator, w2 represents the weight coefficient of the second key performance indicator, ΔT represents the first key performance indicator, T base represents the initial value of the first key performance indicator before the change, ΔR represents the change value of the second key performance indicator, R base represents the initial value of the second key performance indicator before the change;
[0076] Obtaining the impact analysis result and risk prediction result based on the data analysis report, anomaly detection algorithm, and linear regression analysis;
[0077] Generate a visual reliability report based on the reliability score, impact analysis results, and risk prediction results to complete the reliability detection of the cluster.
[0078] Optionally, generate a visual reliability report based on the reliability score, impact analysis results, and risk prediction results to complete the reliability detection of the cluster, including:
[0079] Generate optimization suggestions through the LLM model based on the reliability score, impact analysis results, and risk prediction results;
[0080] Generate a visual reliability report based on the reliability score, impact analysis results, risk prediction results, and optimization suggestions to complete the reliability detection of the cluster.
[0081] In the specific implementation process, calculate the average change trend, change amplitude, and change speed of each indicator unit, use anomaly detection algorithms to analyze the impact of faults on each indicator, such as a 10ms increase in network latency resulting in a 20% increase in response time, and combine the trend of test data to predict potential risks. Finally, for different performance bottlenecks and abnormal manifestations, corresponding prompt words can be constructed in the report. For example, within the process parameter range [specific range], when [specific performance bottleneck or abnormal manifestation 1] occurs, the prompt word is [prompt word 1], when [specific performance bottleneck or abnormal manifestation 2] occurs, the prompt word is [prompt word 2], and so on; and with the help of LLM model reasoning, generate targeted optimization suggestions; and finally output the analysis results in the form of charts and text, and through the large model combined with the existing knowledge graph related to operation, reliability detection related data and scoring, automatically output a visual reliability report, and the analysis report designs reliability evaluation indicators, reliability scores, impact analysis, risk prediction and optimization suggestions, improvement plans, and intuitively displays the test results and optimization suggestions through the icons in the report.
[0082] In this embodiment, first, fault simulation is carried out in a real cluster running environment through corresponding fault simulation technologies, which supports multiple types of fault injection and combined tests of complex scenarios, significantly improving the coverage rate. At the same time, through model analysis, the interaction effects between multiple nodes can be discovered, realizing efficient, comprehensive, and accurate fault simulation of complex scenarios. Meanwhile, the running state of the cluster is monitored and fed back in real time, and abnormal recovery is performed when the fault simulation is abnormal, ensuring the controllability and safety of the fault simulation and subsequent test processes. Furthermore, by collecting the corresponding key performance indicator record information in real time, based on the reliability detection model optimized by training, on the basis of complex fault simulation, data detection and analysis are carried out through key performance indicator records, and finally, reliability scores, impact analysis results, and risk prediction results are generated through data analysis reports, realizing the precise quantification of reliability detection results. The reliability detection results obtained by this method can intuitively reflect the reliability level of the cluster system, facilitating comparative analysis between different cluster systems, realizing the automation of the entire process of fault injection, data collection, and analysis, significantly reducing manual intervention, and improving test efficiency.
[0083] Embodiment 2:
[0084] This embodiment provides a cluster reliability test system based on fault simulation for implementing a cluster reliability test method based on fault simulation, including:
[0085] A fault simulation module, configured to obtain injection parameter information and perform fault simulation in the cluster based on the injection parameter information and a fault transfer model to obtain a complex fault scenario;
[0086] A data collection module, configured to obtain key performance indicator records in real time based on the complex fault scenario;
[0087] A reliability detection module, configured to perform reliability detection steps based on a preset reliability detection model and key performance indicator records: obtain the time series feature change result based on the key performance indicator records and the long short-term memory network algorithm; perform classification and fitting regression prediction based on the key performance indicator records, a support vector machine, and an attention mechanism to obtain a classification prediction result; obtain the index mutation situation based on the key performance indicator records and an anomaly detection algorithm; obtain a data analysis report based on the time series feature change result, the classification prediction result, and the index mutation situation;
[0088] An intelligent analysis and generation module, configured to obtain reliability scores, impact analysis results, and risk prediction results based on the data analysis report, a weighted scoring algorithm, an anomaly detection algorithm, and linear regression analysis to complete the reliability detection of the cluster.
[0089] The system provided in this embodiment first uses corresponding fault simulation technologies through the fault simulation module to conduct fault simulations in a real cluster running environment, ensuring that the test results are closer to the actual scenario, and realizing efficient, comprehensive, and accurate fault simulations of multiple types of complex fault scenarios in the cluster. At the same time, based on the data acquisition module, key performance indicator record information is collected in real time. Then, through the reliability detection module, data detection and analysis of the key performance indicator records are realized based on the reliability detection model to generate corresponding data analysis reports. Finally, combined with the intelligent analysis generation module, reliability scores, impact analysis results, and risk prediction results are generated based on the data analysis reports, realizing the precise quantification of reliability detection results. The reliability detection results obtained by this method can intuitively reflect the reliability level of the cluster system, facilitating comparative analysis between different cluster systems. At the same time, it significantly reduces manual intervention, improves test efficiency and detection capabilities. Moreover, the system design is modular and general, and can support adaptation to clusters of different scales to expand support for more fault types and data analysis capabilities according to requirements.
[0090] Embodiment Three:
[0091] This embodiment provides a cluster reliability evaluation framework based on fault simulation, which is implemented by using the reliability test method provided in the embodiment. As Figure 2 shown, its test method includes the following steps:
[0092] S31: Target configuration and parameter setting, obtaining injection parameter information: Select the target cluster through the tool interface, configure the fault type, injection range, test duration, and preset abnormal recovery strategy, and define key performance indicators and evaluation criteria;
[0093] S32: Fault injection and simulation execution, obtaining complex fault scenarios: Inject simulated faults based on the injection parameter information into the specified nodes to obtain complex fault scenarios, and monitor the cluster status in real time;
[0094] S33: Data collection and analysis calculation: Collect various key performance indicators, use the weighted scoring algorithm and anomaly detection analysis to evaluate the fault impact, and generate reliability scores and performance trend analysis in combination with the preset reliability detection model;
[0095] S34: Report generation and result evaluation: Automatically generate a visual reliability report based on the reliability scores and performance trend analysis by means of the inference generation technology of the LLM model, including reliability scores, fault impact analysis, and improvement suggestions, such as optimizing resource configuration, improving network performance, or increasing redundant design;
[0096] S35: Recovery and Optimization Adjustment: If an exception occurs during the fault injection simulation process, execute the preset exception recovery strategy to restore the normal operating state, and further adjust the system configuration according to the improvement suggestions to optimize stability and fault tolerance.
[0097] Embodiment 4:
[0098] This embodiment provides a cluster reliability test application solution based on single-point faults, including the following process: In a distributed cluster composed of 100 nodes, randomly select one node to inject a 500ms network latency fault, and use the reliability detection model for analysis and evaluation. First, collect key indicators such as system throughput and response time, and calculate the change rate before and after the fault. Quantify the reliability score through a weighted scoring algorithm. The specific formula is:
[0099]
[0100] Among them, R0 represents the reliability score, w 10 represents the weight coefficient of the response time, w 20 represents the weight coefficient of the throughput, ΔT0 represents the response time, T base0 represents the initial value before the change in the response time, ΔR0 represents the change value of the throughput, R base0 represents the initial value before the change in the throughput;
[0101] When the analysis shows that the reliability score drops by 85 points month-on-month, it indicates that the system has a certain performance decline in the current fault scenario. Based on the evaluation results, it is recommended to optimize the network routing configuration to reduce latency and improve the stability of the cluster.
[0102] In the specific implementation process, perform multi-point combined fault testing based on the reliability testing method. The example is as follows: In the same cluster, simultaneously simulate the CPU resource occupancy (90%) of two nodes and the disk I / O bottleneck (90%) of one node. The test tool records the decrease ratio of the task completion rate and the increase in the response time, and finally inputs them into the preset reliability detection model to generate a reliability score and a visual reliability report, automatically generate a recommendation to increase the load balancing strategy, and improve the model through the visual reliability report.
[0103] In the specific implementation process, when conducting large-scale cluster fault testing based on the reliability testing method, the example is as follows: In a distributed system composed of 500 nodes, simulate the fault that the memory of 10 nodes is full (100%), collect relevant data, and generate a detailed report; the report includes a performance change curve, showing the changes in key indicators such as the memory usage, response time, and throughput of the system when the fault occurs, and records the performance of the system during the recovery period. For example, when the memory is full, the response time may increase significantly, and the throughput may drop sharply, while during the recovery process, these indicators gradually return to normal levels; based on these data, the report puts forward suggestions for optimizing the memory management strategy, such as increasing the memory resource allocation, improving the memory recycling mechanism, or adjusting the load balancing strategy, to improve the stability and fault recovery ability of the system.
[0104] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A cluster reliability testing method based on fault simulation, characterized in that: The following steps are involved: Obtain injection parameter information, and perform fault simulation in the cluster based on the injection parameter information and the failover model to obtain complex fault scenarios; Obtain key performance indicator records in real time based on complex fault scenarios; Perform reliability testing steps based on the preset reliability testing model and key performance indicator records: Obtain the results of time series feature changes based on key performance indicator records and long short-term memory network algorithms; Based on key performance indicator records, vector machines and attention mechanisms, classification and fitting regression prediction are performed to obtain classification prediction results; based on key performance indicator records and anomaly detection algorithms, indicator mutation conditions are obtained; based on time series feature change results, classification prediction results and indicator mutation conditions, data analysis reports are obtained; Based on data analysis reports, weighted scoring algorithms, anomaly detection algorithms and linear regression analysis, reliability scores, impact analysis results and risk prediction results are obtained to complete the reliability detection of the cluster.
2. A cluster reliability testing method based on fault simulation according to claim 1, characterized in that: The step of obtaining the injection parameter information and performing a fault simulation in the cluster based on the injection parameter information and the fault transfer model to obtain a complex fault scenario includes: Building a failover model based on Markov chains; The injection parameter information is obtained, and the parallel injection technology is used to perform fault simulation in the cluster based on the injection parameter information and the failover model to obtain complex fault scenarios.
3. A cluster reliability testing method based on fault simulation according to claim 1, characterized in that: When obtaining the injection parameter information and performing fault simulation in the cluster based on the injection parameter information and the fault transfer model to obtain a complex fault scenario, it also includes: obtaining fault data information based on a real-time monitoring and feedback mechanism, and when the fault data information does not conform to a preset controllable range, adjusting the injection parameter information based on the fault data information to obtain a final complex fault scenario.
4. A cluster reliability testing method based on fault simulation according to claim 3, characterized in that: When the injection parameter information is obtained and a fault simulation is performed in the cluster based on the injection parameter information and the fault transfer model to obtain a complex fault scenario, the method further includes: Based on the real-time monitoring and feedback mechanism, the fault data information is adjusted and acquired. When the fault data information meets the preset abnormal range, the injection parameter information is adjusted based on the fault data information and the preset abnormal recovery strategy to restore the cluster to normal operation.
5. The cluster reliability testing method based on fault simulation according to claim 1 is characterized in that: The real-time acquisition of key performance indicator records based on complex fault scenarios includes: Create custom data collection scripts; Obtain fault injection signals, and obtain source data information in real time based on fault injection signals, custom data collection scripts, and complex fault scenarios; Obtain key performance indicator records based on source data information and preset data processing methods.
6. A cluster reliability testing method based on fault simulation according to claim 1, characterized in that: In the step of performing reliability detection based on the preset reliability detection model and key performance indicator records, the reliability detection model preset process includes: Obtain training and test sets based on key performance indicator records; Construct a reliability detection model, perform model training steps and model optimization steps based on the training set, test set and reliability detection model, and obtain a preset reliability detection model.
7. A cluster reliability testing method based on fault simulation according to claim 6, characterized in that: The reliability detection model is constructed, and a model training step and a model optimization step are performed based on the training set, the test set and the reliability detection model to obtain a preset reliability detection model, including: Build a reliability detection model based on deep learning framework and machine learning algorithm; Perform model training steps based on the training set, test set, and reliability detection model: In the reliability detection model, the training results of time series feature changes are obtained based on the training set and the long short-term memory network algorithm; the mutation of training indicators is obtained based on the training set and the anomaly detection algorithm; in the initial detection model, classification and fitting regression prediction are performed based on the training results of time series feature changes, the mutation of training indicators, vector machines and attention mechanisms to obtain the reliability training model and training evaluation results; Perform a model optimization step based on the reliability training model: obtain an evaluation accuracy and a loss function based on the training evaluation results, and iteratively optimize the reliability training model based on a cross-validation method, the evaluation accuracy and the loss function to obtain a reliability optimization model; The test evaluation results are obtained based on the test set and the reliability optimization model, and the preset reliability detection model is obtained when it is determined that the test evaluation results meet the preset requirements.
8. The cluster reliability testing method based on fault simulation according to claim 1 is characterized in that: The reliability score, impact analysis result and risk prediction result are obtained based on the data analysis report, weighted scoring algorithm, anomaly detection algorithm and linear regression analysis to complete the reliability detection of the cluster, including: The reliability score is calculated based on the data analysis report and the weighted scoring algorithm, and the calculation method satisfies the following formula: Where: R represents the reliability score, w1 represents the weight coefficient of the first key performance indicator, w2 represents the weight coefficient of the second key performance indicator, ΔT represents the first key performance indicator, T base represents the initial value of the first key performance indicator before the change, ΔR represents the change value of the second key performance indicator, and R base represents the initial value of the second key performance indicator before the change; Obtain impact analysis results and risk prediction results based on data analysis reports, anomaly detection algorithms and linear regression analysis; Generate a visual reliability report based on the reliability score, impact analysis results, and risk prediction results to complete the reliability detection of the cluster.
9. A cluster reliability testing method based on fault simulation according to claim 8, characterized in that: The visual reliability report is generated based on the reliability score, impact analysis results and risk prediction results to complete the reliability detection of the cluster, including: Generate optimization suggestions based on reliability scores, impact analysis results, and risk prediction results through the LLM model; Generate a visual reliability report based on reliability scores, impact analysis results, risk prediction results, and optimization suggestions to complete reliability testing of the cluster.
10. A cluster reliability testing system based on fault simulation, characterized in that: A cluster reliability testing method based on fault simulation according to any one of claims 1 to 9 is implemented, comprising: The fault simulation module is used to obtain the injection parameter information and perform fault simulation in the cluster based on the injection parameter information and the fault transfer model to obtain complex fault scenarios; Data acquisition module, used to obtain key performance indicator records in real time based on complex fault scenarios; The reliability detection module is used to perform reliability detection steps based on the preset reliability detection model and key performance indicator records: obtain the time series feature change results based on the key performance indicator records and the long short-term memory network algorithm; perform classification and fitting regression prediction based on the key performance indicator records, vector machines and attention mechanisms to obtain classification prediction results; obtain indicator mutations based on the key performance indicator records and anomaly detection algorithms; obtain data analysis reports based on the time series feature change results, classification prediction results and indicator mutations; The intelligent analysis generation module is used to obtain reliability scores, impact analysis results and risk prediction results based on data analysis reports, weighted scoring algorithms, anomaly detection algorithms and linear regression analysis, and complete the reliability detection of the cluster.
Citation Information
Cited By
Intelligent control system with AI identification in traditional Chinese medicine extraction and concentration process
CN120370660A
Notebook computer multi-scene pressure test method and system applied to factory detection
CN120429182A
Water meter data acquisition method based on open source gap
CN120475284A
A water meter data collection method based on open source Hongmeng
CN120475284B
LTE-M rail transit performance test method and system
CN122476371A