Analysis method, device and equipment for operation fault of long-term stable use case and storage medium
By collecting and analyzing node operation data in a distributed storage system, a causal relationship network is constructed, solving the problem of identifying complex causal chains for failures in long-term stable use cases, and achieving efficient fault location and system optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINAN INSPUR DATA TECH CO LTD
- Filing Date
- 2026-01-16
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies struggle to efficiently identify and locate complex causal chains of failures in long-term stable use cases within distributed storage systems, resulting in inefficient fault analysis and a high risk of misjudgment.
By collecting operational data from storage nodes and business operations data, the system uses preset rules to determine business status, identify fault types, construct causal relationship networks, and generate detailed test reports.
It enables efficient analysis and precise location of faults in distributed storage systems, significantly improving the efficiency of fault diagnosis and repair, and enhancing the stability and reliability of the system.
Smart Images

Figure CN122020585A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault detection technology, and in particular to a method, apparatus, device, and storage medium for analyzing faults in long-term stable test cases. Background Technology
[0002] With the rapid development of cloud computing and big data technologies, distributed storage systems, due to their high scalability, high availability, and high fault tolerance, have become the core architecture for modern data processing and storage. In stability testing to evaluate the long-term reliability of storage systems, the continuous operation of long-term stable test cases is crucial. However, during testing, storage systems may experience service interruptions due to various reasons such as hardware failures, software vulnerabilities, and network anomalies. Currently, fault analysis mainly relies on manual experience or simple monitoring threshold judgments, resulting in problems such as untimely fault detection and low analysis efficiency. Traditional methods struggle to intelligently identify whether service interruptions have truly occurred, and given the numerous fault types in distributed systems, manual troubleshooting is not only time-consuming and labor-intensive but also prone to misjudgments. Furthermore, in real-world scenarios, service interruptions are often caused by a combination of factors, and traditional monitoring methods or simple rules are insufficient to unravel the complex causal chains and accurately pinpoint the root cause.
[0003] Therefore, in view of the shortcomings of existing technical solutions, the present invention provides a method for analyzing the failure of long-term stable use cases. Summary of the Invention
[0004] This application provides a method, electronic device, and storage medium for analyzing failures in long-term stable use cases, in order to at least solve the problem of difficulty in clarifying the complex causal chain between failure factors and the difficulty in accurately locating the root cause of the failure.
[0005] This application provides a method for analyzing failures in long-term stable test cases, applicable to distributed storage systems. The distributed storage system includes at least one storage node. The method includes: initiating long-term stable test cases; collecting node operation data from at least one storage node and corresponding business operation data of the distributed storage system; determining the operational status of the business based on preset disconnection rules and business operation data; in response to the business being in a disconnection state, determining the fault type based on preset fault rules and node operation data from at least one storage node; obtaining the causal relationship network corresponding to the fault type based on the fault type and node operation data from at least one storage node; and generating a test report for long-term stable test cases based on the node operation data of the storage node, business operation data, fault type, and causal relationship network.
[0006] This application also provides an electronic device, comprising: a memory for initiating long-term stable test cases, collecting node operation data of at least one storage node and business operation data corresponding to the distributed storage system; determining the operation status of the business based on preset disconnection rules and business operation data; in response to the business being in a disconnection state, determining the fault type based on preset fault rules and node operation data of at least one storage node; obtaining a causal relationship network corresponding to the fault type based on the fault type and node operation data of at least one storage node; and generating a test report for long-term stable test cases based on the node operation data of the storage nodes, business operation data, fault type, and causal relationship network.
[0007] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the following steps: initiating a long-term stable test case, collecting node operation data of at least one storage node and business operation data corresponding to the distributed storage system; determining the operation status of the business based on preset disconnection rules and business operation data; responding to the business being in a disconnection state, determining the fault type based on preset fault rules and node operation data of at least one storage node; obtaining the causal relationship network corresponding to the fault type based on the fault type and node operation data of at least one storage node; and generating a test report for the long-term stable test case based on the node operation data of the storage nodes, business operation data, fault type, and causal relationship network.
[0008] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the following steps: initiating a long-term stable test case; collecting node operation data of at least one storage node and business operation data corresponding to the distributed storage system; determining the operation status of the business based on preset disconnection rules and business operation data; responding to the business being in a disconnection state; determining the fault type based on preset fault rules and node operation data of at least one storage node; obtaining the causal relationship network corresponding to the fault type based on the fault type and node operation data of at least one storage node; and generating a test report for the long-term stable test case based on the node operation data of the storage nodes, business operation data, fault type, and causal relationship network.
[0009] This application enables efficient analysis and accurate localization of storage system faults during the execution of long-term stable test cases. It collects node operation data from at least one storage node and corresponding business operation data from the distributed storage system. Based on preset interruption rules and business operation data, it determines the operational status of the business. In response to a business interruption, it determines the fault type based on preset fault rules and node operation data from at least one storage node. Based on the fault type and node operation data from at least one storage node, it obtains the causal relationship network corresponding to the fault type. Finally, it generates a test report for long-term stable test cases based on the node operation data, business operation data, fault type, and causal relationship network. Therefore, it achieves efficient analysis and accurate localization of storage system faults during the execution of long-term stable test cases and generates comprehensive and standardized test problem reports, providing a reliable basis for storage system optimization. Simultaneously, it significantly saves manpower, greatly improves the efficiency of system problem investigation and repair, and enhances the stability and reliability of the distributed storage system. Attached Figure Description
[0010] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a method for analyzing operational failures of long-term stable use cases, provided in an embodiment of this application; Figure 2 A flowchart illustrating the test report generation process for a method for analyzing long-term stable test case execution failures provided in this application embodiment; Figure 3 A structural block diagram of a device for analyzing stable test case operation failures provided in this application embodiment; Figure 4 This is an internal structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0013] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0014] It should be noted that the terms "S1," "S2," etc., are used only for descriptive purposes and do not specifically refer to the order or sequence, nor are they intended to limit this application. They are merely for the convenience of describing the method of this application and should not be construed as indicating the sequential order of the steps. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0015] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0016] The embodiments of this application provide a method for analyzing the failure of long-term stable test cases, applicable to distributed storage systems. The distributed storage system includes at least one storage node. The method is described in detail below, taking into account the execution flow of the analysis method for long-term stable test case failures.
[0017] S101: Initiate a long-term stability test case to collect node operation data of at least one storage node and business operation data corresponding to the distributed storage system.
[0018] Here, long-term stability test cases are specifically designed to drive the system to run for extended periods (e.g., hours to weeks) under continuous load in a simulated production environment. Their main purpose is to discover resource leaks (memory, connectivity, etc.), performance degradation, and deep-seated, sporadic defects that only trigger during long-term operation or when resources accumulate to a certain level, thereby verifying the system's ability to operate stably and reliably under high load for extended periods.
[0019] Here, distributed storage refers to data being stored across multiple physical or logical storage nodes, with data sharing and management achieved through a network.
[0020] Among them, the analysis method for running failures of long-term stable use cases can be run on a server connected to a distributed storage system.
[0021] Among them, real-time collection of operational data.
[0022] Different long-term stable use cases can correspond to different business operations.
[0023] The node operation data of the storage nodes can include information about the storage nodes themselves, as well as information between different storage nodes, such as CPU utilization, memory utilization, disk I / O read / write speed, network bandwidth utilization, storage transaction processing logs, storage data block verification information, the operating status of each node, the network status between nodes, and the operating status of various services in the storage system.
[0024] Here, the business corresponding to the distributed storage system refers to application scenarios that rely on the distributed storage architecture to realize its functions and services.
[0025] The business operation data may include the initiation time, completion time, operation type, and data transmission volume of the business operation.
[0026] Specifically, the business can include internet content services, social networking and collaboration, e-commerce and transactions, cloud computing platforms, log monitoring, medical imaging, autonomous driving, etc.
[0027] S102: Determine the operational status of the service based on the preset disconnection rules and service operation data.
[0028] Here, the operational status of the service can be normal operation, interrupted operation, or a dangerous state that is between normal operation and interrupted operation.
[0029] The preset disconnection rules can include the logical relationships between multiple indicators.
[0030] These can be achieved through rule engine execution, metric evaluation, logical judgment, or status determination.
[0031] Different types of business can correspond to different operating rules.
[0032] Specifically, when multiple rules are superimposed in the preset disconnection rules, a priority mechanism (such as disconnection, warning, normal) can be adopted: as long as one "disconnection" rule is hit, the overall status is "disconnection"; otherwise, if a "warning" rule is hit, the status is "warning"; if all rules pass, the status is "normal".
[0033] S103: In response to a service interruption, determine the fault type based on preset fault rules and node operation data of at least one storage node.
[0034] Here, "interruption" refers to the interruption of data flow during data transmission, service calls, or user interaction due to network problems, system failures, resource limitations, or other reasons.
[0035] Here, fault types can include hardware faults, network faults, software faults, service faults, etc.
[0036] One or more faults may exist simultaneously. When multiple faults exist, they may be of the same type or of different types.
[0037] S104: Based on the fault type and node operation data of at least one storage node, obtain the causal relationship network corresponding to the fault type.
[0038] Here, a causal relationship network is a graphical model used to represent causal relationships between variables. It uses a directed acyclic graph (DAG) to show the direct or indirect effects between different variables.
[0039] Specifically, causal relationships between different fault types and operational data can be obtained through causal relationship networks.
[0040] In one embodiment, the causal analysis results of each outage (such as "causal network structure, propagation path, and key nodes") are stored in a fault case library, labeled with the fault type (such as "hardware-network coupling fault") and the processing effect (such as "service recovery time after repairing key nodes"). The Bayesian network and GNN model are retrained periodically (e.g., weekly) using the case library data, updating the causal probability calculation parameters (e.g., "after adding 10 cases of service timeout caused by memory leaks, the weight of this causal pair is adjusted from 0.7 to 0.85"), and optimizing the message passing efficiency of the GNN (e.g., shortening the feature update frequency for "low-importance nodes").
[0041] S105: Generate a test report for long-term stable test cases based on the node operation data, business operation data, fault types, and causal relationship network of the storage nodes.
[0042] It should be noted that this application enables efficient analysis and accurate location of storage system faults during the operation of long-term stable test cases, and generates comprehensive and standardized test problem reports, providing a reliable basis for the optimization of the storage system. At the same time, it saves a lot of manpower, significantly improves the efficiency of system problem investigation and repair, and enhances the stability and reliability of the distributed storage system.
[0043] In some specific implementations, the operational status of a service is determined based on preset disconnection rules and service operation data, including: Record the time when the corresponding business operation times out and the business data transmission volume is zero. Statistical analysis of time and determination of whether the time is continuous; if the time is continuous, determination of whether the duration exceeds a time threshold. If the duration exceeds a time threshold, the service is determined to be in a disconnection state.
[0044] Specifically, the business monitoring module analyzes business operation status data in real time according to preset business continuity judgment rules. When X consecutive business operations time out and the business data transmission volume is 0, it is determined that a business interruption has occurred, and the start time of the business interruption is recorded.
[0045] Thus, compared to a single indicator, dual-condition judgment significantly reduces the probability of false alarms and false negatives, and improves the accuracy of monitoring.
[0046] In some specific implementations, the node operation data of the storage node includes hardware data, network data, software data, and service data. Preset fault rules include hardware fault rules, network fault rules, software fault rules, and service fault rules. In response to a service outage, the fault type is determined based on the preset fault rules and the node operation data of at least one storage node, including: Obtain hardware data from at least one storage node, and determine whether there is a hardware fault in the distributed storage system based on the hardware data and hardware fault rules. Obtain network data from at least one storage node, and determine whether a network fault exists in the distributed storage system based on the network data and network fault rules. Obtain software data from at least one storage node, and determine whether there is a storage fault in the distributed storage system based on the software data and software fault rules; Obtain service data from at least one storage node, and determine whether there is a service failure in the distributed storage system based on the service data and service failure rules. Based on the judgment results, at least one fault type corresponding to at least one storage node is determined.
[0047] Here, hardware data may include data such as CPU (Central Processing Unit) utilization, memory utilization, and disk I / O (Input / Output) read / write speed of storage nodes.
[0048] Among them, the hardware fault rule can be a fixed threshold for judgment or a combination of multiple thresholds for judgment.
[0049] Hardware failures can be specifically categorized into CPU hardware failures, memory failures, disk failures, etc.
[0050] Specifically, the hardware data of each storage node is analyzed. If the CPU utilization of a storage node continues to exceed 90% during the service interruption period and is accompanied by frequent process blocking, it is determined that the CPU of the storage node may have a performance bottleneck or overheating failure. If the memory utilization exceeds 85% and there is frequent memory swapping, it is determined that the memory of the storage node is insufficient or leaking. If the disk I / O read / write rate suddenly drops to less than 10% of the normal level and the number of disk error logs increases significantly, it is determined that the disk may have physical damage or read / write queue congestion failure.
[0051] Here, network data can include network bandwidth utilization, packet loss rate, network latency, etc.
[0052] Specifically, by analyzing network bandwidth utilization data, if the network bandwidth utilization between storage system nodes suddenly drops to below 20% of the normal level during a service interruption, and the network latency exceeds three times the normal threshold, while there is a large amount of packet loss, it can be determined that the network link may be interrupted, the network equipment may be malfunctioning, or there may be network congestion.
[0053] For example, network fault rules can be combined with storage transaction processing logs from software fault rules, and with network connection timeout error records from storage transaction processing logs, to further determine the specific location and type of network fault. The specific location of the network fault can be a network link (such as a switch or network interface card) or a network device (router or firewall).
[0054] Here, software data may include stored transaction processing log data and stored data block verification information. The stored transaction processing log may include database transaction rollback records, file system metadata errors, etc.; the stored data block verification information may include verification error rate, distribution of erroneous data blocks, etc.
[0055] Specifically, a deep analysis is performed on the storage transaction processing logs and storage data block verification information. If the logs frequently show records such as database transaction rollbacks and file system metadata errors, it is determined that there may be software vulnerabilities or data consistency issues. If the storage data block verification error rate exceeds 0.1% during the business interruption period, and the distribution of erroneous data blocks has no obvious pattern, it is determined that the storage data may be corrupted or lost.
[0056] Here, service data can include the status of storage high availability services, operation and maintenance monitoring services, OSS (Object Storage Service) services, MDS (Metadata Service) services, etc.
[0057] Specifically, if one or more services are in an abnormal state, and it is determined that the service interruption is caused by the storage service itself, the logs for the corresponding storage service failure period are recorded, and the R&D personnel further determine in the report whether it is a software-level problem.
[0058] In one embodiment, it may also include determining whether the fault is caused by uncontrollable factors. For internal storage faults, such faults are caused by controllable factors, including faults in various storage services and other faults caused by them. For such faults, R&D personnel need to analyze the causes in detail based on the report and fix the potential faults in the storage. For non-internal storage faults, such faults are caused by uncontrollable factors, such as network faults caused by switch problems, node faults caused by power outages, etc. For service interruptions caused by such faults, it is necessary to analyze the interruption duration and determine whether the current interruption time is a normal service interruption based on the interruption duration specifications of various faults in distributed storage. That is, it supports the accurate distinction between expected and unexpected service interruptions.
[0059] In this way, intelligent fault identification can be achieved. By conducting a comprehensive analysis at different levels, the specific cause of the fault can be located more accurately, reducing misjudgments caused by the interplay of multiple factors.
[0060] In some specific implementations, the preset fault rules include fault thresholds, and the method further includes: Obtain business load data, which includes time, concurrency, and outage tags; Discretize the time and concurrency to obtain discretized time and discretized concurrency; The fault threshold, disconnection label, discretized time, and discretized concurrency are used as input parameters and fed into a reinforcement learning model trained based on historical business load data to obtain the adjusted fault threshold.
[0061] Here, discretization is the process of converting continuous data or signals into a discrete form.
[0062] Among them, time discretization can be used to divide a continuous time axis into windows of fixed size (such as one window per minute or one window per hour), and aggregate the data in each window (such as calculating the average or maximum value).
[0063] Among them, concurrency discretization divides continuous metric values (such as CPU utilization) into several intervals (for example, 0-30% is low load, 30%-70% is normal, and above 70% is high load).
[0064] Specifically, during high-load periods (such as 9:00-18:00 during the day), the fault threshold is increased (allowing more temporary timeouts) to reduce false alarms; during low-load periods (such as early morning), the fault threshold is decreased to improve the sensitivity of current outage detection.
[0065] In this way, the limitations of static rules are broken, enabling the interruption detection to have scene awareness capabilities and improving the detection accuracy under complex loads.
[0066] In some specific implementations, a causal relationship network corresponding to the fault type is obtained based on the fault type and node operation data of at least one storage node, including: Obtain historical node operation data and real-time node operation data of at least one storage node, align the historical node operation data and real-time node operation data with the time axis, and slice them according to the preset time length to obtain time slices; Extract features from time slices, including trend features and mutation features; By using a sliding time window, the conditional probability between the fault type and the corresponding historical node operation data is calculated based on the characteristics of the historical node operation data. The initial causal relationship network is constructed by taking the types of faults that exist in the historical operation as nodes, the causal dependencies between fault types as edges, and the conditional probabilities between fault types and the corresponding historical node operation data as the initial weights of the edges. By using a sliding time window, the conditional probability between the fault type and the corresponding real-time node operation data is calculated based on the characteristics of the real-time node operation data. Based on the conditional probabilities between fault types and corresponding real-time node operation data, and between fault types and corresponding historical node operation data, the weights of the edges in the initial causal relationship network are updated to obtain the causal relationship network.
[0067] Here, trend characteristics refer to the overall direction, rate, and stability of data changes within a time slice, reflecting a relatively slow and continuous evolution pattern of data during that period.
[0068] Trend characteristics can be calculated using methods such as slope trend, mean, variance, standard value, and autocorrelation.
[0069] Specifically, trend characteristics can include the slope change of CPU utilization, the variance of network latency, etc.
[0070] Here, abrupt changes are sudden, drastic, and brief changes (peaks, drops, breaks) that occur in data within a time slice, reflecting outliers or state transitions in the data during that time period.
[0071] The abrupt change trend can be obtained through methods such as the absolute value of the first-order difference, the absolute value of the second-order difference, the extreme value, the peak value, the proportion exceeding the threshold, and the quantile difference.
[0072] Specifically, abrupt change characteristics can be, for example, the absolute value of the first-order difference in disk I / O rates.
[0073] In one embodiment, multiple trend features and multiple mutation features of a time slice can be calculated to form a feature vector.
[0074] Here, the edge is a directed edge pointing from node A to node B, indicating that the occurrence of fault type A may directly lead to the occurrence of fault type B.
[0075] Here, the weight of an edge represents the probability that fault B will occur within a specific time window given that fault A has occurred.
[0076] The conditional probability between fault type and historical node operation data can be obtained by counting the total number of times fault A occurred in the historical node operation data and the number of times fault B occurred within Δt after fault A occurred, thus obtaining the initial weight. For example, if network latency occurred 100 times in the historical data, and 80 of them caused service timeouts within 3 seconds, the corresponding edge weight W(network latency → service timeout) = 80 / 100 = 0.8.
[0077] In one embodiment, memory overload (A) and network congestion (C) are observed to occur together in the real-time window, and service interruption (B) occurs 2 seconds later. The joint conditional probability is calculated as: P(B|A∩C) = N(A∩C→B) / N(A∩C), and the edge weight W(Memory overload ∩ Network congestion → Service interruption) is updated.
[0078] Specifically, using causal analysis algorithms, correlation analysis is performed on the various faults obtained from the above analysis. First, the collected multi-dimensional data (such as hardware indicators and network status) is aligned on the time axis and features are extracted. The continuously running data is sliced into millisecond-level time slices (such as generating a data snapshot every 100ms). The time dimension of each indicator is unified (such as mapping CPU utilization, network latency, and service response time to the same time axis) to solve the causal misalignment problem caused by "inconsistent collection frequency of different indicators". For each time slice, trend features and mutation features are extracted. For example, if the CPU utilization of a node increases from 60% to 90% at a rate of 0.5% / s in the 5 minutes before the network outage, this "slope feature" can serve as a key input for subsequent causal analysis. Based on this, a fault causal relationship network can be constructed, with fault type (hardware, network, software, service) as nodes and the "probability of fault occurrence order" in historical fault data as the initial edge weights (e.g., "the probability of service timeout is 80% within 3 seconds after network latency exceeds the threshold", then the initial weight of "network latency → service timeout" is 0.8). For the time-series data during the network outage, the conditional probability between nodes is calculated through a "sliding time window" (e.g., window size of 5 seconds).
[0079] In one embodiment, when a fault type that is not present in the causal relationship network occurs, the analysis parameters of similar cases are reused through transfer learning, such as the propagation path characteristics of data consistency faults, to quickly learn preliminary causal relationships and add the new fault type and corresponding causal relationship to the causal relationship network.
[0080] In one embodiment, for a sliding window of real-time node running data, the conditional probability between the fault type and the corresponding real-time node running data within the sliding window is calculated to obtain the current weight. The difference between the current weight and the historical weight is calculated to obtain the difference value. The number of samples within the sliding window of real-time node running data is obtained. The confidence factor is calculated based on the number of samples. A preset basic smoothing factor α is obtained. The difference-sensitive part is calculated based on the difference value and the sensitivity coefficient. The effective smoothing factor is calculated based on the confidence factor, the basic smoothing factor, and the difference-sensitive part. The weight of the edge is updated using the effective smoothing factor, the conditional probability, and the current weight.
[0081] For example, suppose the current weight is W1, the updated weight is W2, the historical conditional probability is P0, the conditional probability is P1, the difference is diff, the base smoothing factor is α, the sample size is n, the confidence factor is β(n), the sensitivity coefficient is k, and the effective smoothing factor is α. e In the initial update, diff = |P1 - P0|, and in subsequent updates, diff = |P1 – W1|, β(n) = min(1, n / N), where N is a preset threshold for the number of samples, and α e =min(α m ,α*β(n)+k*diff), where α m As a preset value, W2 = (1 - α) e )* W1+α e *P1.
[0082] Specifically, when the amount of real-time data is small, β(n) is small, therefore α e The value is small, with small update increments, mainly relying on historical weights; when the amount of real-time data is large, β(n) is large, and the base component becomes larger; when the difference between the real-time probability and the current weight is large, α... e It grows larger, the update range increases, and it adapts to changes more quickly; at the same time, through α m The maximum update size is limited to avoid excessively large single updates. This ensures that weight updates are both responsive to rapid changes and maintain stability.
[0083] In this way, timeline alignment can solve the problem of inconsistent data collection frequencies and unify the time dimension of each data point; it can also automate root cause analysis and reduce manual investigation.
[0084] In some specific implementations, the method further includes: In response to multiple concurrent fault types, the marginal contribution of each fault type is calculated using a Bayesian network. Adjust the weights of the edges in the causal network based on the marginal contribution.
[0085] Here, Bayesian networks are directed acyclic graphs used to represent probabilistic dependencies between variables. Combining graph theory and probability theory, they provide a powerful framework for handling uncertain information.
[0086] Here, marginal contribution is used to measure the incremental contribution of an individual, factor, or variable to the whole.
[0087] Specifically, for concurrent scenarios involving CPU overload, network congestion, and service anomalies, the marginal contribution of each factor is calculated using the "belief propagation algorithm" of Bayesian networks (e.g., "CPU overload contributes 40% to the outage, network congestion contributes 35%, and service anomalies contribute 25%").
[0088] In this way, quantitative attribution analysis can be achieved under multiple concurrent causes, solving the problem that a single rule cannot distinguish between primary and secondary factors.
[0089] In some specific implementations, test reports for long-term stable test cases are generated based on the node operation data, service operation data, fault types, and causal relationship networks of the storage nodes, including: Based on the node operation data and causal relationship network of the storage nodes, node failure analysis data is recorded. The node failure analysis data includes the faulty storage node, the time of failure, and the scope of failure impact. Based on business operation data, record business interruption analysis data, which includes the start time, duration and scope of the interruption. Based on the fault type, corresponding fault repair suggestions are matched, with different fault types corresponding to different repair suggestion libraries; Fill in the node fault analysis data, service interruption analysis data, and fault repair suggestions into the test report template to generate a test report.
[0090] Specifically, the report includes: an overview of the service outage, detailing the start time, duration, and affected service scope; a detailed analysis of the storage system faults, classifying and describing various detected faults, including fault type, fault node, fault occurrence time, and assessment of the fault's impact; fault repair recommendations, proposing specific repair measures and optimization schemes for different fault types, and for faults that are merely storage service-related, proactively restarting the service after clearly recording the fault cause and error logs, and continuing to run long-term stable use cases; and related operational data charts, visually demonstrating the changes in the storage system's operational status before and after the service outage.
[0091] Thus, the test problem reports generated according to the preset template cover key information on business interruptions and storage system failures, providing a comprehensive and intuitive basis for the optimization and improvement of the storage system, and helping technical personnel to quickly understand the problems and develop solutions.
[0092] In some specific implementations, the method further includes: In response to the fact that the business is not interrupted, the node operation data of the storage nodes within the preset time period is cleaned. Based on the trend prediction model, the node operation data of the cleaned storage nodes is used to predict the trend and determine whether a pre-disruption warning is triggered.
[0093] Data cleaning can include operations such as handling missing values, smoothing noise, and normalization.
[0094] This includes acquiring historical business operation data under normal operating conditions, statistically analyzing the normal fluctuation range of business operation data, and constructing a normal operating baseline model.
[0095] The trend prediction model can be tailored to different time periods. For example, for short-term predictions (less than 1 hour), Holt-Winters exponential smoothing can be used, while for medium-term predictions (1-4 hours), Prophet or LSTM can be employed.
[0096] The logic for triggering a pre-disconnection warning can include: judging whether the threshold is approaching, for example, the predicted value is greater than 80% of the disconnection threshold; judging the trend slope, for example, the upward slope in the past 10 minutes is greater than twice the historical average; deteriorating through the coordinated action of multiple indicators, for example, CPU is greater than 85% and disk write speed is greater than 90% threshold; and I / O wait continues to increase, for example, %iowait is greater than 30% and lasts for more than 5 minutes.
[0097] Specifically, based on time series analysis, the system predicts the trend of the operating data (such as IOPS, bandwidth, memory usage, etc.) over the past hour. When the indicators deviate from the normal range but do not reach the interruption threshold, a "pre-interruption warning" is triggered, and potential risk warnings are pushed in advance (such as "interruption may occur in the next 10 minutes, the reason is related to the continuous increase in memory usage of node A").
[0098] In one embodiment, a causal relationship network can be used to provide a power outage warning.
[0099] In this way, proactive early warning of faults can be achieved, reducing the impact of power outages.
[0100] In one embodiment, Figure 2 This is a schematic diagram of the process for generating test reports in the embodiments of this application, such as... Figure 2As shown, the test report generation process in this application includes: running stable test cases on the test server, configuring the client address through the configuration file, and executing the script; starting the stable test case detection script to monitor the running status of stable test cases and the cluster status in real time; determining whether there is a flow interruption phenomenon in the execution of stable test cases; if not, outputting a test report; if so, performing fault analysis and localization; determining whether the fault is caused by uncontrollable factors; if not, analyzing the faults of each storage service and other faults caused by them, and outputting a test report; if so, analyzing the flow interruption duration, determining whether the current flow interruption time is a normal flow interruption service according to the flow interruption duration rules of each fault in distributed storage, and outputting a test report.
[0101] In one embodiment, the method includes the following modules: a data acquisition and monitoring module, used to initiate continuous operation of long-term stable storage test cases and collect real-time storage system operation data and business operation status data; a business interruption detection module, used to determine the business interruption situation according to preset rules and trigger the fault analysis process; a fault analysis and location module, used to analyze and locate the faults of the storage system during the business interruption period, determine the fault type, fault node and fault causal relationship, accumulate historical fault cases, and identify new faults; and a test problem report generation module, used to generate a test problem report based on the fault analysis results.
[0102] It should be understood that, although Figure 1 and 2 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 and 2 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0103] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0104] Embodiments of this application also provide a historical alarm merging device, applicable to a distributed storage system. The distributed storage system includes at least one storage node. The device includes: a first processing module 301, used to initiate long-term stability test cases and collect node operation data of at least one storage node and corresponding business operation data of the distributed storage system; a second processing module 302, used to determine the operation status of the business based on preset disconnection rules and business operation data; a third processing module 303, used to determine the fault type based on preset fault rules and node operation data of at least one storage node in response to the business being in a disconnection state; a fourth processing module 304, used to obtain the causal relationship network corresponding to the fault type based on the fault type and node operation data of at least one storage node; and a fifth processing module 305, used to generate a test report of the long-term stability test cases based on the node operation data of the storage node, business operation data, fault type, and causal relationship network.
[0105] As a preferred implementation, in this embodiment of the application, the second processing module 302 is specifically used to: record the time when the business operation corresponding to the business times out and the business data transmission volume is zero; count the time and determine whether the time is continuous; in response to the time being continuous, determine whether the duration is greater than a time threshold; in response to the duration being greater than the time threshold, determine that the business operation status is a disconnection state.
[0106] In a preferred implementation, in this embodiment, the node operation data of the storage node includes hardware data, network data, software data, and service data. The preset fault rules include hardware fault rules, network fault rules, software fault rules, and service fault rules. The third processing module 303 is specifically used for: acquiring hardware data of at least one storage node; determining whether a hardware fault exists in the distributed storage system based on the hardware data and hardware fault rules; acquiring network data of at least one storage node; determining whether a network fault exists in the distributed storage system based on the network data and network fault rules; acquiring software data of at least one storage node; determining whether a storage fault exists in the distributed storage system based on the software data and software fault rules; acquiring service data of at least one storage node; determining whether a service fault exists in the distributed storage system based on the service data and service fault rules; and determining at least one fault type corresponding to at least one storage node based on the determination result.
[0107] In a preferred embodiment of this application, the device further includes an adjustment module, which is specifically used to: acquire service load data, including time, concurrency, and disconnection label; discretize the time and concurrency to obtain discretized time and discretized concurrency; and input the fault threshold, disconnection label, discretized time, and discretized concurrency as input parameters into a reinforcement learning model trained based on historical service load data to obtain the adjusted fault threshold.
[0108] In a preferred embodiment of this application, the fourth processing module 304 is specifically configured to: acquire historical node operation data and real-time node operation data of at least one storage node; align the historical node operation data and real-time node operation data along a time axis; slice the data according to a preset time length to obtain time slices; extract features from the time slices, including trend features and mutation features; calculate the conditional probability between fault types and corresponding historical node operation data based on the features of the historical node operation data using a sliding time window; construct an initial causal relationship network by using fault types existing in the historical operation process as nodes, causal dependencies between fault types as edges, and the conditional probability between fault types and corresponding historical node operation data as the initial weights of the edges; calculate the conditional probability between fault types and corresponding real-time node operation data based on the features of the real-time node operation data using a sliding time window; and update the weights of the edges of the initial causal relationship network based on the conditional probabilities between fault types and corresponding real-time node operation data and between fault types and corresponding historical node operation data to obtain the causal relationship network.
[0109] As a preferred implementation, in this embodiment of the application, the fourth processing module 304 is further configured to: in response to multiple fault types occurring concurrently, calculate the marginal contribution of multiple fault types through a Bayesian network; and adjust the weights of the edges of the causal relationship network based on the marginal contribution.
[0110] In a preferred implementation, in this embodiment, the fifth processing module 305 is specifically used to: record node fault analysis data based on the node operation data and causal relationship network of the storage node, wherein the node fault analysis data includes the faulty storage node, the fault occurrence time, and the fault impact range; record service interruption analysis data based on the service operation data, wherein the service interruption analysis data includes the interruption start time, duration, and impact range; match corresponding fault repair suggestions based on the fault type, wherein different fault types correspond to different repair suggestion libraries; and fill the node fault analysis data, service interruption analysis data, and fault repair suggestions into the test report template to generate a test report.
[0111] For a description of the features in the embodiment corresponding to the analysis device for long-term stable test case operation failure, please refer to the relevant description of the embodiment corresponding to the analysis method for long-term stable test case operation failure, which will not be repeated here.
[0112] Embodiments of this application also provide an electronic device, which may be a terminal, and its internal structure diagram may be as follows: Figure 3 As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for analyzing long-term stable use case operation failures. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.
[0113] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. Specifically, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0114] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it performs the following steps: S1: Initiating a long-term stability test case, collecting node operation data of at least one storage node and business operation data corresponding to the distributed storage system; S2: Determining the operating status of the business based on preset disconnection rules and business operation data; S3: In response to the business being in a disconnection state, determining the fault type based on preset fault rules and node operation data of at least one storage node; S4: Obtaining the causal relationship network corresponding to the fault type based on the fault type and node operation data of at least one storage node; S5: Generating a test report for the long-term stability test case based on the node operation data of the storage nodes, business operation data, fault type, and causal relationship network.
[0115] In one embodiment, when the processor executes the computer program, it further performs the following steps: recording the time when the business operation corresponding to the business times out and the business data transmission volume is zero; calculating the time and determining whether the time is continuous; in response to the time being continuous, determining whether the duration is greater than a time threshold; in response to the duration being greater than the time threshold, determining that the business operation state is a disconnection state.
[0116] In one embodiment, when the processor executes the computer program, it further performs the following steps: acquiring hardware data of at least one storage node, and determining whether a hardware fault exists in the distributed storage system based on the hardware data and hardware fault rules; acquiring network data of at least one storage node, and determining whether a network fault exists in the distributed storage system based on the network data and network fault rules; acquiring software data of at least one storage node, and determining whether a storage fault exists in the distributed storage system based on the software data and software fault rules; acquiring service data of at least one storage node, and determining whether a service fault exists in the distributed storage system based on the service data and service fault rules; and determining at least one fault type corresponding to at least one storage node based on the determination result.
[0117] In one embodiment, when the processor executes the computer program, it further performs the following steps: acquiring business load data, which includes time, concurrency, and outage label; discretizing the time and concurrency to obtain discretized time and discretized concurrency; and inputting the fault threshold, outage label, discretized time, and discretized concurrency as input parameters into a reinforcement learning model trained based on historical business load data to obtain an adjusted fault threshold.
[0118] In one embodiment, when the processor executes the computer program, it further performs the following steps: acquiring historical node operation data and real-time node operation data of at least one storage node; aligning the historical node operation data and real-time node operation data along a time axis; slicing the data according to a preset time length to obtain time slices; extracting features from the time slices, wherein the features include trend features and abrupt change features; calculating the conditional probability between a fault type and the corresponding historical node operation data based on the features of the historical node operation data using a sliding time window; constructing an initial causal relationship network by using fault types existing in the historical operation process as nodes, causal dependencies between fault types as edges, and the conditional probability between a fault type and the corresponding historical node operation data as the initial weights of the edges; calculating the conditional probability between a fault type and the corresponding real-time node operation data based on the features of the real-time node operation data using a sliding time window; and updating the weights of the edges of the initial causal relationship network based on the conditional probabilities between the fault type and the corresponding real-time node operation data and the conditional probabilities between the fault type and the corresponding historical node operation data to obtain the causal relationship network.
[0119] In one embodiment, when the processor executes the computer program, it further performs the following steps: in response to multiple concurrent fault types, calculating the marginal contribution of multiple fault types through a Bayesian network; and adjusting the weights of the edges of the causal relationship network based on the marginal contribution.
[0120] In one embodiment, when the processor executes the computer program, it further performs the following steps: recording node fault analysis data based on the node operation data and causal relationship network of the storage node, wherein the node fault analysis data includes the faulty storage node, the fault occurrence time, and the fault impact range; recording service interruption analysis data based on the service operation data, wherein the service interruption analysis data includes the interruption start time, duration, and impact range; matching corresponding fault repair suggestions based on the fault type, wherein different fault types correspond to different repair suggestion libraries; and filling the node fault analysis data, service interruption analysis data, and fault repair suggestions into a test report template to generate a test report.
[0121] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it performs the following steps: S1: Initiating a long-term stability test case, collecting node operation data of at least one storage node and business operation data corresponding to the distributed storage system; S2: Determining the operation status of the business based on preset disconnection rules and business operation data; S3: In response to the business being in a disconnection state, determining the fault type based on preset fault rules and node operation data of at least one storage node; S4: Obtaining the causal relationship network corresponding to the fault type based on the fault type and node operation data of at least one storage node; S5: Generating a test report for the long-term stability test case based on the node operation data of the storage nodes, business operation data, fault type, and causal relationship network.
[0122] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: recording the time when the business operation corresponding to the business times out and the business data transmission volume is zero; calculating the time and determining whether the time is continuous; in response to the time being continuous, determining whether the duration is greater than a time threshold; in response to the duration being greater than the time threshold, determining that the business operation state is a disconnection state.
[0123] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: acquiring hardware data of at least one storage node, and determining whether a hardware fault exists in the distributed storage system based on the hardware data and hardware fault rules; acquiring network data of at least one storage node, and determining whether a network fault exists in the distributed storage system based on the network data and network fault rules; acquiring software data of at least one storage node, and determining whether a storage fault exists in the distributed storage system based on the software data and software fault rules; acquiring service data of at least one storage node, and determining whether a service fault exists in the distributed storage system based on the service data and service fault rules; and determining at least one fault type corresponding to at least one storage node based on the determination result.
[0124] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: acquiring business load data, which includes time, concurrency, and outage label; discretizing the time and concurrency to obtain discretized time and discretized concurrency; and inputting the fault threshold, outage label, discretized time, and discretized concurrency as input parameters into a reinforcement learning model trained based on historical business load data to obtain an adjusted fault threshold.
[0125] In one embodiment, when the computer program is executed by a processor, it further performs the following steps: acquiring historical node operation data and real-time node operation data of at least one storage node; aligning the historical node operation data and real-time node operation data along a time axis; slicing the data according to a preset time length to obtain time slices; extracting features from the time slices, wherein the features include trend features and abrupt change features; calculating the conditional probability between a fault type and the corresponding historical node operation data based on the features of the historical node operation data using a sliding time window; constructing an initial causal relationship network by using fault types existing in the historical operation process as nodes, causal dependencies between fault types as edges, and the conditional probability between a fault type and the corresponding historical node operation data as the initial weights of the edges; calculating the conditional probability between a fault type and the corresponding real-time node operation data based on the features of the real-time node operation data using a sliding time window; and updating the weights of the edges of the initial causal relationship network based on the conditional probabilities between the fault type and the corresponding real-time node operation data and the conditional probabilities between the fault type and the corresponding historical node operation data to obtain the causal relationship network.
[0126] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: in response to the concurrent occurrence of multiple fault types, calculating the marginal contribution of multiple fault types through a Bayesian network; and adjusting the weights of the edges of the causal relationship network based on the marginal contribution.
[0127] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: recording node fault analysis data based on the node operation data and causal relationship network of the storage node, wherein the node fault analysis data includes the faulty storage node, the fault occurrence time, and the fault impact range; recording service outage analysis data based on the service operation data, wherein the service outage analysis data includes the outage start time, duration, and impact range; matching corresponding fault repair suggestions based on the fault type, wherein different fault types correspond to different repair suggestion libraries; and filling the node fault analysis data, service outage analysis data, and fault repair suggestions into a test report template to generate a test report.
[0128] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0129] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0130] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application.
Claims
1. A method for analyzing failures in long-term stable test cases, characterized in that, Applicable to a distributed storage system, the distributed storage system including at least one storage node, the method includes: Initiate a long-term stable use case to collect node operation data of at least one storage node and business operation data corresponding to the distributed storage system; The operating status of the service is determined based on the preset disconnection rules and the service operation data; In response to the service being interrupted, the fault type is determined based on preset fault rules and the node operation data of the at least one storage node; Based on the fault type and the node operation data of the at least one storage node, a causal relationship network corresponding to the fault type is obtained; A test report for the long-term stable use case is generated based on the node operation data of the storage node, the business operation data, the fault type, and the causal relationship network.
2. The method for analyzing operational failures of long-term stable test cases according to claim 1, characterized in that, The step of determining the operating status of the service based on the preset disconnection rules and the service operation data includes: Record the time when the business operation corresponding to the aforementioned business times out and the business data transmission volume is zero. The time is statistically analyzed and it is determined whether the time is continuous. In response to the fact that the time is continuous, it is determined whether the duration is greater than a time threshold. If the duration exceeds the time threshold, the service is determined to be in a disconnection state.
3. The method for analyzing operational failures of long-term stable test cases according to claim 1, characterized in that, The node operation data of the storage node includes hardware data, network data, software data and service data, and the preset fault rules include hardware fault rules, network fault rules, software fault rules and service fault rules. In response to the service being interrupted, the fault type is determined based on preset fault rules and node operation data of the at least one storage node, including: Obtain hardware data of at least one storage node, and determine whether there is a hardware fault in the distributed storage system based on the hardware data and hardware fault rules; Obtain network data from at least one storage node, and determine whether a network fault exists in the distributed storage system based on the network data and network fault rules; Obtain software data from at least one storage node, and determine whether the distributed storage system has a storage fault based on the software data and software fault rules; Obtain service data from at least one storage node, and determine whether there is a service failure in the distributed storage system based on the service data and service failure rules; Based on the judgment result, at least one fault type corresponding to the at least one storage node is determined.
4. The method for analyzing operational failures of long-term stable test cases according to claim 1, characterized in that, The preset fault rules include fault thresholds, and the method further includes: Obtain business load data, which includes time, concurrency, and outage tags; Discretize the time and the concurrency to obtain discretized time and discretized concurrency; The fault threshold, the disconnection label, the discretized time, and the discretized concurrency are used as input parameters and input into a reinforcement learning model trained based on historical business load data to obtain the adjusted fault threshold.
5. The method for analyzing operational failures of long-term stable test cases according to claim 1, characterized in that, The step of obtaining the causal relationship network corresponding to the fault type based on the fault type and the node operation data of the at least one storage node includes: Obtain historical node operation data and real-time node operation data of the at least one storage node, align the historical node operation data and the real-time node operation data with the time axis, and slice them according to a preset time length to obtain a time slice; Extract features from the time slice, wherein the features include trend features and abrupt change features; By using a sliding time window, the conditional probability between the fault type and the corresponding historical node operation data is calculated based on the characteristics of the historical node operation data. The fault types existing in the historical operation process are used as nodes, the causal dependencies between fault types are used as edges, and the conditional probabilities between the fault types and the corresponding historical node operation data are used as the initial weights of the edges to construct an initial causal relationship network. By using a sliding time window, the conditional probability between the fault type and the corresponding real-time node operation data is calculated based on the characteristics of the real-time node operation data. Based on the conditional probability between the fault type and the corresponding real-time node operation data and the conditional probability between the fault type and the corresponding historical node operation data, the weights of the edges in the initial causal relationship network are updated to obtain the causal relationship network.
6. The method for analyzing operational failures of long-term stable test cases according to claim 5, characterized in that, The method further includes: In response to multiple concurrent fault types, the marginal contribution of each fault type is calculated using a Bayesian network. The weights of the edges in the causal network are adjusted based on the marginal contribution.
7. The method for analyzing operational failures of long-term stable test cases according to claim 1, characterized in that, The step of generating a test report for the long-term stable use case based on the node operation data of the storage node, the service operation data, the fault type, and the causal relationship network includes: Based on the node operation data of the storage node and the causal relationship network, node fault analysis data is recorded, including the faulty storage node, the time of fault occurrence, and the scope of fault impact. Based on the business operation data, record business interruption analysis data, wherein the business interruption analysis data includes the start time, duration and scope of the interruption; Based on the fault type, a corresponding fault repair suggestion is matched, where different fault types correspond to different repair suggestion libraries; Fill the node fault analysis data, the service outage analysis data, and the fault repair suggestions into the test report template to generate a test report.
8. An analysis device for long-term stable test case operation failure, characterized in that, Suitable for distributed storage systems, the distributed storage system including at least one storage node, the device includes: The first processing module is used to initiate long-term stable test cases and collect node operation data of at least one storage node and business operation data corresponding to the distributed storage system. The second processing module is used to determine the operating status of the service based on the preset disconnection rules and the service operation data; The third processing module is used to determine the fault type in response to the service being in a disconnection state, based on preset fault rules and node operation data of the at least one storage node. The fourth processing module is used to obtain the causal relationship network corresponding to the fault type based on the fault type and the node operation data of the at least one storage node; The fifth processing module is used to generate a test report for the long-term stable use case based on the node operation data of the storage node, the business operation data, the fault type, and the causal relationship network.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program to implement the steps of the analysis method for long-term stable use case operation failures as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the method for analyzing operational failures of long-term stable use cases as claimed in any one of claims 1 to 7.