Database fault processing method and system based on reinforcement learning
By building a causal decision tree and propagation link based on reinforcement learning, identifying the root cause of database failures and optimizing the processing strategy, the problems of inaccurate fault identification and slow recovery in the existing technology are solved, and efficient failure recovery and system reliability are achieved.
Patent Information
- Application Number
- CN202510838856.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-23
AI Technical Summary
When existing database systems deal with the failure of complex, dynamic, and multi-source causal links, it is difficult to accurately identify the root cause of the failure, slow response speed, lack dynamic optimization capabilities, resulting in a long system recovery cycle and high cost.
Using a reinforcement learning-based method, a causal decision tree is built by collecting performance indicators, execution plans and error log data, a cause and effect decision tree is built, the root cause of the failure is identified and the propagation link is generated, the fault processing priority list is generated, the reinforcement learning action sequence is performed, and the processing strategy is adjusted in real time until the failure is recovered.
Accurate location and adaptive optimization of database failures are realized, fault recovery efficiency and system reliability are improved, and manual intervention and recovery time are reduced.
Smart Images

Figure CN120371675A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to database management technology, and in particular to a database fault handling method and system based on reinforcement learning. Background Art
[0002] During the operation of existing database systems, with the continuous improvement of business complexity and data processing requirements, the frequency and types of system failures have become increasingly diverse. Problems such as deadlocks, resource bottlenecks, index invalidation, and SQL execution exceptions occur frequently. Traditional database fault handling methods rely on static rules or manual experience, making it difficult to identify the root cause of faults in a timely and accurate manner, resulting in a long system recovery period and high costs, seriously affecting business continuity and service quality.
[0003] Although some studies have attempted to introduce technologies such as log analysis, fault diagnosis graphs, or causal reasoning to enhance the automated fault diagnosis ability of the system, they still have problems such as low accuracy, slow response speed, and difficulty in dynamically adjusting the processing priority in scenarios involving complex, dynamic, and multi-source causal links. At the same time, such solutions generally lack the ability to provide real-time feedback and adaptive optimization of the action effects during the fault handling process, and are difficult to meet the intelligent maintenance requirements of modern database systems.
[0004] In recent years, reinforcement learning technology has shown significant advantages in the fields of decision optimization and dynamic feedback. However, its application in database fault handling is still in its initial stage, and there is no effective systematic method that combines causal chain identification, action selection, and performance feedback regulation. Therefore, there is an urgent need for a database fault handling method and system based on reinforcement learning to achieve intelligent closed-loop optimization from root cause identification to fault action execution. Summary of the Invention
[0005] Embodiments of the present invention provide a database fault handling method and system based on reinforcement learning, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention, a database fault handling method based on reinforcement learning is provided, including: Collect performance metric data, execution plan data, and error log data at the moment of database failure, generate a fault state vector, calculate a time series correlation coefficient matrix based on the fault state vector, construct a causal decision tree according to the time series correlation coefficient matrix, identify the fault root cause node from the causal decision tree, and generate a propagation link from the fault root cause node to the performance anomaly node; Generate a fault handling priority list based on the correlation coefficient magnitude and resource occupancy degree of the nodes in the propagation link, and construct a reinforcement learning action sequence according to the fault handling priority list; Execute the reinforcement learning action sequence, and calculate in real time the degree of influence of the executed actions on the node states in the propagation link. According to the degree of influence, divide the actions into a master action set and an adjustment action set according to a preset influence threshold; Collect performance monitoring metric data to calculate the performance improvement ratio, adjust the execution order of the master action set and the execution parameters of the adjustment action set according to the performance improvement ratio, generate an optimized reinforcement learning action sequence, and continue to execute the optimized reinforcement learning action sequence until the fault recovery is completed.
[0007] In an alternative embodiment, Collect performance metric data, execution plan data, and error log data at the moment of database failure, generate a fault state vector, calculate a time series correlation coefficient matrix based on the fault state vector, and construct a causal decision tree according to the time series correlation coefficient matrix, including: Collect performance metric data, execution plan data, and error log data at the moment of database failure; Perform periodic sampling on the performance metric data to obtain the resource usage, query response time, and data processing volume in each sampling period; extract the access path, parallelism, and resource prediction value of the query from the execution plan data; extract the fault type identifier and associated object identifier from the error log data; Calculate the deviation degree between the resource usage and the corresponding predicted value, calculate the delay degree between the query response time and the reference time, and perform an association mapping between the fault type identifier and the associated object identifier; generate a fault state vector based on the deviation degree, delay degree, and association mapping result; Perform a sliding window process on the fault state vector, calculate the time series correlation coefficient between each index in the window, and construct a time series correlation coefficient matrix according to the index correspondence relationship of the time series correlation coefficient; weight the correlation coefficients in the time series correlation coefficient matrix according to the resource prediction value, and construct a causal decision tree based on the weighted time series correlation coefficient matrix, where the index pair with the strongest correlation coefficient is determined as the root node of the tree, the child nodes are constructed in turn according to the correlation coefficient strength, and the resource prediction value is used as the constraint condition for the connection between nodes.
[0008] In an alternative embodiment, Identify the fault root cause node from the causal decision tree, and generate a propagation link from the fault root cause node to the performance anomaly node, including: Calculate the fault influence degree of each node in the causal decision tree, and perform a combined operation on the out-degree value, correlation coefficient strength, and resource occupancy rate of the node to obtain the fault influence degree value of the node; Starting from the leaf nodes of the causal decision tree, traverse backward, calculate the influence transfer value of each node based on the connection strength and fault impact degree value between nodes, and mark the nodes where the influence transfer value mutates as candidate root cause nodes; Taking the candidate root cause node as the starting point, calculate the abnormality value by propagating downward along the causal decision tree based on the fault impact degree, and compare the abnormality value with the abnormal value in the performance index data. When the comparison result meets the matching condition, confirm the current candidate root cause node as the fault root cause node; Starting from the fault root cause node, adopt a depth - first search strategy, determine the propagation direction according to the correlation coefficient strength between nodes, and prune the currently obtained propagation link when the temporal correlation coefficient between nodes is lower than the preset correlation threshold to obtain an initial propagation link; Verify the initial propagation link, confirm the link validity based on the performance index change trend and temporal relationship between nodes, and output the fault propagation link from the fault root cause node to the performance abnormal node.
[0009] In an alternative embodiment, Starting from the fault root cause node, adopting a depth - first search strategy, determining the propagation direction according to the correlation coefficient strength between nodes, and pruning the currently obtained propagation link when the temporal correlation coefficient between nodes is lower than the preset correlation threshold to obtain an initial propagation link includes: Obtain the performance index time - series data of the fault root cause node and its connected nodes, analyze the performance index time - series data using a sliding time window, and calculate the temporal correlation coefficient between adjacent nodes; Determine the propagation direction between nodes based on the magnitude of the temporal correlation coefficient, record all propagation directions of the fault root cause node, and set a preset correlation threshold for pruning the propagation link; Starting from the fault root cause node, traverse the nodes using a depth - first search strategy, preferentially select the propagation direction with the largest temporal correlation coefficient, sequentially add adjacent nodes with a temporal correlation coefficient greater than the preset correlation threshold to the current propagation link, and use the newly added node as the current search node to continue the search until the temporal correlation coefficients of all adjacent nodes of the current search node are lower than the preset correlation threshold; Trace back from the current search node to the previous search nodes in sequence. For the nodes among the previous search nodes that have un - searched propagation directions, select the propagation direction with the second - largest temporal correlation coefficient to continue the depth - first search. After completing the search of all propagation directions, merge the retained propagation links to obtain the initial propagation link.
[0010] In an alternative embodiment, Generate a fault handling priority list based on the correlation coefficient magnitude and resource occupancy degree of nodes in the propagation link. Constructing a reinforcement learning action sequence according to the fault handling priority list includes: Obtain the correlation coefficients between nodes and the resource occupancy data of each node in the fault propagation link; Calculate the correlation coefficient weight between the current node and adjacent nodes based on the correlation coefficients, calculate the resource occupancy intensity of the nodes based on the resource occupancy data, count the number of downstream nodes affected by the nodes and their importance levels, and obtain the downstream influence scope of the nodes; Perform a weighted combination of the correlation coefficient weight, resource occupancy intensity, and downstream influence scope to obtain the processing priority of the nodes, and generate a fault handling priority list according to the processing priority of the nodes; Construct a state space with the performance metrics of the nodes as state variables, construct an action space with the resource adjustment operations of the nodes as action variables, and sort the operations in the action space based on the fault handling priority list to obtain an initial action sequence; Extract state transition information from the fault handling historical data, calculate the performance benefits of different operation sequences based on the state transition information, and adjust the initial action sequence in combination with the resource cost of the operations to obtain the final reinforcement learning action sequence.
[0011] In an alternative embodiment, Execute the reinforcement learning action sequence, and calculate in real time the degree of influence of the executed actions on the states of the nodes in the propagation link. Divide the actions into a main control action set and an adjustment action set according to the degree of influence based on a preset influence threshold, including: Execute the actions in the reinforcement learning action sequence, collect in real time the performance metric data of the nodes in the propagation link before and after the execution of the actions, and construct a performance influence evaluation matrix with the performance metric data according to the positive change values and negative change values; Calculate the initial influence vector of the actions on the states of the nodes in the propagation link based on the performance influence evaluation matrix, and perform temporal weighting on the initial influence vector using a time decay function to obtain a temporal change sequence of the node states; Determine the node weight coefficients according to the positional relationships of the nodes in the propagation link, perform weighted calculation on the temporal change sequence to obtain the degree of influence of the actions on the node states, and construct a node state transition probability matrix based on the degree of influence; Calculate the influence degree value of the actions in combination with the node state transition probability matrix and the state transition benefits, compare the influence degree value with the preset influence threshold, and divide the executed actions into a main control action set and an adjustment action set according to the comparison result of the influence degree value.
[0012] In an alternative embodiment, Collect performance monitoring metric data, calculate the performance improvement ratio, adjust the execution order of the master action set and the execution parameters of the adjustment action set according to the performance improvement ratio, generate an optimized reinforcement learning action sequence, and continue to execute the optimized reinforcement learning action sequence until the fault recovery is completed, including: Collect performance monitoring metric data, where the performance monitoring metric data includes the current moment metric value and the reference moment metric value, assign weight coefficients to different monitoring metrics according to the system resource occupancy, and multiply the weight coefficient by the change amount of the monitoring metric value to obtain the performance improvement ratio; Calculate the difference in the performance improvement ratio before and after the execution of each action in the master action set to obtain the performance change amount, and calculate the performance change rate according to the ratio of the performance change amount to the execution time interval; Based on the performance change rate, prioritize the actions in the master action set to obtain the execution order of the master actions. At the same time, calculate the correlation degree between the execution parameters of each action in the adjustment action set and the performance improvement ratio, determine the parameter adjustment range according to the correlation degree, and update the execution parameters of the actions in the adjustment action set according to the parameter adjustment range to generate the execution configuration of the adjustment actions; Combine the sorted execution order of the master actions with the updated execution configuration of the adjustment actions to generate an optimized reinforcement learning action sequence, execute the optimized reinforcement learning action sequence and update the performance improvement ratio in real time, and adjust the execution order and execution parameters of the unexecuted actions according to the updated performance improvement ratio until the fault recovery is completed.
[0013] In the second aspect of the embodiments of the present invention, a database fault handling system based on reinforcement learning is provided, including: A first unit for collecting performance metric data, execution plan data, and error log data at the database fault moment, generating a fault state vector, calculating a time series correlation coefficient matrix based on the fault state vector, constructing a causal decision tree according to the time series correlation coefficient matrix, identifying a fault root cause node from the causal decision tree, and generating a propagation link from the fault root cause node to the performance anomaly node; A second unit for generating a fault handling priority list based on the correlation coefficient magnitude and resource occupancy degree of the nodes in the propagation link, and constructing a reinforcement learning action sequence according to the fault handling priority list; A third unit for executing the reinforcement learning action sequence, and calculating in real time the influence degree of the executed actions on the node states in the propagation link, and dividing the actions into a master action set and an adjustment action set according to the influence degree according to a preset influence threshold; The fourth unit is used to collect performance monitoring index data, calculate the performance improvement ratio, adjust the execution order of the master action set and the execution parameters of the adjustment action set according to the performance improvement ratio, generate an optimized reinforcement learning action sequence, and continue to execute the optimized reinforcement learning action sequence until the fault recovery is completed.
[0014] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0015] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0016] In this embodiment, by establishing a causal decision tree to identify the root cause of the fault and generate a propagation link, the source of the fault can be accurately located and the fault impact path can be understood, avoiding the problem that traditional methods only focus on surface phenomena and ignore the root cause, and improving the accuracy and efficiency of fault location. Generating a processing priority list and constructing a reinforcement learning action sequence according to the correlation coefficient and resource occupancy degree realizes intelligent decision-making in the fault handling process. The system can automatically determine the optimal order of processing steps, reduce manual intervention, and improve the automation degree and response speed of fault handling. By real-time evaluating the action execution effect and dynamically adjusting the processing strategy, the actions are divided into master actions and adjustment actions, and the execution order and parameters are optimized according to the performance improvement ratio, realizing the adaptive optimization of fault handling, enabling the system to continuously adjust the recovery strategy according to the actual effect, greatly improving the success rate and efficiency of fault recovery, and reducing the impact time of database faults on services. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flow chart of the method for handling database faults based on reinforcement learning according to the embodiments of the present invention; Figure 2 It is a schematic diagram of a causal decision tree constructed based on a weighted correlation coefficient; Figure 3 It is a schematic diagram of the simulation results of fault handling priority and operation benefit. DETAILED DESCRIPTION
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0020] Figure 1 The flowchart of the database fault handling method based on reinforcement learning according to the embodiment of the present invention is shown as Figure 1 shown, and the method includes: Collect performance metric data, execution plan data, and error log data at the moment of database failure, generate a fault state vector, calculate a time series correlation coefficient matrix based on the fault state vector, construct a causal decision tree according to the time series correlation coefficient matrix, identify the fault root cause node from the causal decision tree, and generate a propagation link from the fault root cause node to the performance anomaly node; Generate a fault handling priority list based on the correlation coefficient magnitude and resource occupancy degree of the nodes in the propagation link, and construct a reinforcement learning action sequence according to the fault handling priority list; Execute the reinforcement learning action sequence, and calculate in real time the influence degree of the executed actions on the node states in the propagation link. According to the influence degree, divide the actions into a main control action set and an adjustment action set according to a preset influence threshold; Collect performance monitoring metric data to calculate the performance improvement ratio, adjust the execution order of the main control action set and the execution parameters of the adjustment action set according to the performance improvement ratio, generate an optimized reinforcement learning action sequence, and continue to execute the optimized reinforcement learning action sequence until the fault recovery is completed.
[0021] In an alternative embodiment, collecting performance metric data, execution plan data, and error log data at the moment of database failure, generating a fault state vector, calculating a time series correlation coefficient matrix based on the fault state vector, and constructing a causal decision tree according to the time series correlation coefficient matrix includes: Collect performance metric data, execution plan data, and error log data at the moment of database failure; Perform periodic sampling on the performance metric data to obtain the resource usage, query response time, and data processing volume within each sampling period; extract the access path, parallelism, and resource pre-estimation value of the query from the execution plan data; extract the fault type identifier and associated object identifier from the error log data; Calculate the deviation degree between the resource usage and the corresponding pre-estimation value, calculate the delay degree between the query response time and the reference time, and perform an association mapping between the fault type identifier and the associated object identifier; generate a fault status vector based on the deviation degree, delay degree, and association mapping result; Perform a sliding window process on the fault status vector, calculate the temporal correlation coefficient between each index within the window, and construct a temporal correlation coefficient matrix according to the index correspondence relationship of the temporal correlation coefficient; weight the correlation coefficients in the temporal correlation coefficient matrix according to the resource pre-estimation value, and construct a causal decision tree based on the weighted temporal correlation coefficient matrix. Among them, determine the index pair with the strongest correlation coefficient as the root node of the tree, construct child nodes in sequence according to the correlation coefficient strength, and use the resource pre-estimation value as the constraint condition for the connection between nodes.
[0022] Figure 2 As a schematic diagram of a causal decision tree constructed based on weighted correlation coefficients, this embodiment provides a database fault diagnosis method, which collects multi-dimensional data during database faults and constructs a causal decision tree for root cause analysis of faults.
[0023] Automatically start collecting performance metric data, execution plan data, and error log data when a database system fails. The performance metric data includes system resource metrics such as CPU usage rate, memory occupancy, IO throughput, and network transmission volume; the execution plan data includes information such as access method, scanned row count, and returned row count in the SQL statement execution plan; the error log data includes fault information such as error codes, error descriptions, and exception stacks.
[0024] Perform periodic sampling processing on the collected performance metric data, and set the sampling period to 10 seconds. Taking the CPU usage rate as an example, collect 30 sampling points before and after the fault, and record the usage rate value of each sampling point. For example, in a certain fault case, it is monitored that the CPU usage rate rises sharply from 25% to 95%, stays at a high level for 2 minutes, and then drops back to 35%. The memory usage increases from the original stable occupancy of 8GB to 14.5GB, approaching the threshold of the system's total memory of 16GB. The IO throughput increases from 200MB per second to 800MB per second, and the query response time extends from an average of 50ms to 350ms.
[0025] In terms of the processing of execution plan data, extract the execution plan of the SQL statement currently being executed and analyze its access path characteristics. For example, in a failure case, a full table scan operation was detected, with an estimated number of scanned rows of 5 million, but the actual number of scanned rows reached 7.2 million; the degree of parallelism was set to 8, but actually only 3 parallel threads were working; the estimated execution time was 2 seconds, but the actual execution time reached 12 seconds. In addition, key information such as index usage, JOIN method, and sorting method in the execution plan was also extracted.
[0026] When processing error log data, first extract the error code and error description information. For example, "ORA-01555" indicates a snapshot too old error, and "ORA-04031" indicates insufficient shared pool memory. Further extract the object identifiers associated with the error, such as the table name "USER_TRANSACTION" and the index name "IDX_USER_ID", etc. In the example failure, multiple error logs of "insufficient shared pool memory" were detected, and the associated objects were large complex queries and temporary tablespaces.
[0027] Based on the processed data, calculate the deviation degree of resource usage. Compare the difference between the actual CPU usage rate of 95% and the estimated value of 40%, and calculate the deviation rate as 137.5%; the memory usage deviation rate is 81.25%; the IO throughput deviation rate is 300%. The query response time delay degree is calculated as 600%, which is much higher than the system-acceptable threshold of 50%. The error log association mapping shows that 90% of the errors are related to memory allocation problems, mainly associated with the shared pool and temporary tablespaces.
[0028] Based on the above analysis results, generate a 36-dimensional failure state vector, which includes indicators such as the deviation values of various resource usages, query delay degree, and error association degree. In this vector, the values of memory usage deviation and shared pool error association degree are significantly higher than other indicators.
[0029] Process the failure state vector using a 30-second sliding window with a window step of 5 seconds, and calculate the temporal correlation between various indicators within the window. For example, the correlation coefficient between CPU usage rate and query response time is 0.78, indicating a strong positive correlation; the correlation coefficient between memory usage rate and shared pool errors is 0.92, indicating an extremely strong positive correlation; the correlation coefficient between disk IO and query response time is 0.45, indicating a medium positive correlation.
[0030] Organize the calculated correlation coefficients into a 36×36 time-series correlation coefficient matrix, where each cell in the matrix represents the time-series correlation strength between the corresponding two metrics. Subsequently, adjust the correlation coefficients weighted according to the resource estimation values in the execution plan. For example, if the estimated memory resource ratio is 60%, the CPU estimation ratio is 25%, and the IO estimation ratio is 15%, then adjust the coefficient weights related to these resources accordingly. After weighting, the correlation coefficient between memory utilization and shared pool errors increases from 0.92 to 0.95, becoming the highest value in the matrix.
[0031] Construct a causal decision tree based on the weighted time-series correlation coefficient matrix. Take the metric pair with the highest correlation coefficient, "memory utilization - shared pool errors", as the root node, with a correlation coefficient of 0.95. From this node downwards, construct child nodes in sequence according to the correlation coefficient strength: "shared pool errors - temporary tablespace utilization" (0.88), "temporary tablespace utilization - large complex queries" (0.86), "large complex queries - full table scan operations" (0.82). At the same time, use the resource estimation value as the node connection constraint condition. For example, only when the estimated memory usage exceeds 50%, establish the connection between "memory utilization" and other metrics.
[0032] The constructed causal decision tree clearly shows the fault propagation path from "full table scan operations" to "shared pool errors" and then to "abnormal memory utilization", providing the database administrator with accurate fault root cause analysis results, indicating that the fault source is the exhaustion of memory resources caused by a certain large query.
[0033] In this embodiment, by fusing performance metrics, execution plans, and error log data, constructing a fault state vector, and introducing time-series correlation analysis and weighted causal modeling, it is possible to achieve multi-dimensional accurate characterization of database faults and mining of causal chains. On the one hand, resource anomalies and query bottlenecks can be accurately identified through deviation and delay metrics, improving the accuracy of fault location; on the other hand, key correlation relationships in the time series are extracted through sliding windows and weighted correlation coefficient matrices, enhancing the structural rationality and interpretability of the causal decision tree, thereby realizing visual traceability and intelligent analysis of fault root causes, providing effective support for the adaptive diagnosis and optimization of the database system.
[0034] In an alternative implementation, identifying the fault root cause node from the causal decision tree and generating the propagation link from the fault root cause node to the performance anomaly node includes: Calculate the fault impact degree of each node in the causal decision tree, and combine the out-degree value, correlation coefficient strength, and resource occupancy rate of the node to obtain the fault impact degree value of the node; Start traversing backward from the leaf nodes of the causal decision tree, calculate the influence transfer value of each node based on the connection strength and fault impact degree value between nodes, and mark the nodes where the influence transfer value shows a mutation as candidate root cause nodes; Starting from the candidate root cause nodes, calculate the abnormality value by propagating downward along the causal decision tree based on the fault impact degree, and compare the abnormality value with the abnormal value in the performance metric data. When the comparison result meets the matching condition, confirm the current candidate root cause node as the fault root cause node; Starting from the fault root cause node, adopt a depth-first search strategy, determine the propagation direction according to the correlation coefficient strength between nodes, and prune the currently obtained propagation link when the temporal correlation coefficient between nodes is lower than the preset correlation threshold to obtain the initial propagation link; Exemplarily, in a cloud computing environment, a server cluster runs multiple microservices applications, and a system monitoring platform continuously collects performance metric data such as CPU usage, memory occupancy, and network traffic. When performance anomalies are detected, the system constructs a causal decision tree based on historical monitoring data to analyze the fault root cause and its propagation path.
[0035] After constructing the causal decision tree, it is necessary to calculate the fault impact degree of each node in the tree. The fault impact degree is determined by combining the out-degree value, correlation coefficient strength, and resource occupancy rate of the node. Specifically, for node N, its out-degree value represents the number of downstream nodes directly connected to this node; the correlation coefficient strength represents the average value of the Pearson correlation coefficients between this node and its downstream nodes; the resource occupancy rate represents the percentage of the system resources represented by this node used. The fault impact degree FI(N) of node N is calculated as the weighted sum of these three factors. In practical applications, the weights can be set to 0.3, 0.4, and 0.3 respectively. For example, if a node has an out-degree value of 4, a correlation coefficient strength of 0.75, and a resource occupancy rate of 85%, then its fault impact degree is 0.3×4 + 0.4×0.75 + 0.3×0.85 = 1.94.
[0036] After calculating the fault impact degree of the nodes, start traversing backward from the leaf nodes of the causal decision tree and calculate the influence transfer value of each node. For leaf node L, its influence transfer value IT(L) is initially set to its fault impact degree FI(L). For non-leaf node P, its influence transfer value IT(P) is the sum of its fault impact degree FI(P) and the influence transfer values of all its child nodes multiplied by the corresponding connection strength. The connection strength CS(P, C) is defined as the absolute value of the correlation coefficient between node P and its child node C. During the backward traversal process, record the influence transfer value of each node and calculate the change rate of the influence transfer values between adjacent nodes. When the change rate exceeds the preset threshold (such as 30%), mark this node as a candidate root cause node.
[0037] Taking a causal decision tree with 10 nodes as an example, the influence transfer values of nodes A to J are 2.1, 1.8, 1.4, 1.9, 2.5, 1.2, 0.9, 1.6, 1.3, and 2.0 respectively. By calculating the change rate between adjacent nodes, it is found that the change rate from node D to E is (2.5 - 1.9) / 1.9 ≈ 31.6%, exceeding the threshold of 30%. Therefore, node E is marked as a candidate root cause node.
[0038] Starting from the candidate root cause node, calculate the abnormality value by propagating downward along the causal decision tree. For the root cause node R, its abnormality is initialized to 1.0. For other nodes X, its abnormality AD(X) is the abnormality AD(P) of its parent node P multiplied by the connection strength CS(P, X) and then multiplied by the fault impact degree FI(X) of node X. Compare the calculated abnormality value with the actual abnormality value in the performance metric data. The comparison method is to calculate the similarity between the abnormality sequence and the actual abnormality sequence, and the cosine similarity or Pearson correlation coefficient can be used. When the similarity exceeds the preset threshold (such as 0.8), confirm the current candidate root cause node as the fault root cause node.
[0039] Calculate the abnormality propagation starting from node E, and the abnormality values from node E to J are 1.0, 0.75, 0.62, 0.85, 0.70, and 0.91 respectively. Compare these values with the actually monitored performance abnormality data (after normalization): 1.0, 0.78, 0.65, 0.82, 0.68, and 0.88. Calculate the cosine similarity of the two sequences to be 0.997, exceeding the threshold of 0.8. Therefore, confirm node E as the fault root cause node. After determining the fault root cause node, generate the fault propagation link from this node using the depth-first search strategy. During the search process, determine the propagation direction according to the correlation coefficient strength between nodes, and preferentially select the connection with the largest absolute value of the correlation coefficient. To improve accuracy, when the temporal correlation coefficient between nodes is lower than the preset threshold (such as 0.3), prune the current search path.
[0040] In the example, start searching from node E and find the connected nodes F, G, and H, whose correlation coefficients are 0.65, 0.28, and 0.70 respectively. Since the correlation coefficient of node G, 0.28, is lower than the threshold of 0.3, the path leading to G is pruned. Continue to search for nodes F and H, and find that node H is connected to node J with a correlation coefficient of 0.82; node F is connected to node I with a correlation coefficient of 0.41. Finally, obtain the initial propagation link: E→H→J and E→F→I. Finally, verify the initial propagation link and confirm the link validity based on the performance metric change trend and temporal relationship between nodes. The verification method is to check whether there is a reasonable temporal change pattern in the performance metrics of adjacent nodes in the link. For example, the abnormality of the upstream node should occur earlier in time than the downstream node.
[0041] In the link E→H→J, by analyzing the time series of performance metrics of the three nodes, it is found that the abnormal occurrence time of node E is T1, the abnormal occurrence time of node H is T1 + 2 seconds, and the abnormal occurrence time of node J is T1 + 5 seconds, which conforms to the timing logic of the fault propagating from E to J. In the link E→F→I, the abnormal occurrence time of node F is T1 + 15 seconds, and the abnormal occurrence time of node I is T1 + 10 seconds, which does not conform to the timing logic of propagating from F to I. Therefore, this link is considered invalid. The final output is the valid fault propagation link E→H→J, indicating that the fault propagates from the root cause node E to the performance abnormal node J.
[0042] In this embodiment, by constructing a causal decision tree and introducing a fault impact degree and propagation path recognition mechanism, automatic identification of the root cause of database faults and accurate construction of the propagation link are achieved. Existing technologies mostly rely on static rules or manual experience for fault analysis, making it difficult to dynamically depict the causal relationship between multiple metrics, and lacking a systematic modeling of the influence intensity and propagation path effectiveness during the root cause tracing process, resulting in unstable positioning results, redundant or missing links. To address the above problems, this application proposes to fuse the out-degree, correlation intensity, and resource occupancy rate of nodes to calculate the fault impact degree, quantitatively evaluate the abnormality dominance of nodes, and identify candidate root causes by combining the change of influence transfer values, enhancing the accuracy of root cause identification. At the same time, a pruning strategy based on relevant thresholds and a performance trend verification mechanism are adopted to optimize the propagation chain construction process, avoid interference from low-correlation paths, and effectively improve the accuracy and reliability of the link. Starting from enhancing the discriminant ability of causal path recognition, this improvement significantly improves the degree of automation and analysis effect of fault location in complex multi-metric scenarios.
[0043] In an alternative embodiment, starting from the fault root cause node, a depth-first search strategy is adopted, and the propagation direction is determined according to the correlation coefficient strength between nodes. When the temporal correlation coefficient between nodes is lower than the preset correlation threshold, pruning is performed on the currently obtained propagation link. The initial propagation link includes: Obtain the time series data of performance metrics of the fault root cause node and its connected nodes, analyze the time series data of performance metrics using a sliding time window, and calculate the temporal correlation coefficient between adjacent nodes; Determine the propagation direction between nodes based on the magnitude of the temporal correlation coefficient, record all propagation directions of the fault root cause node, and set a preset correlation threshold for pruning the propagation link; Starting from the fault root cause node, traverse the nodes using the depth - first search strategy, preferentially select the propagation direction with the largest time - series correlation coefficient, successively add adjacent nodes with time - series correlation coefficients greater than the preset correlation threshold to the current propagation link, and use the newly added node as the current search node to continue the search until the time - series correlation coefficients of all adjacent nodes of the current search node are lower than the preset correlation threshold; Trace back from the current search node to the previous search nodes in sequence. For nodes among the previous search nodes that have un - searched propagation directions, select the propagation direction with the second - largest time - series correlation coefficient to continue the depth - first search. After completing the search of all propagation directions, merge the retained propagation links to obtain the initial propagation link.
[0044] In a specific embodiment, the detailed process of the fault propagation link analysis method is as follows: When a fault is detected in the system, first determine the fault root cause node. In this embodiment, the fault root cause node is node S1 in the server cluster, and the CPU usage rate of this node has increased abnormally, triggering a system alarm.
[0045] Obtain the time - series data of the performance metrics of the fault root cause node S1 and its connected nodes. In this embodiment, collect the performance metric data of node S1 and its adjacent nodes S2, S3, and S4 in the most recent 30 minutes, including key metrics such as CPU usage rate, memory usage rate, network throughput, and disk I / O. These data are recorded at a sampling frequency of one point every 10 seconds, forming a time - series data set.
[0046] Apply a sliding time window to analyze the collected time - series data of the performance metrics. The size of the sliding window is set to 5 minutes, and the step size is 1 minute, that is, it slides forward 1 minute each time. Within each sliding window, calculate the time - series correlation coefficient between adjacent nodes. In this embodiment, calculate the time - series correlation coefficients of each metric between node S1 and its adjacent nodes S2, S3, and S4. For example, the time - series correlation coefficient of the CPU usage rate between S1 and S2 is 0.85, the time - series correlation coefficient of the CPU usage rate between S1 and S3 is 0.62, and the time - series correlation coefficient of the CPU usage rate between S1 and S4 is 0.35. Determine the propagation direction between nodes based on the calculated time - series correlation coefficients. Set the preset correlation threshold to 0.5 for subsequent pruning of the propagation link. In this embodiment, there are two propagation directions for node S1: S1→S2 and S1→S3, because the correlation coefficients between these two pairs of nodes are both greater than the preset threshold of 0.5, while the correlation coefficient of S1→S4 is 0.35, which is less than the threshold and is not used as a valid propagation direction.
[0047] Starting from the fault root cause node S1, the nodes in the network topology are traversed using the depth-first search strategy. The propagation direction with the largest temporal correlation coefficient is preferentially selected, that is, first select S1→S2 (correlation coefficient 0.85), add the node S2 to the current propagation link, and form the link S1→S2.
[0048] Take the newly added node S2 as the current search node and continue the search. Analyze the temporal correlation coefficients between S2 and its adjacent nodes S5 and S6, and obtain that the correlation coefficient of S2→S5 is 0.78, and the correlation coefficient of S2→S6 is 0.42. Since the correlation coefficient of S2→S5 is greater than the threshold 0.5, add S5 to the propagation link to form the link S1→S2→S5. The correlation coefficient of S2→S6 is less than the threshold, so S6 is not added to the link. Continue to search with S5 as the current search node, analyze the temporal correlation coefficients between S5 and its adjacent nodes S7 and S8, and obtain that the correlation coefficient of S5→S7 is 0.33, and the correlation coefficient of S5→S8 is 0.25. Since both of these correlation coefficients are less than the threshold 0.5, S7 and S8 are not added to the propagation link, and the current link search ends here, obtaining a propagation link S1→S2→S5.
[0049] Backtrack to the node S1 and continue to search for the sub-optimal propagation direction S1→S3 (correlation coefficient 0.62). Add S3 to the propagation link to form a new link S1→S3. Analyze the temporal correlation coefficients between S3 and its adjacent nodes S9 and S10, and obtain that the correlation coefficient of S3→S9 is 0.71, and the correlation coefficient of S3→S10 is 0.29. Since the correlation coefficient of S3→S9 is greater than the threshold, add S9 to the propagation link to form the link S1→S3→S9. Take S9 as the current search node and continue the search. Analyze the temporal correlation coefficient between S9 and its adjacent node S11, and obtain that the correlation coefficient of S9→S11 is 0.65. Therefore, add S11 to the propagation link to form the link S1→S3→S9→S11. Continue to take S11 as the current search node, analyze the temporal correlation coefficient between S11 and its adjacent node S12, and obtain that the correlation coefficient of S11→S12 is 0.41. Since the correlation coefficient is less than the threshold 0.5, S12 is not added to the propagation link, and the current link search ends here, obtaining the second propagation link S1→S3→S9→S11.
[0050] After completing the search for all propagation directions, the two obtained propagation links S1→S2→S5 and S1→S3→S9→S11 are merged to form the initial propagation link: starting from the root cause node S1, on the one hand, it affects S5 through S2, and on the other hand, it affects S9 and S11 through S3.
[0051] In practical applications, to improve the accuracy of analysis, the preset relevant thresholds can be adjusted according to the system characteristics. A higher threshold (such as 0.7) will result in a shorter but more reliable propagation link, which is suitable for scenarios with high requirements for the accuracy of fault propagation; a lower threshold (such as 0.4) will result in a longer propagation link, which may include more affected nodes and is suitable for scenarios that require a comprehensive understanding of the scope of fault impact. In this embodiment, a medium threshold of 0.5 is selected, which not only ensures the reliability of the propagation link but also does not overly limit the exploration scope of fault propagation.
[0052] The initial propagation link obtained through the above method provides a visual representation of the fault propagation path for system operation and maintenance personnel, helping to quickly locate the affected system components and improve the efficiency of fault troubleshooting and repair. This method is particularly suitable for fault analysis in large-scale distributed systems, which can effectively reduce the fault handling time and minimize the losses caused by system unavailability.
[0053] By introducing a depth-first search strategy driven by temporal correlation coefficients, a propagation path supported by the association strength is constructed in database fault analysis. Most existing technologies adopt static topological structures or path inference methods based on prior rules, lacking the modeling of the temporal variation law in actual operation data, making it difficult to accurately identify the true influence path between nodes in the fault propagation chain, and prone to path omission or misjudgment. This solution starts from the fault root cause node, combines the temporal correlation coefficients between performance indicators calculated by a sliding time window to dynamically determine the propagation direction, and completes link pruning by setting thresholds to effectively exclude the interference of weakly correlated nodes. By preferentially searching for the propagation path with the greatest correlation, it not only improves the accuracy of link construction but also enhances the interpretability of the path structure. This solution ensures high correlation and practical utility of the path while maintaining the search efficiency, providing a reliable basis for subsequent fault evolution analysis and performance optimization, and significantly improving the deficiencies of traditional methods in terms of dynamics and relevance.
[0054] In an alternative embodiment, based on the correlation coefficient magnitude and resource occupancy degree of nodes in the propagation link, a fault handling priority list is generated. Constructing a reinforcement learning action sequence according to the fault handling priority list includes: Obtain the correlation coefficients between nodes in the fault propagation link and the resource occupancy data of each node; Based on the correlation coefficients, calculate the correlation coefficient weights between the current node and adjacent nodes. Based on the resource occupancy data, calculate the resource occupancy intensity of the nodes, and count the number of downstream nodes affected by the nodes and their importance to obtain the downstream influence scope of the nodes; Perform weighted combination of the correlation coefficient weights, resource occupancy intensity, and downstream influence scope to obtain the processing priority of the nodes, and generate a fault handling priority list according to the processing priority of the nodes; Construct a state space with the performance metrics of nodes as state variables, construct an action space with the resource adjustment operations of nodes as action variables, and sort the operations in the action space based on the fault handling priority list to obtain an initial action sequence; Extract state transition information from the fault handling historical data, calculate the performance benefits of different operation sequences based on the state transition information, and adjust the initial action sequence in combination with the resource cost of the operations to obtain the final reinforcement learning action sequence.
[0055] In this embodiment, when obtaining the correlation coefficients and resource occupancy data between nodes in the fault propagation link, a distributed monitoring collector is used to collect the performance metrics and resource usage of each node in the network topology. For a set of nodes {A, B, C, D, E} in a cloud platform environment, the correlation coefficient matrix between nodes is recorded. For example, the correlation coefficient between A and B is 0.85, the correlation coefficient between A and C is 0.72, the correlation coefficient between B and D is 0.68, and the correlation coefficient between C and E is 0.79. At the same time, the resource occupancy data of each node is collected, including CPU usage rate, memory occupancy rate, and I / O throughput. For example, the CPU usage rate of node A is 85%, the memory occupancy rate is 73%, and the I / O throughput is 620 MB / s; the CPU usage rate of node B is 65%, the memory occupancy rate is 58%, and the I / O throughput is 480 MB / s.
[0056] When calculating the correlation coefficient weight between the current node and adjacent nodes, calculate the sum of the correlation coefficients between each node and all its adjacent nodes, and divide the single correlation coefficient by the sum to obtain the normalized weight. Taking node A as an example, its correlation coefficients with adjacent nodes B and C are 0.85 and 0.72 respectively, and the sum of the correlation coefficients is 1.57. Then the correlation coefficient weight between A and B is 0.85 / 1.57 = 0.54, and the correlation coefficient weight between A and C is 0.72 / 1.57 = 0.46.
[0057] When calculating the node resource occupancy intensity, the system normalizes each resource occupancy index of the node and assigns different weights for weighted summation. For node A, the CPU usage rate of 85%, the memory occupancy rate of 73%, and the I / O throughput of 620 MB / s (assuming the maximum value is 1000 MB / s, normalized to 0.62) are respectively assigned weights of 0.4, 0.4, and 0.2 for weighted summation, and the resource occupancy intensity is obtained as 85% × 0.4 + 73% × 0.4 + 62% × 0.2 = 76.9%.
[0058] When counting the number and importance of downstream nodes affected by a statistical node, the system traverses the fault propagation link, records all downstream nodes that each node may affect, and assigns different weights according to the business importance of the downstream nodes. Taking node A as an example, it directly affects nodes B and C and indirectly affects nodes D and E. Assuming the node importance weights are B = 0.8, C = 0.9, D = 0.7, and E = 0.6 respectively, the downstream influence range of A is calculated as 0.8 + 0.9 + 0.7×0.68 (correlation coefficient from B to D) + 0.6×0.79 (correlation coefficient from C to E) = 2.23.
[0059] When obtaining the node processing priority by weighted combination of the correlation coefficient weight, resource occupation intensity, and downstream influence range, the system sets the weights of the three factors to 0.3, 0.3, and 0.4 respectively. For node A, its processing priority = 0.54×0.3 + 0.769×0.3 + 2.23×0.4 = 1.25. Similarly, calculate the processing priorities of other nodes. For example, the processing priority of node B is 0.92. Sort all nodes according to the processing priority from high to low based on the calculation results to generate a fault handling priority list. When constructing the state space, use the performance indicators of the nodes as state variables, including indicators such as response time, throughput, and error rate. For example, the state of node A can be expressed as {response time = 120ms, throughput = 2000QPS, error rate = 0.5%}. When constructing the action space, use the resource adjustment operations of the nodes as action variables, including increasing the number of CPU cores, expanding the memory capacity, adjusting the I / O priority, etc. For example, the action for node A can be {increase 2 CPU cores, increase 4GB of memory, increase the I / O priority}.
[0060] When sorting the operations in the action space based on the fault handling priority list, the system combines and sorts the adjustment operations that can be taken for each node in the order of the nodes in the priority list. For the node A with the highest priority, the possible action sequences are {increasing the number of CPU cores, expanding the memory capacity, adjusting the I / O priority}; for the node C, the action sequences are {increasing the memory capacity, adjusting the network bandwidth, optimizing the cache configuration}. Sort the action sequences of each node according to the processing priority to obtain the initial action sequence. When extracting the state transition information from the fault handling historical data, analyze the impact of various operations on the system state in past similar fault scenarios. For example, when adding 2 CPU cores in a state similar to node A, the response time drops from 120 ms to 80 ms, the throughput increases from 2000 QPS to 2800 QPS, and the error rate drops from 0.5% to 0.1%. Calculate the performance benefits of different operation sequences based on the state transition information, and the system evaluates the improvement degree of each operation on the key performance indicators. For example, the benefit score of the operation of increasing the number of CPU cores is (120 - 80) / 120×0.4+(2800 - 2000) / 2000×0.4+(0.5 - 0.1) / 0.5×0.4 = 0.59.
[0061] When adjusting the initial action sequence in combination with the resource cost of the operations, consider the resource cost required for each operation. For example, the resource cost of adding 2 CPU cores is 0.3, and the resource cost of adding 4GB of memory is 0.2. Calculate the comprehensive benefit of the operation = performance benefit / resource cost. For example, the comprehensive benefit of increasing the number of CPU cores = 0.59 / 0.3 = 1.97, and the comprehensive benefit of increasing the memory capacity = 0.45 / 0.2 = 2.25. Re-sort the action sequence according to the comprehensive benefit to obtain the final reinforcement learning action sequence as.
[0062] In this embodiment, by introducing multi-dimensional indicators such as the correlation coefficient, resource occupancy intensity, and downstream influence range of each node in the propagation link, a refined fault handling priority list is constructed, and on this basis, a dynamically optimized action sequence is generated in combination with the reinforcement learning mechanism, effectively improving the efficiency and accuracy of database fault handling. Based on the propagation chain structure, quantitatively analyze the influence of nodes in the system, and mine the relationship between the performance benefits and resource costs of operations based on historical state transition information, realizing the adaptive optimization and adjustment of the initial action sequence. This improvement aims to improve the fault recovery efficiency and resource usage rationality, not only enhancing the pertinence and executability of the strategy, but also providing intelligent decision support with learning and optimization capabilities for the database system.
[0063] Figure 3It is a schematic diagram of the simulation results of the fault handling priority and operation benefit. The red circles represent the operations of node A (processing priority 1.25), which contain three data points and are distributed in the upper left area, indicating that the operations of node A have high performance benefits and low resource costs. The blue triangles represent the operations of node B (processing priority 0.92), which are mainly concentrated in the low-benefit area. The orange squares represent the operations of node C (processing priority 0.98), which are more dispersed. The small gray dots in the figure represent the random sample points generated during the Monte Carlo simulation.
[0064] The dotted line connects the optimal operation sequence path, indicating that the system no longer strictly executes operations according to the node priority (A > C > B), but reorders according to the comprehensive benefit (performance benefit / resource cost). For example, some operations of node B or node C in the figure may be preferentially executed due to higher comprehensive benefits, although the processing priority of their corresponding nodes is lower.
[0065] This optimization method based on reinforcement learning breaks through the limitation of the traditional method that only sorts according to the importance of nodes, can more effectively balance performance improvement and resource consumption, and maximize the fault recovery efficiency under limited resources.
[0066] In an optional implementation manner, execute the reinforcement learning action sequence, and calculate in real time the influence degree of the executed actions on the node states in the propagation link. According to the influence degree, divide the actions into a main control action set and an adjustment action set, including: Execute the actions in the reinforcement learning action sequence, collect the performance index data of the nodes in the propagation link before and after the execution of the actions in real time, and construct a performance influence evaluation matrix based on the positive change values and negative change values of the performance index data; Calculate the initial influence vector of the actions on the node states in the propagation link based on the performance influence evaluation matrix, and perform temporal weighting on the initial influence vector using a time decay function to obtain the temporal change sequence of the node states; Determine the node weight coefficients according to the positional relationship of the nodes in the propagation link, perform weighted calculation on the temporal change sequence to obtain the influence degree of the actions on the node states, and construct a node state transition probability matrix based on the influence degree; Calculate the influence degree value of the actions by combining the node state transition probability matrix and the state transition benefit, compare the influence degree value with the preset influence threshold, and divide the executed actions into a main control action set and an adjustment action set according to the comparison result of the influence degree value.
[0067] In this embodiment, a method for executing a reinforcement learning action sequence and partitioning an action set is provided. When executing the reinforcement learning action sequence, the influence degree of the executed actions on the node states in the propagation link is calculated in real time, and the actions are partitioned into a master control action set and an adjustment action set according to the influence degree.
[0068] During the process of executing the reinforcement learning action sequence, the system first needs to execute actions and collect relevant data. For example, in a network transmission scenario, the actions can be adjusting the routing policy, changing the bandwidth allocation, or modifying the caching mechanism, etc. For each executed action, the performance index data of each node in the propagation link before and after the execution is collected. These indexes can include the latency time, throughput, packet loss rate, CPU utilization rate, etc. Specifically, the system records that the latency of node A before the execution of the action is 20 ms and after the execution is 15 ms, so the positive change value is 5 ms; the throughput of node B before the execution of the action is 100 Mbps and after the execution is 95 Mbps, so the negative change value is 5 Mbps. Based on the collected data, the system constructs a performance impact evaluation matrix. The rows of this matrix represent different nodes, the columns represent different performance indexes, and the matrix element values are the change values of the corresponding indexes.
[0069] To evaluate the initial influence of the actions on the node states, an initial influence vector is calculated based on the performance impact evaluation matrix. For each node, the system comprehensively considers the changes of multiple performance indexes to form an initial influence value. For example, the changes of each performance index of the node can be weighted and summed according to their importance to obtain the initial influence value of the node. If the weight of the latency improvement of node A is 0.6 and the weight of the throughput improvement is 0.4, then its initial influence value is calculated as 0.6×5 + 0.4×(-2) = 2.2, indicating that the action has a positive overall impact on node A.
[0070] Considering that the influence of the actions will decay over time, a time decay function is used to perform temporal weighting on the initial influence vector. The time decay function can be in the form of exponential decay, making the influence weight of the more recent time larger. For example, if an exponential decay function with a decay coefficient of 0.9 is used, then the influence weight at time t is 0.9^t. For the data collected at 10 consecutive time points after the execution of the action, the influence values at each time point are calculated respectively and the corresponding weights are applied to obtain the temporal change sequence of the node state. Taking node A as an example, if the initial influence value at t = 0 is 2.2, at t = 1 is 1.8, and at t = 2 is 1.5, then the weighted temporal change value is 2.2×1 + 1.8×0.9 + 1.5×0.81 = 5.335.
[0071] In the propagation link, the importance of different nodes is usually different. The node weight coefficient is determined according to the positional relationship of the nodes in the propagation link. Nodes in critical positions (such as ingress nodes, aggregation nodes) have higher weights, while edge nodes have lower weights. For example, in a five-node network, the weight of the central node is 0.3, each of the two aggregation nodes is 0.25, and each of the two edge nodes is 0.1. These weight coefficients are used to perform weighted calculations on the time-series change sequence to obtain the final degree of influence of the action on the state of each node. If the time-series change values of nodes A, B, C, D, and E are 5.335, 4.221, 3.872, 2.563, and 1.998 respectively, the comprehensive degree of influence is 0.3×5.335 + 0.25×4.221 + 0.25×3.872 + 0.1×2.563 + 0.1×1.998 = 3.950.
[0072] Based on the calculated degree of influence, the system constructs a node state transition probability matrix. This matrix describes the probability that a node transitions from one state to another after performing a specific action. For example, the node states are divided into four categories: "good", "normal", "warning", and "dangerous". The system calculates that after performing a certain action, the probability that a node transitions from the "warning" state to the "normal" state is 0.75, the probability of transitioning from the "warning" state to the "good" state is 0.15, the probability of remaining in the "warning" state is 0.08, and the probability of transitioning to the "dangerous" state is 0.02.
[0073] The influence degree value of an action is a key indicator for measuring the importance of the action, and the system calculates it by combining the node state transition probability matrix and the state transition benefit. The state transition benefit reflects the value of different state transitions. For example, the benefit value of transitioning from the "dangerous" state to the "good" state is 10, and the benefit value of transitioning from the "warning" state to the "normal" state is 5. The system multiplies the state transition probability by the corresponding state transition benefit and sums them up to obtain the influence degree value of the action. Finally, the calculated influence degree value of the action is compared with a preset influence threshold. If the influence degree value is greater than or equal to the preset influence threshold (such as 4.0), the action is classified into the main control action set; otherwise, it is classified into the adjustment action set. Actions in the main control action set have a significant impact on the system state, while actions in the adjustment action set have a relatively small impact. In the above example, the influence degree value of this action, 4.75, is greater than the preset threshold of 4.0, so it is classified into the main control action set.
[0074] By performing real-time quantitative analysis on the changes in node states during the execution of reinforcement learning actions, a dynamic evaluation mechanism for the impact of actions on the system state is constructed, which can effectively identify the scope of action and priority of key control actions and auxiliary adjustment actions. When executing the fault repair strategy in the prior art, all operations are usually treated equally, and the primary and secondary impacts of actions in the actual control process are not distinguished, resulting in inaccurate resource allocation, long execution paths, or redundant repair strategies, thereby affecting the recovery efficiency and system stability. In contrast, this solution constructs a performance impact evaluation matrix based on the positive and negative changes of performance indicators, combines time decay and node topology weights to accurately reflect the temporal and structural impacts of each action on the system state. At the same time, through the analysis of node state transition probability and benefit, the comprehensive impact degree of actions is calculated, and the master control and adjustment action sets are divided according to the preset threshold, providing a hierarchical basis for subsequent policy optimization and action adjustment. This improvement aims to improve control accuracy and execution efficiency, realizes quantifiable feedback and refined classification of action effects, and helps to construct a more robust and intelligent fault handling system.
[0075] In an alternative embodiment, collecting performance monitoring index data to calculate the performance improvement ratio, adjusting the execution order of the master control action set and the execution parameters of the adjustment action set according to the performance improvement ratio, generating an optimized reinforcement learning action sequence, and continuing to execute the optimized reinforcement learning action sequence until the fault recovery is completed includes: Collect performance monitoring index data, where the performance monitoring index data includes the index value at the current moment and the index value at the reference moment, assign weight coefficients to different monitoring indexes according to the system resource occupancy, and multiply the weight coefficient by the change amount of the monitoring index value to obtain the performance improvement ratio; Calculate the difference in the performance improvement ratio before and after the execution of each action in the master control action set to obtain the performance change amount, and calculate the performance change rate according to the ratio of the performance change amount to the execution time interval; Based on the performance change rate, rank the actions in the master control action set to obtain the execution order of the master control actions. At the same time, calculate the correlation degree between the execution parameters of each action in the adjustment action set and the performance improvement ratio, determine the parameter adjustment range according to the correlation degree, and update the execution parameters of the actions in the adjustment action set according to the parameter adjustment range to generate the execution configuration of the adjustment actions; Combine the sorted execution order of the master control actions and the updated execution configuration of the adjustment actions to generate an optimized reinforcement learning action sequence, execute the optimized reinforcement learning action sequence and update the performance improvement ratio in real time, and adjust the execution order and execution parameters of the unexecuted actions according to the updated performance improvement ratio until the fault recovery is completed.
[0076] In this embodiment, performance monitoring metric data is first collected. These data include the metric values at the current moment and the metric values at the reference moment. The reference moment can be the state when the system is running normally or the state before the last fault recovery operation. Performance monitoring metrics can include CPU usage rate, memory occupancy rate, network throughput, disk I / O rate, service response time, etc. For example, the current CPU usage rate is 85%, and the CPU usage rate at the reference moment is 45%; the current memory occupancy rate is 70%, and the memory occupancy rate at the reference moment is 30%; the current network throughput is 20 MB / s, and the network throughput at the reference moment is 60 MB / s.
[0077] Weight coefficients are assigned to different monitoring metrics according to the resource occupancy situation. In high-load situations, CPU usage rate and memory occupancy rate may be more important, so higher weights are given; while in network-intensive applications, network throughput may be more critical and should be given a higher weight. For example, the weight of CPU usage rate is 0.4, the weight of memory occupancy rate is 0.3, the weight of network throughput is 0.2, and the weight of disk I / O rate is 0.1. Multiply the weight coefficient by the change amount of the monitoring metric value to obtain the performance improvement ratio. For example, the change amount of CPU usage rate is -40% (decreased by 40%), multiplied by the weight 0.4 to get -16%; the change amount of memory occupancy rate is -40%, multiplied by the weight 0.3 to get -12%; the change amount of network throughput is -66.7%, multiplied by the weight 0.2 to get -13.34%; the change amount of disk I / O is 10%, multiplied by the weight 0.1 to get 1%. The total performance improvement ratio is -16% - 12% - 13.34% + 1% = -40.34%, indicating that the system performance has deteriorated by 40.34%.
[0078] For each action in the master control action set, calculate the difference in the performance improvement ratio before and after the execution of each action to obtain the performance change amount. Assume that the master control action set contains four actions: restart the application service, clear the cache, release the connection pool, and reconfigure the load balancer. Before the execution of the "restart the application service" action, the performance improvement ratio is -40.34%, and after the execution, it is -15.20%, and the performance change amount is 25.14%. Before and after the execution of the "clear the cache" action, the performance improvement ratio changes from -40.34% to -35.10%, and the performance change amount is 5.24%. Before and after the execution of the "release the connection pool" action, the performance improvement ratio changes from -40.34% to -30.25%, and the performance change amount is 10.09%. Before and after the execution of the "reconfigure the load balancer" action, the performance improvement ratio changes from -40.34% to -18.65%, and the performance change amount is 21.69%.
[0079] Calculate the performance change rate based on the ratio of the performance change amount to the execution time interval. Assume that the execution time of "restart application service" is 20 seconds, and the performance change rate is 25.14% / 20 = 1.257% / second; the execution time of "clean cache" is 5 seconds, and the performance change rate is 5.24% / 5 = 1.048% / second; the execution time of "release connection pool" is 8 seconds, and the performance change rate is 10.09% / 8 = 1.261% / second; the execution time of "reconfigure load balancing" is 15 seconds, and the performance change rate is 21.69% / 15 = 1.446% / second.
[0080] Based on the performance change rate, sort the actions in the master action set by priority to obtain the execution order of the master actions. According to the above calculation results, the execution order is: reconfigure load balancing (1.446% / second), release connection pool (1.261% / second), restart application service (1.257% / second), clean cache (1.048% / second). At the same time, calculate the correlation between the execution parameters of each action in the adjustment action set and the performance improvement ratio. Assume that the adjustment action set includes: connection pool size adjustment, thread pool size adjustment, cache capacity adjustment, timeout setting. The system observes the change of the performance improvement ratio by changing different parameter values and calculates the correlation coefficient. For example, when the connection pool size increases from 50 to 100, the performance improvement ratio changes from -40.34% to -32.15%, and the correlation degree is 0.82; when the thread pool size increases from 20 to 40, the performance improvement ratio changes from -40.34% to -28.70%, and the correlation degree is 0.91; when the cache capacity increases from 200MB to 500MB, the performance improvement ratio changes from -40.34% to -38.20%, and the correlation degree is 0.31; when the timeout increases from 3 seconds to 5 seconds, the performance improvement ratio changes from -40.34% to -37.50%, and the correlation degree is 0.27.
[0081] Determine the parameter adjustment range according to the correlation degree. The higher the correlation degree, the greater the adjustment range. For example, the correlation degree of the thread pool size is 0.91, and it can be increased by 100%; the correlation degree of the connection pool size is 0.82, and it can be increased by 80%; the correlation degree of the cache capacity is 0.31, and it can be increased by 30%; the correlation degree of the timeout is 0.27, and it can be increased by 20%. Accordingly, update the execution parameters of the actions in the adjustment action set: the thread pool size is set to 40, the connection pool size is set to 90, the cache capacity is set to 260MB, and the timeout is set to 3.6 seconds.
[0082] Combine the sorted main control action execution order with the updated adjustment action execution configuration to generate an optimized reinforcement learning action sequence. The optimized action sequence is: reconfigure the load balancer, release the connection pool (set the connection pool size to 90), restart the application service (set the thread pool size to 40), clear the cache (set the cache capacity to 260MB), and set the timeout to 3.6 seconds.
[0083] Execute the optimized reinforcement learning action sequence and update the performance improvement ratio in real time. For example, after executing "reconfigure the load balancer", the performance improvement ratio is updated from -40.34% to -18.65%; after executing "release the connection pool", the performance improvement ratio is updated from -18.65% to -8.56%; after executing "restart the application service", the performance improvement ratio is updated from -8.56% to 2.30%; after executing "clear the cache", the performance improvement ratio is updated from 2.30% to 7.54%; after executing "set the timeout", the performance improvement ratio is updated from 7.54% to 9.20%.
[0084] Based on the updated performance improvement ratio, the execution order and execution parameters of the unexecuted actions can also be adjusted. For example, if after executing "release the connection pool", the performance improvement ratio only reaches -15.40% instead of the expected -8.56%, the system will recalculate the priorities and parameter settings of the remaining actions. The system continuously executes and adjusts the action sequence until the performance improvement ratio reaches the preset threshold, indicating that the fault recovery is completed. In this example, the final performance improvement ratio is 9.20%, indicating that the system performance has increased by 9.20% compared to the baseline moment, and the fault has been successfully recovered.
[0085] By introducing the performance improvement ratio as a dynamic feedback metric, a quantitative correlation mechanism between action effects and system performance is constructed, realizing the adaptive optimization execution of the reinforcement learning action sequence. In the prior art, the repair operations mostly rely on static rules or preset strategies, lacking a real-time feedback mechanism, and it is difficult to meet the dynamic intervention requirements of multiple nodes and multiple actions in complex systems, often resulting in problems such as action redundancy, unreasonable execution order, or resource waste. In contrast, this solution calculates the performance improvement ratio by weighted calculation of performance metric changes and resource occupancy, and evaluates the actual effectiveness of the main control actions by combining the evolution trend of this ratio in different time windows, thereby dynamically adjusting the action execution order; at the same time, based on the correlation between adjustment actions and performance improvement, the parameter adjustment range is set to finely control the action behavior and improve the flexibility and accuracy of the adjustment response. This solution is driven by real-time performance feedback, establishes a closed-loop optimization mechanism of action-metric-feedback, can significantly improve the efficiency, adaptability, and intelligence level of the fault handling process, promote the system to evolve from static response to dynamic closed-loop, and achieve efficient and reliable automated fault recovery.
[0086] In the second aspect of the embodiments of the present invention, a database fault handling system based on reinforcement learning is provided. The system includes: A first unit, configured to collect performance metric data, execution plan data, and error log data at the moment of database failure, generate a fault status vector, calculate a time series correlation coefficient matrix based on the fault status vector, construct a causal decision tree according to the time series correlation coefficient matrix, identify a fault root cause node from the causal decision tree, and generate a propagation link from the fault root cause node to a performance anomaly node; A second unit, configured to generate a fault handling priority list based on the correlation coefficient magnitudes and resource occupancy degrees of the nodes in the propagation link, and construct a reinforcement learning action sequence according to the fault handling priority list; A third unit, configured to execute the reinforcement learning action sequence, and calculate in real time the influence degree of the executed actions on the node states in the propagation link, and divide the actions into a main control action set and an adjustment action set according to the influence degree according to a preset influence threshold; A fourth unit, configured to collect performance monitoring metric data to calculate a performance improvement ratio, adjust the execution order of the main control action set and the execution parameters of the adjustment action set according to the performance improvement ratio, generate an optimized reinforcement learning action sequence, and continue to execute the optimized reinforcement learning action sequence until the fault recovery is completed.
[0087] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0088] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0089] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.
[0090] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A database fault handling method based on reinforcement learning, characterized in that Including: Collect performance metric data, execution plan data, and error log data at the moment of database failure, generate a failure status vector, calculate a time series correlation coefficient matrix based on the failure status vector, construct a causal decision tree according to the time series correlation coefficient matrix, identify the root cause node of the failure from the causal decision tree, and generate a propagation link from the root cause node of the failure to the performance anomaly node; Generate a failure handling priority list based on the correlation coefficient magnitude and resource occupancy degree of the nodes in the propagation link, and construct a reinforcement learning action sequence according to the failure handling priority list; Execute the reinforcement learning action sequence, and calculate in real time the influence degree of the executed actions on the node states in the propagation link. Divide the actions into a main control action set and an adjustment action set according to the influence degree according to a preset influence threshold; Collect performance monitoring metric data to calculate the performance improvement ratio, adjust the execution order of the main control action set and the execution parameters of the adjustment action set according to the performance improvement ratio, generate an optimized reinforcement learning action sequence, and continue to execute the optimized reinforcement learning action sequence until the failure recovery is completed.
2. The method according to claim 1, wherein Collect performance metric data, execution plan data, and error log data at the moment of database failure, generate a failure status vector, calculate a time series correlation coefficient matrix based on the failure status vector, and construct a causal decision tree according to the time series correlation coefficient matrix including: Collect performance metric data, execution plan data, and error log data at the moment of database failure; Perform periodic sampling on the performance metric data to obtain the resource usage, query response time, and data processing volume in each sampling period; extract the access path, parallelism, and resource pre-estimation value of the query from the execution plan data; extract the failure type identifier and associated object identifier from the error log data; Calculate the deviation degree between the resource usage and the corresponding pre-estimation value, calculate the delay degree between the query response time and the reference time, and perform an association mapping between the failure type identifier and the associated object identifier; generate a failure status vector based on the deviation degree, delay degree, and association mapping result; Perform a sliding window process on the failure status vector, calculate the time series correlation coefficient between each index within the window, and construct a time series correlation coefficient matrix according to the index correspondence of the time series correlation coefficient; weight the correlation coefficients in the time series correlation coefficient matrix according to the resource pre-estimation value, and construct a causal decision tree based on the weighted time series correlation coefficient matrix. Among them, determine the index pair with the strongest correlation coefficient as the root node of the tree, construct sub-nodes in sequence according to the correlation coefficient strength, and use the resource pre-estimation value as the constraint condition for the connection between nodes.
3. The method according to claim 1, wherein Identifying the root cause node of the failure from the causal decision tree and generating a propagation link from the root cause node of the failure to the performance anomaly node includes: Calculate the failure influence degree of each node in the causal decision tree, and perform a combined operation on the out-degree value, correlation coefficient strength, and resource occupancy rate of the node to obtain the failure influence degree value of the node; Traverse backward from the leaf nodes of the causal decision tree, calculate the influence transfer value of each node based on the connection strength and failure influence degree value between nodes, and mark the node where the influence transfer value changes suddenly as a candidate root cause node; Starting from the candidate root cause node, calculate the abnormality value by propagating downward along the causal decision tree based on the fault impact degree, and compare the abnormality value with the abnormal value in the performance index data. When the comparison result meets the matching condition, confirm the current candidate root cause node as the fault root cause node; Starting from the fault root cause node, adopt a depth-first search strategy, determine the propagation direction according to the strength of the correlation coefficient between nodes, and prune the currently obtained propagation link when the temporal correlation coefficient between nodes is lower than the preset correlation threshold to obtain the initial propagation link; Verify the initial propagation link, confirm the link validity based on the change trend and temporal relationship of the performance indexes between nodes, and output the fault propagation link from the fault root cause node to the performance abnormal node.
4. The method according to claim 3, characterized in that, Starting from the fault root cause node, adopt a depth-first search strategy, determine the propagation direction according to the strength of the correlation coefficient between nodes, and prune the currently obtained propagation link when the temporal correlation coefficient between nodes is lower than the preset correlation threshold to obtain the initial propagation link, including: Obtain the performance index time series data of the fault root cause node and its connected nodes, analyze the performance index time series data by using a sliding time window, and calculate the temporal correlation coefficient between adjacent nodes; Determine the propagation direction between nodes based on the magnitude of the temporal correlation coefficient, record all propagation directions of the fault root cause node, and set a preset correlation threshold for propagation link pruning; Starting from the fault root cause node, traverse the nodes by adopting a depth-first search strategy, preferentially select the propagation direction with the largest temporal correlation coefficient, sequentially add the adjacent nodes with the temporal correlation coefficient greater than the preset correlation threshold to the current propagation link, and use the newly added node as the current search node to continue the search until the temporal correlation coefficients of all adjacent nodes of the current search node are lower than the preset correlation threshold; Trace back from the current search node to the previous search nodes in sequence. For the nodes in the previous search nodes that have un-searched propagation directions, select the propagation direction with the second largest temporal correlation coefficient and continue to execute the depth-first search. After completing the search of all propagation directions, merge the retained propagation links to obtain the initial propagation link.
5. The method according to claim 1, characterized in that, Generate a fault handling priority list based on the correlation coefficient magnitude and resource occupancy degree of the nodes in the propagation link, and construct a reinforcement learning action sequence according to the fault handling priority list, including: Obtain the correlation coefficient between each pair of nodes and the resource occupancy data of each node in the fault propagation link; Calculate the correlation coefficient weight between the current node and the adjacent node based on the correlation coefficient, calculate the resource occupancy intensity of the node based on the resource occupancy data, count the number of downstream nodes affected by the node and its importance degree to obtain the downstream influence range of the node; Perform a weighted combination of the correlation coefficient weight, resource occupancy intensity, and downstream influence range to obtain the processing priority of the node, and generate a fault handling priority list according to the processing priority of the node; Construct a state space with the performance metrics of the nodes as state variables, construct an action space with the resource adjustment operations of the nodes as action variables, and sort the operations in the action space based on the fault handling priority list to obtain an initial action sequence; Extract state transition information from the fault handling historical data, calculate the performance benefits of different operation sequences based on the state transition information, and adjust the initial action sequence in combination with the resource cost of the operations to obtain the final reinforcement learning action sequence.
6. The method according to claim 1, characterized in that Execute the reinforcement learning action sequence, and calculate in real time the degree of influence of the executed actions on the states of the nodes in the propagation link. According to the degree of influence, divide the actions into a main control action set and an adjustment action set including: Execute the actions in the reinforcement learning action sequence, collect in real time the performance metric data of the nodes in the propagation link before and after the execution of the actions, and construct a performance impact evaluation matrix with the performance metric data according to the positive change value and the negative change value; Calculate the initial influence vector of the actions on the states of the nodes in the propagation link based on the performance impact evaluation matrix, and perform temporal weighting on the initial influence vector using a time decay function to obtain the temporal change sequence of the node states; Determine the node weight coefficients according to the positional relationship of the nodes in the propagation link, perform weighted calculation on the temporal change sequence to obtain the degree of influence of the actions on the node states, and construct a node state transition probability matrix based on the degree of influence; Calculate the influence degree value of the actions by combining the node state transition probability matrix and the state transition benefits, compare the influence degree value with a preset influence threshold, and divide the executed actions into a main control action set and an adjustment action set according to the comparison result of the influence degree value.
7. The method according to claim 1, wherein Collect performance monitoring metric data to calculate the performance improvement ratio, adjust the execution order of the main control action set and the execution parameters of the adjustment action set according to the performance improvement ratio, generate an optimized reinforcement learning action sequence, and continue to execute the optimized reinforcement learning action sequence until the fault recovery is completed including: Collect performance monitoring metric data, where the performance monitoring metric data includes the current moment metric value and the reference moment metric value, allocate weight coefficients for different monitoring metrics according to the system resource occupancy, and multiply the weight coefficients by the change amount of the monitoring metric values to obtain the performance improvement ratio; Calculate the difference in the performance improvement ratio before and after the execution of each action in the main control action set to obtain the performance change amount, and calculate the performance change rate according to the ratio of the performance change amount to the execution time interval; Based on the performance change rate, perform priority sorting on the actions in the main control action set to obtain the execution order of the main control actions. At the same time, calculate the correlation degree between the execution parameters of each action in the adjustment action set and the performance improvement ratio, determine the parameter adjustment amplitude according to the correlation degree, and update the execution parameters of the actions in the adjustment action set according to the parameter adjustment amplitude to generate the execution configuration of the adjustment actions. Combine the sorted main control action execution order with the updated adjustment action execution configuration to generate an optimized reinforcement learning action sequence, execute the optimized reinforcement learning action sequence and update the performance improvement ratio in real time, and adjust the execution order and execution parameters of the unexecuted actions according to the updated performance improvement ratio until the fault recovery is completed.
8. A database fault handling system based on reinforcement learning, which is used to implement the method described in any one of the foregoing claims 1-7, characterized in that Including: The first unit is used to collect performance metric data, execution plan data, and error log data at the moment of database failure, generate a fault status vector, calculate a time series correlation coefficient matrix based on the fault status vector, construct a causal decision tree according to the time series correlation coefficient matrix, identify the fault root cause node from the causal decision tree, and generate a propagation link from the fault root cause node to the performance anomaly node; The second unit is used to generate a fault handling priority list based on the correlation coefficient size and resource occupancy degree of the nodes in the propagation link, and construct a reinforcement learning action sequence according to the fault handling priority list; The third unit is used to execute the reinforcement learning action sequence, and calculate in real time the influence degree of the executed actions on the node states in the propagation link, and divide the actions into a main control action set and an adjustment action set according to the influence degree with a preset influence threshold; The fourth unit is used to collect performance monitoring metric data to calculate the performance improvement ratio, adjust the execution order of the main control action set and the execution parameters of the adjustment action set according to the performance improvement ratio, generate an optimized reinforcement learning action sequence, and continue to execute the optimized reinforcement learning action sequence until the fault recovery is completed.
9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Equipment fault cause tracing method based on reinforcement learning and knowledge graph
CN112100392A
Reinforcement learning-based power grid regulation and control strategy optimization method
CN113988508A
Dynamic fault tree simplification method and device
CN115599586A
Method and system for self-detection and self-recovery of software fault of energy controller
CN118152172A
Fault reason analysis method for electric power informatization system
CN119806873A
Cited By
Safe and intelligent power distribution system based on AI algorithm and power distribution box
CN121150061A
An AI algorithm-based safe intelligent power distribution system and a power distribution box
CN121150061B
Intelligent business index data analysis method and system based on knowledge graph and large model
CN121329237A