Root cause analysis method based on IT operation and maintenance system
By constructing a spatial topology matrix and spatiotemporal joint analysis of IT equipment, and combining reinforcement learning and knowledge graph models, the problem of lagging root cause analysis in IT operation and maintenance systems is solved, and the automation and accuracy of IT equipment failure root cause analysis are achieved.
Patent Information
- Application Number
- CN202511109970.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-11
AI Technical Summary
The root cause analysis methods of existing IT operations and maintenance systems are unable to cope with the periodic and sudden changes in load, resulting in delayed early warnings, frequent false alarms and missed alarms, and an inability to determine the root cause of failures in a timely and accurate manner, with an over-reliance on human experience.
By acquiring the physical location and real-time environmental data of IT equipment, a spatial topology matrix is constructed. Combined with operational status data, spatiotemporal joint analysis is performed. Reinforcement learning algorithms are used to generate fault propagation paths. Finally, multivariate fault knowledge graphs and causal graph models are combined to determine the root causes of faults.
It enables accurate and timely identification of the root cause of IT equipment failures without relying on manual analysis, avoiding missing the best opportunity to handle failures and improving the automated analysis capabilities of IT operation and maintenance systems.
Smart Images

Figure CN120929293A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of equipment failure root cause analysis technology, and in particular to a root cause analysis method based on an IT operation and maintenance system. Background Technology
[0002] With the rapid development of information technology, the scale of modern enterprise IT infrastructure is constantly expanding, and the architecture of IT systems is becoming increasingly complex. Traditional IT systems typically rely on static threshold alarm mechanisms for single devices, determining whether to trigger an alarm based solely on whether the operating status indicators of a single device exceed preset thresholds. This mechanism ignores the physical connections and logical dependencies between devices, leading to frequent false alarms and missed alarms in actual operation and maintenance.
[0003] In existing technologies, root cause analysis of IT operations and maintenance systems typically uses static threshold analysis. However, dynamic threshold analysis cannot adapt to dynamically changing load environments. In actual use, the load of IT systems often has significant periodic and sudden characteristics. Static thresholds are too rigid in this scenario, which may frequently trigger false alarms during peak load periods or fail to capture anomalies in a timely manner during off-peak load periods, resulting in delayed warnings and missing the best opportunity for fault handling.
[0004] Therefore, it is necessary to provide a root cause analysis method for IT operations and maintenance systems that can cope with the periodicity and suddenness of load. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing root cause analysis, which is difficult to eliminate reliance on human experience and cannot cope with dynamic load changes. It provides a root cause analysis method based on IT operation and maintenance system, which obtains the IT equipment status prediction curve through the IT equipment spatial topology matrix, then obtains the fault propagation path of the IT equipment with abnormal status, and performs root cause analysis based on the fault propagation path, thereby accurately and timely determining the root cause of IT equipment failure. This solves the technical problems of delayed early warning and over-reliance on manual root cause location in IT operation and maintenance system.
[0006] The present invention provides a root cause analysis method based on an IT operations and maintenance system, comprising:
[0007] The physical location data and real-time environmental data of IT equipment are acquired. Based on the physical location data and the real-time environmental data, a matrix is constructed through dual topological relationships to obtain a spatial topology matrix of IT equipment that includes environmental impact factors.
[0008] Acquire the operating status data of IT equipment, perform spatiotemporal joint analysis and dimensionality reduction based on the operating status data and the spatial topology matrix of IT equipment, obtain dynamic prediction data of IT equipment, and obtain dynamic prediction curve of IT equipment status based on the dynamic prediction data of IT equipment.
[0009] If the dynamic prediction curve of the IT equipment status detects an anomaly, the spatiotemporal correlation weight matrix is processed starting from the abnormal IT equipment. A fault propagation path is generated through reinforcement learning algorithm, and the root cause of the IT equipment failure is determined by combining a multivariate fault knowledge graph and a causal graph model.
[0010] In one of the optional technical solutions, the step of acquiring the operating status data of IT equipment, performing spatiotemporal joint analysis and dimensionality reduction based on the operating status data and the spatial topology matrix of the IT equipment to obtain dynamic prediction data of the IT equipment, and obtaining a dynamic prediction curve of the IT equipment status based on the dynamic prediction data of the IT equipment includes:
[0011] Obtain the operating status data of IT equipment, and perform multi-source feature fusion on the operating status data to obtain fused data;
[0012] By performing spatiotemporal joint analysis on the fused data and the spatial topology matrix of the IT equipment, a spatiotemporal correlation weight matrix of IT equipment with time-series weights is obtained;
[0013] The dimensionality of the fused data is reduced to obtain initial prediction data of IT equipment status that includes uncertainty estimation;
[0014] The initial prediction data of IT equipment status is combined with the spatiotemporal correlation weight matrix of IT equipment to obtain dynamic prediction data of IT equipment.
[0015] Based on the initial prediction data of the IT equipment status, a dynamic prediction curve of the IT equipment status that is dynamically adjusted over time is obtained through the analysis of the corrected confidence interval.
[0016] In one of the optional technical solutions, if the dynamic prediction curve of the IT equipment status detects an anomaly, the spatiotemporal correlation weight matrix is processed starting from the abnormal IT equipment, a fault propagation path is generated through a reinforcement learning algorithm, and the root cause of the IT equipment failure is determined by combining a multivariate fault knowledge graph and a causal graph model, including:
[0017] In response to the detection of an abnormal IT equipment status by the dynamic prediction curve of the IT equipment status, starting from the IT equipment with the abnormal status, the business importance factor of the IT equipment is introduced into the spatiotemporal correlation weight matrix of the IT equipment and the edge weight is normalized. A fault propagation path containing the cumulative risk value and business impact is generated by the reinforcement learning path search algorithm.
[0018] Construct a multi-dimensional fault knowledge graph that integrates historical fault cases, equipment dependencies, and expert experience, and model the knowledge graph using a causal graph model;
[0019] By combining the topology of the fault propagation path, the causal effect value of each potential root cause is calculated through counterfactual reasoning, and the root cause of the IT equipment failure with the highest probability is determined based on the causal effect value.
[0020] In one of the optional technical solutions, the step of acquiring the physical location data and real-time environmental data of IT equipment, and constructing a matrix based on the physical location data and the real-time environmental data through a dual topological relationship to obtain an IT equipment spatial topology matrix containing environmental impact factors, includes:
[0021] The temperature, humidity and energy consumption parameters of IT equipment are acquired in real time to form real-time environmental data, and the signal transmission delay data of IT equipment is measured based on the real-time environmental data.
[0022] The physical location data is transformed into three-dimensional coordinate data through spatial vector field modeling; the three-dimensional coordinate data and the signal transmission delay data of the IT equipment are subjected to dual topological relationship mapping, including physical adjacency topology and signal transmission topology, to obtain the spatial coordinate matrix of the IT equipment.
[0023] By performing data correlation analysis, including physical adjacency analysis and device status relationship analysis, on the spatial coordinate matrix of the IT devices using graph convolutional networks, an IT device spatial topology matrix containing environmental impact factors is obtained.
[0024] In one of the optional technical solutions, the specific conditions for determining the abnormal status of the IT equipment are as follows:
[0025] If the actual operating status index value of any IT device exceeds the confidence interval of the corresponding time point in the dynamic prediction curve of the IT device status for A consecutive sampling points, and the association weight of the device with at least B neighboring devices in the spatiotemporal association weight matrix exceeds the dynamic threshold, then the IT device is determined to be in an abnormal state, where A and B are both preset positive integers.
[0026] In one of the optional technical solutions, the step of obtaining a dynamically adjusted IT equipment status dynamic prediction curve over time based on the initial IT equipment status prediction data and through modified confidence interval analysis includes:
[0027] Retrieve the set of associated IT devices;
[0028] An anomaly propagation sample set is constructed based on historical anomaly events. The influence coefficient of the state deviation of the neighboring IT devices on the target IT device is obtained by model training based on the anomaly propagation sample set.
[0029] The state difference is calculated between the actual state value of the associated IT device set and the initial predicted state data of the IT devices to obtain the state prediction deviation.
[0030] Based on the state prediction deviation and the influence coefficient, the initial state prediction data of the IT equipment is corrected by neighborhood propagation to obtain the corrected state prediction value.
[0031] Uncertainty estimation sampling is performed on the corrected state prediction values to obtain a set of predicted state probability distributions;
[0032] Based on the set of predicted state probability distributions, statistical analysis is performed at time points to determine the standard deviation data.
[0033] Confidence intervals are constructed based on the standard deviation data and the corrected state prediction values to form the dynamic prediction curve of the IT equipment state.
[0034] In one optional technical solution, in response to the detection of an IT equipment status anomaly by the dynamic prediction curve of the IT equipment status, starting from the IT equipment with the anomaly, an IT equipment business importance factor is introduced into the spatiotemporal correlation weight matrix of the IT equipment and edge weight normalization is performed. A fault propagation path containing cumulative risk value and business impact is generated through a reinforcement learning path search algorithm, including:
[0035] Define the IT equipment business importance factor as (business priority * traffic share) / (equipment redundancy + 1) and calculate the IT equipment business importance factor.
[0036] The edge weights in the spatiotemporal correlation matrix of the IT equipment are normalized.
[0037] Based on the spatiotemporal association weight matrix of the IT equipment, the path distance of the IT equipment is determined after edge weight normalization.
[0038] Build or update the dynamic propagation path map of IT devices based on the path distance of the IT devices;
[0039] The cumulative risk value is calculated using the formula: Cumulative Risk Value = (Cumulative Weight * IT Equipment Business Importance Factor).
[0040] Select an IT device with an abnormal status as the source node, and perform risk path search and sorting on the dynamic propagation path map of the IT device based on the cumulative risk value to obtain the path node sequence, cumulative weight value and path risk level.
[0041] A fault propagation path is generated based on the path node sequence, cumulative weight value, and path risk level.
[0042] In one of the optional technical solutions, the construction of a multivariate fault knowledge graph integrating historical fault cases, equipment dependencies, and expert experience, the modeling of the knowledge graph using a causal graph model, the calculation of the causal effect value of each potential root cause through counterfactual reasoning based on the topological structure of the fault propagation path, and the determination of the IT equipment fault root cause with the highest probability based on the causal effect value, includes:
[0043] A multi-dimensional fault knowledge graph is constructed based on node type and edge type. The node type includes device entity, fault mode, alarm event and business indicator. The edge type includes causal relationship, dependency relationship and similarity relationship.
[0044] The multivariate fault knowledge graph is embedded and learned using a causal graph model. The causal effects of each potential root cause are calculated through intervention analysis to obtain fault similarity data. The fault similarity data includes historical rule confidence and real-time state similarity.
[0045] Based on the propagation topology represented by the fault propagation path and the fault similarity data, the root cause hypothesis is verified by counterfactual reasoning, and the node with the highest causal effect value and that satisfies the topology propagation logic is selected as the root cause of the IT equipment failure.
[0046] The present invention provides an electronic device, including a memory, a processor, and an electronic device program on the memory, wherein the processor executes the electronic device program to implement any of the steps of the aforementioned root cause analysis method based on an IT operation and maintenance system.
[0047] The present invention provides an electronic device readable storage medium storing an electronic device program / instruction thereon, characterized in that, when the electronic device program / instruction is executed by a processor, it implements any of the steps of the aforementioned root cause analysis method based on an IT operation and maintenance system.
[0048] The above technical solution has the following beneficial effects:
[0049] The root cause analysis method based on IT operations and maintenance systems provided by this invention constructs a spatial topology matrix of IT devices based on their physical locations. Then, it performs spatiotemporal joint analysis of the operational status data of the IT devices with this spatial topology matrix to obtain dynamic predictive data for the IT devices. This dynamic predictive data is then used to identify IT devices exhibiting abnormal states. Starting from these abnormal devices, the method determines the fault propagation path. Finally, it combines the fault propagation path with a knowledge graph for reasoning analysis. This allows for accurate and timely identification of the root causes of IT device failures without relying on manual analysis, alerting operators and preventing the loss of the optimal time for troubleshooting. Attached Figure Description
[0050] The disclosure of this invention will become more readily understood by referring to the accompanying drawings. It should be understood that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings:
[0051] Figure 1 A flowchart illustrating the root cause analysis method based on an IT operations and maintenance system provided in an embodiment of the present invention;
[0052] Figure 2 A flowchart of a root cause analysis method based on an IT operations and maintenance system provided in another embodiment of the present invention;
[0053] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0054] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0055] like Figure 1 The image shows a root cause analysis method based on an IT operations and maintenance system provided by an embodiment of the present invention, comprising the following steps:
[0056] Step S101: Obtain the physical location data and real-time environmental data of the IT equipment, and construct a matrix based on the physical location data and the real-time environmental data through dual topological relationships to obtain the spatial topology matrix of the IT equipment containing environmental impact factors.
[0057] Step S102: Obtain the operating status data of the IT equipment, perform spatiotemporal joint analysis and dimensionality reduction based on the operating status data and the spatial topology matrix of the IT equipment to obtain dynamic prediction data of the IT equipment, and obtain the dynamic prediction curve of the IT equipment status based on the dynamic prediction data of the IT equipment.
[0058] Step S103: If the dynamic prediction curve of the IT equipment status detects an anomaly, the spatiotemporal correlation weight matrix is processed starting from the abnormal IT equipment, a fault propagation path is generated through reinforcement learning algorithm, and the root cause of the IT equipment failure is determined by combining the multivariate fault knowledge graph and causal graph model.
[0059] Preferably, step S101 includes the following specific steps:
[0060] Obtain the operating status data of IT equipment, and perform multi-source feature fusion on the operating status data to obtain fused data;
[0061] By performing spatiotemporal joint analysis on the fused data and the spatial topology matrix of the IT equipment, a spatiotemporal correlation weight matrix of IT equipment with time-series weights is obtained;
[0062] The dimensionality of the fused data is reduced to obtain initial prediction data of IT equipment status that includes uncertainty estimation;
[0063] The initial prediction data of IT equipment status is combined with the spatiotemporal correlation weight matrix of IT equipment to obtain dynamic prediction data of IT equipment.
[0064] Based on the initial prediction data of the IT equipment status, a dynamic prediction curve of the IT equipment status that is dynamically adjusted over time is obtained through the analysis of the corrected confidence interval.
[0065] Step S102 includes the following specific steps:
[0066] In response to the detection of an abnormal IT equipment status by the dynamic prediction curve of the IT equipment status, starting from the IT equipment with the abnormal status, the business importance factor of the IT equipment is introduced into the spatiotemporal correlation weight matrix of the IT equipment and the edge weight is normalized. A fault propagation path containing the cumulative risk value and business impact is generated by the reinforcement learning path search algorithm.
[0067] Construct a multi-dimensional fault knowledge graph that integrates historical fault cases, equipment dependencies, and expert experience, and model the knowledge graph using a causal graph model;
[0068] By combining the topology of the fault propagation path, the causal effect value of each potential root cause is calculated through counterfactual reasoning, and the root cause of the IT equipment failure with the highest probability is determined based on the causal effect value.
[0069] In step S101, after acquiring the physical location data and real-time environmental data of the TI device, the physical location data is converted into three-dimensional coordinate data. Real-time temperature, humidity, and energy consumption parameters of the IT device are acquired to form real-time environmental data. Based on the real-time environmental data, the signal transmission delay data of the IT device is measured. A topological relationship mapping is performed on the three-dimensional coordinate data through spatial vector field modeling. A dual topological relationship mapping, including physical adjacency topology and signal transmission topology, is performed on the three-dimensional coordinate data and the IT device signal transmission delay data to obtain the IT device spatial coordinate matrix. Then, data correlation analysis can be performed on the IT device spatial coordinate matrix to obtain the IT device spatial topology matrix. Data correlation analysis can include physical adjacency analysis and device state relationship analysis, etc.
[0070] In cases involving multiple data centers and multiple IT devices in each data center, it is necessary to analyze the physical location data of each IT device. On the one hand, it is necessary to locate the physical location of the IT device to obtain its relevant position in the IT operation and maintenance system. On the other hand, it is necessary to construct a spatial topology matrix based on the relevant position of the IT device, that is, to construct a dual topology relationship, so as to realize the spatial positioning and positional relationship analysis of the IT device, thereby improving the efficiency of root cause analysis of IT devices.
[0071] Physical location data includes the latitude and longitude of the data center and the specific location of the IT equipment within it. The three-dimensional coordinates are defined as follows: X-axis represents the data center column number, Y-axis represents the rack row number, and Z-axis represents the equipment's unit height (U-axis). This physical location data is used to determine the coordinates of the equipment in three-dimensional space, providing a basis for topology analysis. The three-dimensional coordinates of the IT equipment can be calculated using wireless positioning technology or WiFi signal fingerprinting by accurately measuring the time difference and signal strength between the equipment and multiple base stations. Each IT device is assigned a unique three-dimensional coordinate value based on its actual location within the data center.
[0072] Preferably, the method for constructing the spatial topology matrix of IT equipment can be as follows:
[0073] The physical location of IT devices is obtained through laser point cloud scanning technology in three dimensions, and a dynamic coordinate mapping is constructed by combining it with real-time positioning data. Then, spatial vector field analysis is performed on the three-dimensional coordinate data to calculate the physical adjacency between IT devices based on Euclidean distance and the signal transmission adjacency based on latency jitter. The dual adjacency matrix is then fused using a graph convolutional network to generate a spatial topology matrix of IT devices that includes the influence of environmental temperature factors.
[0074] Because existing technologies generally rely on static predictions for IT equipment status, it is difficult to dynamically adjust key thresholds. Step S102 addresses this by constructing a spatiotemporal joint analysis model for IT equipment, enabling dynamic prediction of IT equipment data. This improves the dynamic analysis capability of IT equipment status and avoids maintenance errors caused by delayed early warnings.
[0075] Operational status data includes log events, performance metrics, configuration change records, etc. When extracting multimodal features, semantic encoding can be performed on the log text, followed by extraction of temporal features from performance metrics, and graph embedding algorithms can be used to extract dependency features from configuration changes. Then, a cross-modal attention mechanism is used to fuse these three types of features to generate a unified feature vector, which is then captured by a Transformer encoder to capture long-term temporal dependencies.
[0076] Preferably, step S102 performs Gaussian kernel function operation on the spatial topology matrix of IT equipment to determine the physical location correlation matrix, then converts the operating state data into a state matrix, and performs standardization operation on the state matrix to obtain a standardized state matrix. Then, cosine similarity operation is performed on the standardized state matrix to determine the equipment state correlation matrix. Then, the physical location correlation matrix and the equipment state correlation matrix are fused to obtain the IT equipment spatiotemporal correlation weight matrix. Finally, the standardized state matrix is subjected to fusion network processing to obtain IT equipment dynamic prediction data including the initial prediction data of IT equipment state and the IT equipment spatiotemporal correlation weight matrix.
[0077] Cosine similarity calculation is used to calculate the degree of similarity between the state change trends of two devices. The state of each device is regarded as a vector, and the cosine value of the angle between the vectors is calculated to accurately capture the coordinated changes in the state of the devices, ignoring the differences in absolute values, so as to determine whether IT devices are resonating at the same frequency.
[0078] In step S102, firstly, a set of neighboring devices associated with the abnormal device is obtained, the deviation between their actual state values and the initial predicted data is calculated, and the predicted values are corrected through a neighborhood propagation mechanism. Subsequently, the corrected predicted values are sampled to obtain a set of predicted state probability distributions. Based on a set confidence level, the mean and standard deviation of each time point are statistically analyzed to construct a confidence interval, and finally, a dynamic prediction curve for the IT device status is generated.
[0079] To quantify the error of the initial prediction data of IT equipment status, confidence intervals are constructed and corrected based on the initial prediction data of IT equipment status, thereby obtaining a dynamic prediction curve of IT equipment status for real-time status monitoring, so as to realize IT equipment status prediction and improve the response speed of IT equipment status prediction.
[0080] The confidence interval is the reasonable fluctuation range of the dynamic prediction value. It can be obtained by calculating the 95% probability boundary of each time point based on the prediction distribution of Monte Carlo sampling. The confidence interval can replace the fixed threshold to realize the adaptive alarm of the IT operation and maintenance system. During peak hours, the threshold is automatically relaxed to prevent false alarms, and during off-peak hours, the threshold is tightened to prevent missed alarms.
[0081] In step S103, when the actual index values of multiple consecutive sampling points of any device exceed the corresponding confidence interval, the device is determined to be in an abnormal state. The spatiotemporal correlation weight matrix is normalized and converted into device path distances, and an IT device dynamic propagation path graph is constructed or updated accordingly. Using the abnormal device as the source node, an algorithmic or heuristic risk path search and ranking is employed to filter and rank the top K propagation paths, extracting the path node sequence, cumulative weight value, and path risk level. An IT device business importance factor can be introduced into the IT device spatiotemporal correlation weight matrix, and edge weights can be normalized. A reinforcement learning path search algorithm can then be used to generate a fault propagation path containing cumulative risk values and business impact.
[0082] Cumulative risk value,
[0083] In step S103, a multivariate fault knowledge graph integrating historical fault cases, device dependencies, and expert experience is first constructed. A causal graph model is then used to model the knowledge graph, constructing a multivariate fault knowledge graph covering physical connection links, network topology, and service call chains. A graph neural network is used to embed the knowledge graph, obtaining fault similarity data such as historical rule confidence and real-time state similarity. Then, the causal graph model is used to model the knowledge graph, and combined with the topological structure of the fault propagation path, counterfactual reasoning is used to calculate the causal effect value of each potential root cause. Based on the causal effect value, the root cause of the IT equipment fault with the highest probability is determined. The fault propagation path topology and fault similarity data can be input into the root cause confidence analysis module for comprehensive calculation to obtain the causal effect value, thereby determining the root cause of the IT equipment fault.
[0084] The multivariate fault knowledge graph is a database that stores historical fault relationships, providing an experience base for fault reasoning, similar to a "case study database." It uses a graph structure, with nodes representing IT equipment types and fault types, and edges representing causal relationships, such as service crashes caused by disk exhaustion. Employing a multivariate fault knowledge graph digitizes operational experience, avoiding excessive reliance on operators' personal experience for analyzing the root causes of IT equipment failures.
[0085] Root cause analysis of IT equipment failures is a systematic approach that aims to identify the root causes of IT equipment failures, thereby eliminating or controlling these root causes, preventing the problems from recurring, and gradually improving the IT operations and maintenance system.
[0086] Therefore, IT operations and maintenance systems can automate the analysis and tracking of IT equipment faults. Using this method, only the physical location and operational status data of each IT device need to be determined to construct a dynamic prediction curve for the IT device's status. Then, during operation, the operational status of the IT devices is continuously monitored. By comparing the dynamic prediction curve with the operational status data in real time, IT devices exhibiting abnormal states can be identified promptly. After identifying the abnormal IT device, its physical location and operational status data can be collected specifically. Starting from this abnormal IT device, the fault propagation path within the IT operations and maintenance system can be determined. By combining this fault propagation path with a multivariate fault knowledge graph for inference and analysis, the root cause of the IT device fault can be traced.
[0087] In summary, the root cause analysis method based on IT operation and maintenance systems provided by this invention constructs a spatial topology matrix of IT devices based on their physical locations. Then, it performs spatiotemporal joint analysis of the operational status data of the IT devices with this spatial topology matrix to obtain dynamic predictive data for the IT devices. This dynamic predictive data is then used to identify IT devices exhibiting abnormal states. Starting from these abnormal IT devices, a fault propagation path is determined. Finally, the fault propagation path is analyzed using a knowledge graph, enabling accurate and timely identification of the root cause of IT device failures without relying on manual analysis. This alerts operators and prevents missing the optimal time for troubleshooting IT devices.
[0088] In one embodiment, such as Figure 2 As shown, it also includes:
[0089] Step S201: Based on the root cause of IT equipment failure, determine the IT equipment failure countermeasures.
[0090] Step S202: Optimize the IT equipment failure countermeasures to obtain IT equipment failure optimization countermeasures.
[0091] Step S203: Update the analysis rules of spatiotemporal joint analysis based on the IT equipment failure optimization strategy.
[0092] In this embodiment, step S201, based on the root cause analysis results obtained in step S103, first retrieves standardized processing solutions corresponding to the fault type or fault mode from a pre-built fault countermeasure library. This countermeasure library contains handling procedures and recovery steps for various root cause scenarios such as common hardware failures, network congestion, broken service call chains, and configuration errors. During the retrieval process, the system prioritizes candidate countermeasures based on the confidence data of the fault root cause and historical processing effect scores, selecting the countermeasure with the highest comprehensive score as the initial fault countermeasure.
[0093] Step S202 involves simulating and validating the initial fault mitigation measures in a controlled or sandbox environment, and evaluating the actual effectiveness of the measures using real-time operational data such as device resource utilization, network latency, and service response time. During the simulation, key parameters, including traffic rate limiting thresholds, retry counts, and timeout windows, are automatically adjusted to maximize fault recovery efficiency and minimize service impact. Through multiple iterations and by combining test results and historical feedback data, the system generates optimized fault mitigation measures, including specific parameter configurations, execution procedures, and security checks, to ensure the feasibility and robustness of the measures in a real production environment.
[0094] In step S203, the optimized fault response strategy is not only used for on-site recovery, but its core parameters and strategy elements are also mapped back to the spatiotemporal joint analysis model and rule engine to continuously improve prediction accuracy and response speed. Specifically, the system synchronously writes adjusted parameters such as neighborhood propagation influence radius, confidence interval confidence threshold, and path search risk weight into the analysis rule base, triggering rule reloading or model incremental training processes, thereby achieving closed-loop adaptive updates of the rule engine and deep learning model. In subsequent fault prediction and root cause analysis, the IT operations and maintenance system will automatically incorporate the latest optimized strategies to further reduce the false alarm rate and improve the accuracy of root cause localization and countermeasure recommendations.
[0095] In one embodiment, step S101 includes the following sub-steps:
[0096] The system acquires real-time temperature, humidity, and energy consumption parameters of IT equipment to form real-time environmental data, and measures the signal transmission delay data of IT equipment based on the real-time environmental data.
[0097] Physical location data is transformed into three-dimensional coordinate data through spatial vector field modeling; a dual topological relationship mapping, including physical adjacency topology and signal transmission topology, is performed on the three-dimensional coordinate data and IT equipment signal transmission delay data to obtain the IT equipment spatial coordinate matrix.
[0098] By performing data correlation analysis, including physical adjacency analysis and device status relationship analysis, on the spatial coordinate matrix of IT devices using graph convolutional networks, an IT device spatial topology matrix containing environmental impact factors is obtained.
[0099] In this embodiment, after acquiring the physical location data and real-time environmental data of IT devices, the physical location of the IT devices is stably and accurately converted into three-dimensional coordinates, and the signal transmission delay data of the IT devices is measured based on the real-time environmental data. Based on the three-dimensional coordinates and the signal transmission delay data of the IT devices, a dual topological mapping, including physical adjacency topology and signal transmission topology, is performed on the three-dimensional coordinate data and the IT devices' signal transmission delay data, thus accurately constructing the IT devices' spatial coordinate matrix. Data correlation analysis is performed on the IT devices' spatial coordinate matrix, including physical adjacency analysis and device status relationship analysis. Specifically, physical adjacency analysis is used to analyze the adjacency relationships of IT devices in the rack area and the data center area respectively. Device status relationship analysis is performed through the logical connection strength of the devices, integrating the actual traffic, maximum bandwidth, and network hop count of the IT devices to obtain the IT devices' spatial topology matrix.
[0100] The IT equipment spatial coordinate matrix is a table or matrix that organizes the three-dimensional coordinates of all IT equipment. It converts the physical location into a computer-calculate digital form. Each row represents one device, and the three columns store the coordinate values of the server room column, rack row, and U-position height, respectively. This maps the physical world to the digital space to facilitate subsequent mathematical analysis.
[0101] The IT device spatial topology matrix is a digital model representing the relationship network of IT devices, generated by combining physical adjacency and state relationships. By adding the weights of physical adjacency and state relationships, an N×N matrix is generated, where N is the total number of devices. This matrix is used to quantify the comprehensive correlation strength between devices, transforming complex device relationships into mathematical objects that can be processed by computers.
[0102] Physical adjacency analysis specifically quantifies the physical distance between IT devices to determine whether devices may affect each other due to physical proximity.
[0103] In one embodiment, step S102 includes the following sub-steps:
[0104] Perform Gaussian kernel function operation on the spatial topology matrix of IT equipment to determine the physical location correlation matrix.
[0105] The running state data is transformed into a state matrix, and the state matrix is standardized by Z-score to obtain a standardized state matrix.
[0106] The device state correlation matrix is obtained by performing cosine similarity calculation on the standardized state matrix.
[0107] By fusing the physical location correlation matrix and the device status correlation matrix, the spatiotemporal correlation weight matrix of IT devices is obtained.
[0108] The standardized state matrix is processed by a fusion network to obtain initial prediction data of IT equipment state.
[0109] The steps in this embodiment can be implemented using a spatiotemporal joint analysis model for IT devices. This model includes a temporal convolutional layer (TCN), an LSTM unit, a spatial attention layer, and a GRU unit. Specifically, the temporal convolutional layer extracts the periodic and trend characteristics of device states; the LSTM unit captures long-term dependencies between IT devices, such as daily business peaks or troughs; the spatial attention layer quantifies the correlation strength between devices; and the GRU unit combines spatial weights to adjust predictions and capture short-term mutations.
[0110] To improve the robustness of dynamic analysis, it is necessary to train and optimize the spatiotemporal joint analysis model of IT equipment. After the model has been trained and optimized, in practical applications, a Gaussian kernel function operation needs to be performed on the spatial topology matrix of IT equipment to determine the physical location correlation matrix (N*3).
[0111] Then, the operational status data is transformed into a state matrix, and the state matrix is Z-score standardized to obtain a standardized state matrix (N*T*M). The state correlation degree of the standardized state matrix is calculated based on the cosine similarity of the sliding time window, yielding a physical location correlation matrix. This physical location correlation matrix is then fused with the physical location correlation matrix to obtain the IT device spatiotemporal correlation weight matrix (N*N). The IT device spatiotemporal correlation weight matrix represents the physical location correlation between devices and the state influence relationship between IT devices.
[0112] The physical location correlation matrix is fused with the physical location correlation matrix. The spatial attention layer calculates the spatiotemporal correlation weights between devices. The weights are fused with the Gaussian similarity of physical location and the cosine correlation of operating status to achieve joint analysis of the physical location and device status of IT devices.
[0113] Finally, the standardized state matrix is processed by an LSTM-GRU fusion network to generate an initial prediction curve, which characterizes the initial predicted state changes of the IT device.
[0114] Among them, the Gaussian kernel function operation is used to transform the "distance" in the spatial topology matrix of IT devices into "association degree", thereby smoothly quantifying the physical location relationship.
[0115] Z-score standardization transforms device status data into a standard form with a mean of 0 and a standard deviation of 1. This eliminates differences in the metric units of different IT devices, allowing for fair calculation of similarity for different metrics such as CPU, memory, and traffic.
[0116] LSTM-GRU fusion network processing is a combination of two deep learning models. It uses the output of LSTM as the input of GRU and combines long-term trends with short-term fluctuations to predict the future state of IT equipment. Based on historical data, it predicts the future values of indicators such as CPU usage and memory usage of IT equipment, which significantly advances the prediction time of abnormal states of IT equipment.
[0117] In one embodiment, step S103 includes the following sub-steps:
[0118] Retrieve the set of associated IT devices.
[0119] An anomaly propagation sample set is constructed based on historical anomaly events. The influence coefficient of the state deviation of neighboring IT devices on the target IT device is obtained through model training based on the anomaly propagation sample set.
[0120] The state difference between the actual state values of the associated IT device set and the initial predicted state data of the IT devices is calculated to obtain the state prediction deviation.
[0121] Based on the state prediction bias and the influence coefficient, the initial state prediction data of IT equipment is corrected by neighborhood propagation to obtain the corrected state prediction value.
[0122] Uncertainty estimation sampling is performed on the corrected state prediction values to obtain the set of predicted state probability distributions.
[0123] Based on the set of predicted state probability distributions, statistical analysis is performed at time points to determine the standard deviation data.
[0124] Confidence intervals are constructed based on standard deviation data and corrected state prediction values to form dynamic prediction curves for IT equipment status.
[0125] In this embodiment, the set of predicted state probability distributions is obtained through Monte Carlo Dropout sampling, and the dynamic prediction curve of IT equipment status includes time series, prediction mean vector, and confidence interval matrix.
[0126] Monte Carlo Dropout sampling is performed on the initial predicted values output by the spatiotemporal joint analysis model in step S102. Preferably, during forward propagation, 15%-25% of neurons in the fully connected layers of the model are randomly dropped. The prediction is repeated 45-55 times to generate a set of predicted state probability distributions. Then, the statistics at each time point in the predicted state probability distribution set are analyzed, and the mean and standard deviation of the statistics at each time point are used to construct confidence intervals. The quantiles of the standard normal distribution are determined, and the confidence level is set to 90%-95%. Finally, the above analysis outputs a dynamic prediction curve of IT equipment status with confidence intervals. The data structure in the dynamic prediction curve of IT equipment status includes time series, predicted mean vector, and confidence interval matrix, thereby enabling accurate and timely prediction of IT equipment data fluctuations.
[0127] Monte Carlo Dropout sampling is a technique for assessing prediction uncertainty. For example, it involves randomly shutting down some nodes in a neural network during prediction and repeating this process 50 times to obtain the distribution of predicted values. In this invention, it is used to quantify the reliability of prediction results, generate confidence intervals, and avoid the risk of misjudgment from a single prediction.
[0128] In one embodiment, the specific condition for determining the abnormal status of the IT equipment in step S103 is as follows:
[0129] If the actual operating status index value of any IT device exceeds the confidence interval of the corresponding time point in the dynamic prediction curve of the IT device status for A consecutive sampling points, and the association weight of the device with at least B neighboring devices in the spatiotemporal association weight matrix exceeds the dynamic threshold, then the IT device is determined to be in an abnormal state, where A and B are both preset positive integers.
[0130] The dynamic threshold is adaptively determined by the associated weight distribution of historical abnormal periods.
[0131] Furthermore, step S103 also includes:
[0132] Define the IT equipment business importance factor as (business priority * traffic share) / (equipment redundancy + 1) to calculate the IT equipment business importance factor.
[0133] The edge weights in the spatiotemporal correlation matrix of IT equipment are normalized.
[0134] Based on the spatiotemporal correlation weight matrix of IT devices, the path distance of IT devices is determined after edge weight normalization.
[0135] Build or update the dynamic propagation path map of IT devices based on the path distance of IT devices.
[0136] The cumulative risk value is obtained by calculating the cumulative risk value as (cumulative weight * IT equipment business importance factor).
[0137] Select the IT device with abnormal status as the source node, and search and sort the risk path on the dynamic propagation path map of the IT device based on the cumulative risk value to obtain the path node sequence, cumulative weight value and path risk level.
[0138] A fault propagation path is generated based on the path node sequence, cumulative weight value, and path risk level.
[0139] In this embodiment, edge weights are normalized based on the spatiotemporal correlation weight matrix of IT devices to determine the path distance of IT devices. Based on the path distance, a dynamic propagation path graph of IT devices is obtained through dynamic updates of the propagation path. Risk sequence processing is then applied to the dynamic propagation path graph to obtain the fault propagation path. The fault propagation path includes a path node sequence, cumulative weight values, and path risk levels.
[0140] First, the spatiotemporal correlation weight matrix of IT devices is received, where matrix elements represent the physical location correlation and operational status correlation between devices. Then, IT devices are mapped as graph nodes. If the operational status correlation value is greater than or equal to a preset value, weighted edges are established between nodes, with the operational status correlation value set as the weight. The edge weights are then normalized and converted into IT device path distances. Next, after obtaining the IT device path distances, the spatiotemporal correlation weight matrix of IT devices is periodically recalculated. Edges whose weights decrease to below the preset value are removed, and edges whose weights increase to above the preset value are added. Specific adjustments are made according to actual needs. Finally, risk sequence processing is performed on the dynamic propagation path graph of IT devices to obtain fault propagation paths.
[0141] Preferably, alarm device nodes can also be input, and the Dijkstra algorithm can be used to search for the top K shortest propagation paths, calculating the propagation risk value for each path. Then, the path list is output in descending order of risk value, with each path containing a sequence of path nodes, a cumulative weight value, and a path risk level. Finally, for high-risk paths, a set of critical edges is identified, and based on this set, the effectiveness of the blocking strategy is calculated to generate a propagation blocking strategy to determine the fault propagation path.
[0142] Preferably, in the risk path search and ranking based on the cumulative risk value on the dynamic propagation path map of IT equipment, a reward function can be used for path search. The reward function can be designed as follows:
[0143] R = -(cumulative weight + α * business impact) + β * δ (root cause node);
[0144] Where α and β are weighting coefficients, respectively, and δ (root cause node) is the root cause determination indicator function.
[0145] The propagation path refers to the logical or physical path through which a fault or state change spreads within an IT equipment network. Path discovery is driven by a spatiotemporal weight matrix, and dynamic optimization of the propagation path is used to determine the fault propagation path. By analyzing the fault propagation path of IT equipment, the data requirements for root cause analysis of IT equipment can be met, thereby improving the efficiency of root cause analysis.
[0146] Edge weight normalization is used to convert association weights into propagation distances, making the weights usable for shortest path calculations. Higher weights result in shorter distances, thus converting association strength into computable path distances.
[0147] Dijkstra's algorithm is a computer science algorithm used to find the shortest path. It explores adjacent nodes step by step, always choosing the current shortest path to extend. Starting from the alarm source, it finds the path most likely to propagate the fault, thereby automatically deriving the fault propagation chain, which is far more efficient than manual troubleshooting.
[0148] The path risk level is a quantification of the threat level of a propagation path. It can remind the IT operations and maintenance system to prioritize paths with high risk levels and help the IT operations and maintenance system quickly focus on the core fault chain to improve troubleshooting efficiency.
[0149] In one embodiment, step S103 includes the following sub-steps:
[0150] A multi-dimensional fault knowledge graph is constructed based on node type and edge type. Node types include device entities, fault modes, alarm events, and business metrics, while edge types include causal relationships, dependency relationships, and similarity relationships.
[0151] The fault similarity data is obtained by embedding a multivariate fault knowledge graph into a causal graph model and calculating the causal effects of each potential root cause through intervention analysis. The fault similarity data includes historical rule confidence and real-time state similarity.
[0152] Based on the propagation topology represented by the fault propagation path and the fault similarity data, the root cause hypothesis is verified through counterfactual reasoning, and the node with the highest causal effect value and that satisfies the topological propagation logic is selected as the root cause of the IT equipment failure.
[0153] In this embodiment, firstly, by extracting device configuration information, physical connection relationships and logical dependencies are established. The relationship types include power supply links, network topology, and service call chains. Historical alarm data is analyzed to generate a causal association rule base in order to construct a multi-dimensional fault knowledge graph.
[0154] This embodiment achieves dynamic fusion of real-time status and historical rules through a similarity matrix, forming a comprehensive analysis of the fault location method. It effectively solves the technical problems of existing technologies that rely on human experience and cannot directly locate the root cause of IT equipment.
[0155] After obtaining a multivariate fault knowledge graph, fault similarity data can be obtained through embedding learning. Then, by combining the multivariate fault knowledge graph and fault similarity data for root cause confidence analysis, the root cause of IT equipment failure can be determined. This enables the IT operations and maintenance system to take targeted actions based on the root cause of IT equipment failure, such as configuring countermeasures, optimizing countermeasures, and analyzing rules.
[0156] Among them, graph neural networks are artificial intelligence models specifically designed to process graph-structured data. By learning patterns in the graph through information transmission between neighbor nodes, they can intelligently match the fault modes of current IT equipment in a multi-dimensional fault knowledge graph.
[0157] Root cause confidence is a numerical value representing the credibility of root cause judgments. It is used to assist operations and maintenance personnel in making decisions and can be obtained by comprehensively calculating historical rule confidence and real-time state similarity.
[0158] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0159] like Figure 3 The diagram shows a hardware structure of an electronic device according to the present invention, including a memory 302, a processor 301, and an electronic device program on the memory 302. The processor 301 executes the electronic device program to implement the steps of the root cause analysis method based on the IT operation and maintenance system in any of the above embodiments.
[0160] Figure 3 Take processor 301 as an example.
[0161] The electronic device may also include an input device 303 and a display device 304.
[0162] The processor 301, memory 302, input device 303 and display device 304 can be connected by a bus or other means. The figure shows an example of connection by bus.
[0163] The memory 302, as a non-volatile electronic device readable storage medium, can be used to store non-volatile software programs, non-volatile electronic device executable programs, and modules, such as the program instructions / modules corresponding to the root cause analysis method based on the IT operation and maintenance system in the embodiments of this application. The processor 301 executes various functional applications and data processing by running the non-volatile software programs, instructions, and modules stored in the memory 302, thereby implementing the root cause analysis method based on the IT operation and maintenance system in the above embodiments.
[0164] Memory 302 may include a stored program area and a stored data area. The stored program area may store the operating system and applications required for at least one function; the stored data area may store data created based on the use of the root cause analysis method based on the IT operations system. Furthermore, memory 302 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 302 may optionally include memory remotely located relative to processor 301, and these remote memories may be connected via a network to means of performing the root cause analysis method based on the IT operations system. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0165] Input device 303 can receive user clicks and generate signal inputs related to user settings and function controls based on the root cause analysis method of the IT operations and maintenance system. Display device 304 may include display devices such as a display screen.
[0166] When one or more modules are stored in the memory 302, and are run by one or more processors 301, the root cause analysis method based on the IT operation and maintenance system in any of the above method embodiments is executed.
[0167] The electronic device disclosed in this invention, when in operation, can execute all the steps of the root cause analysis method based on the IT operation and maintenance system described above. It constructs an IT device spatial topology matrix based on the physical location of the IT devices, and then performs spatiotemporal joint analysis between the IT device's operational status data and the IT device spatial topology matrix to obtain dynamic prediction data for the IT devices. It then uses this dynamic prediction data to identify IT devices exhibiting abnormal states, determines the fault propagation path starting from these abnormal IT devices, and finally performs reasoning analysis on the fault propagation path using a knowledge graph. This allows for accurate and timely determination of the root cause of IT device failures without relying on manual analysis.
[0168] One embodiment of the present invention provides an electronic device readable storage medium storing an electronic device program / instruction, which, when executed by a processor 301, implements all the steps of the root cause analysis method based on an IT operation and maintenance system as described above.
[0169] In the context of this disclosure, a storage medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The storage medium can be a machine-readable signal medium or a machine-readable storage medium. Optionally, the storage medium can be a non-transitory electronically readable storage medium, such as a ROM, random access memory (RAM), compact disc ROM (CD-ROM), magnetic tape, floppy disk, and optical data storage device.
[0170] One embodiment of the present invention provides an electronic device program product, including an electronic device program / instruction, which, when executed by a processor, implements the steps of the root cause analysis method based on an IT operations and maintenance system as described above.
[0171] By running the aforementioned electronic device program, all steps of the root cause analysis method based on the IT operations and maintenance system described above can be executed. An IT device spatial topology matrix is constructed based on the physical location of the IT devices. Then, spatiotemporal joint analysis is performed between the operational status data of the IT devices and the IT device spatial topology matrix to obtain dynamic predictive data for the IT devices. This dynamic predictive data is then used to identify IT devices exhibiting abnormal states. Starting from these abnormal IT devices, the fault propagation path is determined. Finally, the fault propagation path is analyzed using a knowledge graph, thereby accurately and promptly identifying the root cause of IT device failures without relying on manual analysis.
[0172] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A root cause analysis method based on IT operations and maintenance systems, characterized in that, include: The physical location data and real-time environmental data of IT equipment are acquired. Based on the physical location data and the real-time environmental data, a matrix is constructed through dual topological relationships to obtain a spatial topology matrix of IT equipment that includes environmental impact factors. Acquire the operating status data of IT equipment, perform spatiotemporal joint analysis and dimensionality reduction based on the operating status data and the spatial topology matrix of IT equipment, obtain dynamic prediction data of IT equipment, and obtain dynamic prediction curve of IT equipment status based on the dynamic prediction data of IT equipment. If the dynamic prediction curve of the IT equipment status detects an anomaly, the spatiotemporal correlation weight matrix is processed starting from the abnormal IT equipment. A fault propagation path is generated through reinforcement learning algorithm, and the root cause of the IT equipment failure is determined by combining a multivariate fault knowledge graph and a causal graph model.
2. The root cause analysis method based on IT operation and maintenance system according to claim 1, characterized in that, The process of acquiring operational status data of IT equipment, performing spatiotemporal joint analysis and dimensionality reduction based on the operational status data and the spatial topology matrix of the IT equipment to obtain dynamic prediction data of the IT equipment, and obtaining dynamic prediction curves of the IT equipment status based on the dynamic prediction data of the IT equipment includes: Obtain the operating status data of IT equipment, and perform multi-source feature fusion on the operating status data to obtain fused data; By performing spatiotemporal joint analysis on the fused data and the spatial topology matrix of the IT equipment, a spatiotemporal correlation weight matrix of IT equipment with time-series weights is obtained; The dimensionality of the fused data is reduced to obtain initial prediction data of IT equipment status that includes uncertainty estimation; The initial prediction data of IT equipment status is combined with the spatiotemporal correlation weight matrix of IT equipment to obtain dynamic prediction data of IT equipment. Based on the initial prediction data of the IT equipment status, a dynamic prediction curve of the IT equipment status that is dynamically adjusted over time is obtained through the analysis of the corrected confidence interval.
3. The root cause analysis method based on IT operation and maintenance system according to claim 1, characterized in that, If the dynamic prediction curve of the IT equipment status detects an anomaly, the spatiotemporal correlation weight matrix is processed starting from the abnormal IT equipment. A fault propagation path is generated through a reinforcement learning algorithm, and the root cause of the IT equipment failure is determined by combining a multivariate fault knowledge graph and a causal graph model, including: In response to the detection of an abnormal IT equipment status by the dynamic prediction curve of the IT equipment status, starting from the IT equipment with the abnormal status, the business importance factor of the IT equipment is introduced into the spatiotemporal correlation weight matrix of the IT equipment and the edge weight is normalized. A fault propagation path containing the cumulative risk value and business impact is generated by the reinforcement learning path search algorithm. Construct a multi-dimensional fault knowledge graph that integrates historical fault cases, equipment dependencies, and expert experience, and model the knowledge graph using a causal graph model; By combining the topology of the fault propagation path, the causal effect value of each potential root cause is calculated through counterfactual reasoning, and the root cause of the IT equipment failure with the highest probability is determined based on the causal effect value.
4. The root cause analysis method based on IT operation and maintenance system according to claim 1, characterized in that, The process of acquiring the physical location data and real-time environmental data of IT devices, and constructing a matrix based on the physical location data and the real-time environmental data through a dual topological relationship, yields an IT device spatial topology matrix that includes environmental impact factors, including: The temperature, humidity and energy consumption parameters of IT equipment are acquired in real time to form real-time environmental data, and the signal transmission delay data of IT equipment is measured based on the real-time environmental data. The physical location data is transformed into three-dimensional coordinate data through spatial vector field modeling; the three-dimensional coordinate data and the signal transmission delay data of the IT equipment are subjected to dual topological relationship mapping, including physical adjacency topology and signal transmission topology, to obtain the spatial coordinate matrix of the IT equipment. By performing data correlation analysis, including physical adjacency analysis and device status relationship analysis, on the spatial coordinate matrix of the IT devices using graph convolutional networks, an IT device spatial topology matrix containing environmental impact factors is obtained.
5. The root cause analysis method based on IT operation and maintenance system according to claim 3, characterized in that: The specific conditions for determining the abnormal status of the IT equipment are as follows: If the actual operating status index value of any IT device exceeds the confidence interval of the corresponding time point in the dynamic prediction curve of the IT device status for A consecutive sampling points, and the association weight of the device with at least B neighboring devices in the spatiotemporal association weight matrix exceeds the dynamic threshold, then the IT device is determined to be in an abnormal state, where A and B are both preset positive integers.
6. The root cause analysis method based on IT operation and maintenance system according to claim 2, characterized in that, The step of obtaining a dynamically adjusted IT equipment status dynamic prediction curve over time based on the initial IT equipment status prediction data and through modified confidence interval analysis includes: Retrieve the set of associated IT devices; An anomaly propagation sample set is constructed based on historical anomaly events. The influence coefficient of the state deviation of the neighboring IT devices on the target IT device is obtained by model training based on the anomaly propagation sample set. The state difference is calculated between the actual state value of the associated IT device set and the initial predicted state data of the IT devices to obtain the state prediction deviation. Based on the state prediction deviation and the influence coefficient, the initial state prediction data of the IT equipment is corrected by neighborhood propagation to obtain the corrected state prediction value. Uncertainty estimation sampling is performed on the corrected state prediction values to obtain a set of predicted state probability distributions; Based on the set of predicted state probability distributions, statistical analysis is performed at time points to determine the standard deviation data. Confidence intervals are constructed based on the standard deviation data and the corrected state prediction values to form the dynamic prediction curve of the IT equipment state.
7. The root cause analysis method based on IT operation and maintenance system according to claim 6, characterized in that, In response to the detection of an IT equipment status anomaly by the dynamic prediction curve of the IT equipment status, starting from the IT equipment with the anomaly, an IT equipment business importance factor is introduced into the spatiotemporal correlation weight matrix of the IT equipment and edge weights are normalized. A fault propagation path containing cumulative risk value and business impact is generated through a reinforcement learning path search algorithm, including: Define the IT equipment business importance factor as (business priority * traffic share) / (equipment redundancy + 1) and calculate the IT equipment business importance factor. The edge weights in the spatiotemporal correlation matrix of the IT equipment are normalized. Based on the spatiotemporal correlation weight matrix of the IT equipment, the path distance of the IT equipment is determined after edge weight normalization. Build or update the dynamic propagation path map of IT devices based on the path distance of the IT devices; The cumulative risk value is obtained by calculating the cumulative risk value as (cumulative weight * IT equipment business importance factor). Select an IT device with an abnormal status as the source node, and perform risk path search and sorting on the dynamic propagation path map of the IT device based on the cumulative risk value to obtain the path node sequence, cumulative weight value and path risk level. A fault propagation path is generated based on the path node sequence, cumulative weight value, and path risk level.
8. The root cause analysis method based on IT operation and maintenance system according to claim 1, characterized in that, The construction of a multivariate fault knowledge graph integrating historical fault cases, equipment dependencies, and expert experience utilizes a causal graph model to model the knowledge graph. Combining this with the topological structure of fault propagation paths, counterfactual reasoning is used to calculate the causal effect value of each potential root cause. Based on these causal effect values, the most probable IT equipment fault root cause is determined, including: A multi-dimensional fault knowledge graph is constructed based on node type and edge type. The node type includes device entity, fault mode, alarm event and business indicator. The edge type includes causal relationship, dependency relationship and similarity relationship. The multivariate fault knowledge graph is embedded and learned using a causal graph model. The causal effects of each potential root cause are calculated through intervention analysis to obtain fault similarity data. The fault similarity data includes historical rule confidence and real-time state similarity. Based on the propagation topology represented by the fault propagation path and the fault similarity data, the root cause hypothesis is verified by counterfactual reasoning, and the node with the highest causal effect value and that satisfies the topology propagation logic is selected as the root cause of the IT equipment failure.
9. An electronic device comprising a memory, a processor, and an electronic device program on the memory, characterized in that, The processor executes the electronic device program to implement the steps of the root cause analysis method based on the IT operation and maintenance system as described in any one of claims 1-8.
10. An electronic device readable storage medium having an electronic device program / instructions stored thereon, characterized in that, When the electronic device program / instruction is executed by the processor, it implements the steps of the root cause analysis method based on the IT operation and maintenance system as described in any one of claims 1-8.
Citation Information
Cited By
Extensible smart community knowledge graph construction method and system
CN121351973A
Fault prediction method and device for subway power supply system
CN121684616A
Equipment fault prediction method and system and electronic equipment
CN121786757A