Abnormity analysis method and device for multi-source operation and maintenance data, equipment, medium and product

By combining joint anomaly modeling of multi-source operation and maintenance data with operation and maintenance knowledge graphs, the problem of low accuracy of anomaly analysis in traditional operation and maintenance monitoring in multi-cloud environments is solved, and efficient fault monitoring and root cause localization are achieved.

CN121765577APending Publication Date: 2026-03-31SHANGHAI SIGE DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In multi-cloud, hybrid cloud, and microservice environments, traditional operation and maintenance monitoring methods are unable to cope with massive amounts of multi-source data and dynamic changes, resulting in low accuracy of anomaly analysis and a lack of intelligent diagnosis and autonomous decision-making capabilities.

Method used

By acquiring multi-source operation and maintenance data, including time-series indicator data, application logs, call chain tracing data, configuration change records, and alarm events, joint anomaly modeling is performed using a multi-model fusion architecture. Furthermore, the operation and maintenance knowledge graph and structured causal model are combined to trace the path of failure impact and locate the root cause.

Benefits of technology

It improves the accuracy of operation and maintenance anomaly analysis, reduces the complexity of manual intervention, improves the efficiency of fault handling, solves the problem of cross-source data correlation analysis, and realizes intelligent and efficient fault monitoring and management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765577A_ABST
    Figure CN121765577A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data analysis, and provides a multi-source operation and maintenance data anomaly analysis method and device, equipment, a medium and a product, the method comprises the steps that multi-source operation and maintenance data is acquired, and the multi-source operation and maintenance data comprises at least two of index time sequence data, application logs, call link tracking data, configuration change records, alarm events and work orders; carrying out joint anomaly modeling on the preprocessed multi-source operation and maintenance data based on a multi-model fusion architecture to identify an abnormal event in the multi-source operation and maintenance data; based on the operation and maintenance knowledge graph and the structured causal model, fault influence path tracing and root cause positioning are carried out on the abnormal event, a root cause analysis result is obtained, and the root cause analysis result is used for indicating a fault root cause node and a fault propagation path in the abnormal event. Therefore, the accuracy of anomaly analysis of the multi-source operation and maintenance data is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data analysis technology, and more specifically, relates to a method, apparatus, equipment, medium and product for anomaly analysis of multi-source operation and maintenance data. Background Technology

[0002] As enterprise IT architectures grow rapidly in scale and complexity, operations and maintenance systems are facing challenges brought by new technologies such as multi-cloud, hybrid cloud, microservices, and containerization. Traditional monitoring methods that rely on manual experience and static rules are struggling to cope with massive amounts of multi-source data and dynamic changes.

[0003] To address these challenges, the concept of Intelligent Operations and Maintenance (AIOps) has gradually emerged, emphasizing the use of machine learning and automation technologies to improve operation and maintenance monitoring capabilities.

[0004] However, most current tools only achieve basic anomaly detection and alarm. When dealing with massive amounts of multi-source data and dynamic changes, they still lack data analysis, intelligent diagnosis, and autonomous decision-making capabilities, which leads to low accuracy in operation and maintenance anomaly analysis. Summary of the Invention

[0005] The purpose of this application is to provide a method, apparatus, device, medium and product for anomaly analysis of multi-source operation and maintenance data, aiming to solve the technical problem of poor accuracy of anomaly analysis in existing operation and maintenance monitoring methods.

[0006] To achieve the above objectives, according to the first aspect of this application, a method for anomaly analysis of multi-source operation and maintenance data is provided, the method comprising:

[0007] Acquire multi-source operation and maintenance data, which includes at least two of the following: time-series indicator data, application logs, call chain tracing data, configuration change records, alarm events, and work orders;

[0008] Based on a multi-model fusion architecture, joint anomaly modeling is performed on the preprocessed multi-source operation and maintenance data to identify abnormal events in the multi-source operation and maintenance data.

[0009] Based on the operation and maintenance knowledge graph and the structured causal model, the abnormal event is traced to find the source of the fault impact path and the root cause, and the root cause analysis results are obtained. The root cause analysis results are used to indicate the root cause node of the fault and the fault propagation path in the abnormal event.

[0010] The beneficial effects of the embodiments in this application compared with the prior art are:

[0011] This application's embodiments acquire multi-source operation and maintenance (O&M) data, specifically including at least two of the following: time-series indicator data, application logs, call chain tracing data, configuration change records, alarm events, and work orders. This comprehensively covers core dimensions such as system resource status, business operation trajectory, component interaction logic, and fault correlation records. Based on a multi-model fusion architecture, joint anomaly modeling is performed on the preprocessed multi-source O&M data to identify abnormal events. Based on O&M knowledge graphs and structured causal models, fault impact path tracing and root cause localization are performed on abnormal events to obtain root cause analysis results. This approach not only adapts to the analysis needs of massive multi-source data but also effectively improves the accuracy of O&M anomaly analysis.

[0012] Furthermore, by presenting the relationships between various entities in a structured manner through an operations and maintenance knowledge graph, the path of fault impact can be quickly traced. Then, through the reasoning ability of the structured causal model, the causal impact of each entity on abnormal events can be quantified, the root cause node of the fault and the fault propagation path can be locked, the complexity of manual intervention can be reduced, and the efficiency of operations and maintenance fault handling can be improved. This solves the technical problem of the lack of intelligent diagnosis and autonomous decision-making capabilities in existing technologies, which leads to the inaccuracy of anomaly analysis in operations and maintenance monitoring methods.

[0013] According to a second aspect of this application, a multi-source monitoring anomaly analysis device is provided for performing the anomaly analysis method for multi-source operation and maintenance data as described in any one of the claims, comprising:

[0014] The acquisition unit is used to acquire multi-source operation and maintenance data, which includes at least two of the following: indicator time series data, application logs, call chain tracing data, configuration change records, alarm events, and work orders.

[0015] The processing unit is used to perform joint anomaly modeling on the preprocessed multi-source operation and maintenance data based on a multi-model fusion architecture, so as to identify abnormal events in the multi-source operation and maintenance data.

[0016] The analysis unit is used to trace the fault impact path and locate the root cause of the abnormal event based on the operation and maintenance knowledge graph and the structured causal model, and obtain the root cause analysis results. The root cause analysis results are used to indicate the fault root cause node and fault propagation path in the abnormal event.

[0017] According to a third aspect of this application, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device causes the electronic device to perform the method as described in any one of the claims.

[0018] According to a fourth aspect of this application, a computer-readable storage medium is provided that stores a computer program, which, when executed by a processor, implements the method as described in any one of the claims.

[0019] According to a fifth aspect of this application, a computer program product is provided that, when run on an electronic device, causes the electronic device to perform the method described in any one of the first aspects above.

[0020] It is understandable that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating an anomaly analysis method for multi-source operation and maintenance data provided in an embodiment of this application;

[0023] Figure 2 This is a flowchart illustrating an optional method for anomaly analysis of multi-source operation and maintenance data provided in an embodiment of this application;

[0024] Figure 3 This is a flowchart illustrating an optional method for anomaly analysis of multi-source operation and maintenance data provided in an embodiment of this application;

[0025] Figure 4 This is a flowchart illustrating an optional method for anomaly analysis of multi-source operation and maintenance data provided in an embodiment of this application;

[0026] Figure 5 This is a flowchart illustrating an optional method for anomaly analysis of multi-source operation and maintenance data provided in an embodiment of this application;

[0027] Figure 6 This is a schematic diagram of the structure of an anomaly analysis device for multi-source operation and maintenance data provided in an embodiment of this application;

[0028] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0030] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0031] It should also be understood that, in the description of this application, unless otherwise stated, the " / " used in the specification and appended claims indicates that the related objects are in an "or" relationship. For example, A / B can mean A or B. The "and / or" in this application is merely a description of the relationship between the related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0032] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, but are only used for distinguishing descriptions, and the terms "first" and "second" do not necessarily imply that they are different, nor should they be construed as indicating or implying relative importance.

[0033] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0034] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0035] As enterprise IT architectures grow rapidly in scale and complexity, operations and maintenance systems are facing challenges brought by new technologies such as multi-cloud, hybrid cloud, microservices, and containerization. Traditional monitoring methods that rely on manual experience and static rules are struggling to cope with massive amounts of multi-source data and dynamic changes: various logs, metrics, tracing data sources are isolated from each other, forming data silos, making cross-system and multimodal data correlation analysis extremely difficult.

[0036] On the one hand, multi-cloud and hybrid deployments expand monitoring targets from a single data center to different cloud platforms. System components are widely distributed and have complex dependencies, leading to a proliferation of alarms, frequent false alarms, and time-consuming fault localization. On the other hand, the types of operational data are diverse, including structured time-series metrics data and unstructured log text. Multi-source heterogeneous data is difficult to integrate and analyze, resulting in high complexity in root cause investigation and persistently high mean time to repair (MTTR).

[0037] To address these challenges, the Intelligent Operations and Maintenance (AIOps) concept is gaining popularity, emphasizing the use of machine learning and automation technologies to improve operational monitoring capabilities and reduce human intervention and subjective misjudgments. Industry practice shows that integrating data collection, anomaly detection, alarm notification, and response handling to build an automated monitoring pipeline can make anomaly detection and handling more timely and effective. However, most current tools only achieve basic anomaly detection and alarms, and still lack capabilities in cross-source data correlation analysis, intelligent diagnosis, and autonomous decision-making.

[0038] Therefore, it is necessary to propose a new technical solution to build a complete closed-loop system for multi-source monitoring data, from anomaly detection to root cause localization, to solve the problems of data silos, false alarms and missed alarms and slow response in existing solutions, and to truly achieve intelligent and efficient fault monitoring and management.

[0039] This application provides an example of an anomaly analysis method for multi-source operation and maintenance data. Please refer to... Figure 1 As shown, Figure 1This illustration shows a schematic flowchart of an anomaly analysis method for multi-source operation and maintenance data provided in this application. It is provided as an example and not a limitation; the method can be applied to or run in electronic devices, such as an anomaly analysis device for multi-source operation and maintenance data. The method includes:

[0040] S101, acquire multi-source operation and maintenance data.

[0041] Multi-source operation and maintenance data includes at least two of the following: time-series indicator data, application logs, call chain tracing data, configuration change records, alarm events, and work orders.

[0042] S102 performs joint anomaly modeling on preprocessed multi-source operation and maintenance data based on a multi-model fusion architecture to identify abnormal events in the multi-source operation and maintenance data.

[0043] S103, based on the operation and maintenance knowledge graph and structured causal model, traces the path of failure impact and locates the root cause of abnormal events, and obtains the root cause analysis results.

[0044] The root cause analysis results are used to indicate the root cause nodes and propagation paths of faults in abnormal events.

[0045] The anomaly analysis method for multi-source operation and maintenance data provided in this application focuses on solving the challenges of anomaly detection and analysis of multi-source data in multi-tenant and multi-cloud environments. In multi-tenant scenarios of public or hybrid clouds, multiple tenants share underlying resources, and high concurrency loads may lead to resource contention and performance anomalies. However, the monitoring data of different tenants are isolated from each other, posing a challenge to locating the source of the fault. In addition, different cloud platforms use different monitoring strategies and tools, which greatly increases the difficulty of cross-cloud fault diagnosis. Through the embodiments of this application, it is possible to distinguish between tenant-level and node-level anomalies in a multi-tenant environment and support global fault correlation analysis under multi-cloud distributed deployment, thereby overcoming the difficulty of anomaly location caused by cross-tenant and cross-cloud environments.

[0046] In some embodiments, multi-source operation and maintenance data can be understood as the collection of operation and maintenance data from different dimensions during the operation of distributed systems, operation and maintenance systems, photovoltaic energy storage and charging systems, power systems, etc. Multi-source operation and maintenance data must include at least two of the following: time-series indicator data, application logs, call chain tracing data, configuration change records, alarm events, and work orders. Furthermore, the collection process must ensure the integrity, timeliness, and relevance of the data.

[0047] For example, time-series metrics are collected in real time by monitoring tools deployed on various system components, covering system resource status (such as CPU utilization and memory usage), application performance (such as interface response time and request success rate), and business operation metrics (such as order transaction volume and data processing throughput). After collection, the data is organized into a continuous time-series sequence by timestamp. Another example is application logs, obtained through log collection components, which contain various event records output during application operation, such as error messages, runtime descriptions, and exception stacks. Yet another example is call tracing data, generated based on a distributed tracing mechanism, which fully records the call relationships between services, call duration, return status, and other interaction details. Furthermore, configuration change records are obtained by monitoring the configuration management system or parsing configuration file modification logs, reflecting historical changes in system configuration. Alarm events are collected from the alarm outputs of various monitoring systems, capturing abnormal signals identified by the system. Finally, work order data is obtained by calling the work order management system interface, containing manually recorded fault phenomena, processing procedures, and results.

[0048] After acquiring multi-source operation and maintenance data, all collected multi-source operation and maintenance data can be uniformly aggregated to the data storage center. The appropriate storage medium can be selected from the data storage center according to the data type to ensure the efficiency of multi-source operation and maintenance data storage and the convenience of subsequent queries.

[0049] Because the collected multi-source operation and maintenance data may have issues such as inconsistent formats, noise interference, and missing data, standardized preprocessing can be performed on the multi-source operation and maintenance data to resolve these problems and ensure that the data meets the modeling requirements. For example, the preprocessing process includes standardization transformation of various types of data, such as time alignment, missing value imputation, and noise filtering for time-series data, and structured processing, irrelevant information removal, and key field extraction for non-time-series data. Ultimately, all data is converted into structured data in a unified format, providing standardized input for joint anomaly modeling.

[0050] Subsequently, a multi-model fusion architecture is used to perform joint anomaly modeling on the preprocessed multi-source operation and maintenance data. This architecture integrates the advantages of different models to mine the temporal and correlation features of multi-source data, avoiding the problems of missed and false judgments in complex scenarios caused by a single model. Specifically, the time-series features of the data can be extracted first through a time-series coding model to capture the trend of data changes over time and long-term dependencies; then, the topological correlation features between different monitored objects can be mined through a graph structure correlation model to reflect the dependencies and interaction logic between entities; finally, the fused features are analyzed through an anomaly detection model, and the presence of anomalies is determined by combining preset judgment criteria. When the anomaly detection model detects that the data deviates from the normal pattern and meets the anomaly detection conditions, it identifies the abnormal event in the multi-source operation and maintenance data, completing the accurate identification of anomalies in the multi-source operation and maintenance data.

[0051] In the root cause localization process based on operation and maintenance knowledge graphs and structured causal models, the knowledge graph's ability to store relationships and the causal model's reasoning capabilities are leveraged to accurately pinpoint the root cause and propagation path of abnormal events. First, a pre-built and real-time updated operation and maintenance knowledge graph is acquired. The nodes of this knowledge graph cover key operation and maintenance entities in the system, and the edges between nodes are established based on the actual relationships between entities. This provides a structured representation of the system's operational architecture and relational logic, offering a foundation for tracing the path of fault impact. The structured causal model is constructed based on historical fault data, expert experience, and causal discovery algorithms, containing causal relationship rules between various entities, providing a logical basis for root cause reasoning.

[0052] In some embodiments, during implementation, the corresponding abnormal node is first accurately located in the operation and maintenance knowledge graph based on the node information involved in the abnormal event. Then, starting from the abnormal node, the fault impact diffusion analysis is performed along the adjacency relationship in the knowledge graph to screen the relevant nodes affected by the fault and form a candidate fault subgraph. Subsequently, causal reasoning is performed on the nodes in the candidate fault subgraph based on the structured causal model to evaluate the degree of influence of each node on the abnormal event and screen out the root cause node of the fault. Finally, the fault propagation path from the root cause node to the initial abnormal node is sorted out to form the root cause analysis result. It should be understood that the root cause analysis result clearly indicates the root cause node of the abnormal event and the fault propagation path, providing accurate guidance for operation and maintenance personnel to quickly handle the fault.

[0053] Based on the example in this application, an anomaly analysis method for multi-source operation and maintenance data is provided. By comprehensively acquiring multi-source operation and maintenance data, performing joint anomaly modeling by fusing multiple models, and combining operation and maintenance knowledge graphs with structured causal models for root cause localization, the method can accurately identify and trace the root causes of abnormal events. It is applicable to complex operation and maintenance scenarios such as various distributed systems and cloud-native application clusters, and can effectively improve operation and maintenance efficiency and reduce losses caused by system failures.

[0054] Traditional solutions often focus on a single type of data (such as metrics or logs only), making it difficult to capture the full picture of complex faults. This application's embodiments integrate multiple data sources, including monitoring metrics, log text, topology relationships, and configuration changes. Through intelligent semantic understanding and correlation analysis, it transforms the originally loosely structured heterogeneous data into a unified knowledge representation. For example, natural language processing technology is used to interpret unstructured text such as logs and work orders, providing contextual information support for anomaly detection and root cause analysis. When anomalies are detected, log error information and metric anomaly patterns can be combined to compensate for the shortcomings of single-data-source analysis, significantly improving the accuracy of cross-source anomaly analysis. In summary, this invention addresses the shortcomings of existing technologies in data fusion and complex correlation analysis in today's multi-tenant, multi-cloud, and multi-source data operation and maintenance environments.

[0055] In some embodiments, such as Figure 2 As shown, a multi-model fusion architecture is used to perform joint anomaly modeling on preprocessed multi-source operation and maintenance data to identify anomalous events in the multi-source operation and maintenance data, including:

[0056] S201, Construct the graph structure of the monitored objects.

[0057] In this graph structure, nodes represent monitoring entities at different granularities. These entities include tenant-level entities, application-level entities, host node entities, database entities, middleware entities, and network device entities. Edges between nodes are established based on the dependencies of the monitored objects. These dependencies include service call relationships, deployment relationships, network connection relationships, and resource dependencies.

[0058] S202 uses a time-series coding model to extract features from the multi-dimensional time-series data corresponding to each node of the graph structure, and obtains the node time-series feature vector.

[0059] The multidimensional time series data includes time series data of indicators and frequency series data of log events from preprocessed multi-source operation and maintenance data.

[0060] S203 uses a graph attention network to perform topological association fusion on the temporal feature vectors of nodes, resulting in a node representation that integrates temporal features and topological relationships.

[0061] S204 uses variational autoencoders and isolated forests to determine anomalies in node representations and generate anomalous events.

[0062] Among them, abnormal events include at least one of the following information: abnormal time window, involved tenant, involved node, abnormal indicator and degree of deviation. The types of abnormal events include tenant-level abnormal events and node-level abnormal events.

[0063] When performing joint anomaly modeling on preprocessed multi-source operation and maintenance data based on a multi-model fusion architecture, a graph structure of monitoring objects can be constructed first to achieve a structured representation of monitoring entities and their relationships at different granularities. Nodes in the graph structure correspond to monitoring entities at different granularities, specifically including tenant-level entities (representing the overall resource and application set of each tenant), application-level entities (representing specific service components or host nodes representing physical / virtual machines, etc.), host node entities, database entities, middleware entities, and network device entities. Each node uniquely corresponds to a type or a single monitoring object in the system, ensuring coverage of the core layers and key components of system operation and maintenance. Edges between nodes are established based on the actual dependencies of the monitoring objects, specifically covering service call relationships, deployment relationships, network connection relationships, and resource dependency relationships. Service call relationships reflect the business interaction logic between different entities; deployment relationships reflect the bearing relationship between applications and underlying hardware; network connection relationships represent the communication links between entities; and resource dependency relationships reflect the entity's usage requirements for various basic resources. Through the construction of these edges, a graph structure foundation reflecting the system's operational relationship logic is formed.

[0064] After constructing the graph structure, a time-series coding model is used to extract features from the multi-dimensional time-series data corresponding to each node in the graph structure to uncover the temporal correlation features of the data. The multi-dimensional time-series data originates from preprocessed multi-source operational data, specifically including indicator time-series data and log event frequency sequence data. The indicator time-series data reflects the quantitative changes in the operational status of the monitored objects corresponding to the nodes, while the log event frequency sequence data is formed by statistically analyzing the frequency of various log events within a preset time window, reflecting the event occurrence patterns of the nodes. The time-series coding model captures the long-term dependencies, trend changes, and abrupt changes in these multi-dimensional time-series data, performing deep feature extraction and ultimately outputting a high-dimensional node time-series feature vector for each node, providing temporal feature support for subsequent topology association fusion.

[0065] Subsequently, a graph attention network is used to fuse the temporal feature vectors of nodes through topological association, fully exploring the spatial association information between different nodes. The graph attention network utilizes an attention mechanism, assigning different attention weights to each node based on the strength of the dependencies between nodes, and weighted aggregation of the temporal feature vectors of its neighboring nodes. During the fusion process, each node not only retains its own temporal features but also integrates key features from neighboring nodes. Simultaneously, the multi-layer network structure expands the feature perception range, achieving effective fusion of multi-hop neighbor node representations. Ultimately, a node representation integrating temporal features and topological relationships is obtained, which comprehensively reflects the node's own operational state and its associated operational characteristics with other nodes.

[0066] Finally, a combination of variational autoencoders and isolated forests is used to collaboratively determine anomalies in node representations, improving the accuracy and reliability of anomaly identification. The variational autoencoder reconstructs node representations by mapping them to a low-dimensional latent space, and the reconstruction error is used as a crucial basis for anomaly determination; a larger reconstruction error indicates a more significant deviation of the node state from the normal pattern. The isolated forest, by randomly partitioning the feature space, quickly isolates anomalous node representations and outputs an anomaly score. Combining the outputs of both models, a pre-defined judgment rule is used to determine the presence of anomalies. When the comprehensive score exceeds a set threshold, an anomaly event is generated. An anomaly event includes at least one key piece of information: an anomaly time window, involved tenants, involved nodes, anomaly indicators, and the degree of deviation, clearly describing the core characteristics of the anomaly. Furthermore, anomaly events are categorized into tenant-level and node-level anomaly events, corresponding to anomalies in monitored entities at different granularities, enabling accurate classification and characterization of anomaly events and facilitating the identification of anomalies in multi-source operational data.

[0067] In this embodiment, a graph neural network is combined with a time-series deep coding model to perform joint anomaly modeling on multi-source data. Specifically, a graph attention network (GAT) is first introduced to model the relationships between monitored objects (e.g., service call topology, resource dependencies, etc.), using the graph attention mechanism to capture the dependency weights between nodes, thereby incorporating contextual information into anomaly detection. Research shows that introducing spatiotemporal graph attention can effectively identify abnormal patterns in multivariate data, significantly improving the accuracy and efficiency of anomaly detection. Simultaneously, a long short-term memory network (LSTM) is fused as a multidimensional time-series encoder to extract features and model the temporal correlation of historical sequences of various monitoring indicators, embedding its output as a graph network node representation. Based on this, a variational autoencoder (VAE) is used to model the high-dimensional features that integrate topological relationships and temporal patterns, learning a low-dimensional probability distribution representation of the system's normal behavior. The reconstruction error of the VAE is used to measure the degree of anomaly, detecting anomalous points deviating from the normal pattern from a probabilistic perspective.

[0068] In addition to deep learning models, unsupervised anomaly detection methods such as Isolation Forest can also be used. Isolation Forest has efficient anomaly detection capabilities for high-dimensional data, discovering anomalies without the need for supervision. By combining deep models with methods like Isolation Forest, anomalies can be identified from different perspectives, improving robustness. It can achieve refined anomaly identification for complex multi-source data, supporting both tenant-level global behavior modeling (e.g., detecting anomalies in a tenant's overall resource usage patterns) and node-level local anomaly detection (e.g., performance anomalies of specific hosts or containers). Graph attention enhances the ability to capture cross-node anomaly associations, LSTM-VAE improves the characterization of multi-dimensional temporal patterns, and Isolation Forest reduces the false negative rate, significantly improving the accuracy and recall of anomaly detection in complex scenarios.

[0069] As a supplement, embodiments of this application allow defining certain business rules or static thresholds to capture specific simple anomalies, such as a single metric exceeding N times the historical standard deviation. Verification can be performed before and after the deep model's determination to prevent obvious anomalies undetected by the model or to filter out misjudgments by the model.

[0070] In some embodiments, such as Figure 3 As shown, a time-series coding model is used to extract features from the multi-dimensional time-series data corresponding to each node of the graph structure, resulting in a node time-series feature vector, including:

[0071] S301 deploys a Long Short-Term Memory (LSTM) network with shared parameters for each node and uses the LSM network as the temporal coding model.

[0072] S302, input the multidimensional time series data of each node within the past preset time period into the corresponding long short-term memory network.

[0073] S303 models the temporal correlation of each node's multidimensional time series data within a preset time period through the hidden layers of the Long Short-Term Memory network, and outputs the node's temporal feature vector.

[0074] When using a time-series coding model to extract features from multidimensional time-series data corresponding to each node in a graph structure, a Long Short-Term Memory (LSTM) network with shared parameters is first deployed for each node in the graph structure, and this LSTM network is used as the core time-series coding model. The shared parameter settings ensure that the model can reuse its learned general time-dimensional modeling capabilities when learning time-series data features from different nodes, avoiding overfitting due to insufficient data for a single node. This also improves model training efficiency and generalization ability, enabling the model to quickly adapt to multidimensional time-series data features from different types of nodes (such as tenant-level entities, host node entities, database entities, etc.).

[0075] Subsequently, for each node, its multidimensional time-series data within a preset time period is selected as model input. The preset time period can be flexibly configured according to the monitoring accuracy requirements of the system operation and maintenance scenario, for example, set to the past 5 minutes, 10 minutes, or 30 minutes, ensuring that the input data can cover the key change cycles of the node's operating status. The multidimensional time-series data for each node is preprocessed data, specifically including the time-series data of the corresponding indicators and the frequency sequence data of log events. This data has undergone standardization processing such as time alignment, missing value imputation, and noise filtering, accurately reflecting the changing patterns of each node's operating status. This data is then organized into an ordered data sequence according to chronological order and input into the shared parameter LSTM network corresponding to the node, providing standardized input for time correlation modeling.

[0076] Finally, the hidden layers of the LSTM network are used to model the temporal correlations of the input multidimensional time series data. Leveraging its unique gating mechanism (input gate, forget gate, output gate), the LSTM network can effectively capture long-term dependencies, trend changes, and abrupt changes in time series data, such as the periodic fluctuations of node indicator time series data and sudden increases in log event frequency. During the modeling process, the hidden neurons of the LSTM network continuously update their states, recording the correlation information of data at different time steps, and performing deep feature mining and integration of the multidimensional time series data. After forward propagation computation, the output layer outputs a high-dimensional node temporal feature vector for each node. This high-dimensional node temporal feature vector condenses the core temporal features of the node within a preset time period of the multidimensional time series data, providing solid temporal dimension feature support for subsequent topological correlation fusion using a graph attention network.

[0077] In some embodiments, such as Figure 4 As shown, a graph attention network is used to fuse the temporal feature vectors of nodes through topological association, resulting in a node representation that fuses temporal features and topological relationships, including:

[0078] S401 inputs the node temporal feature vector of each node into the graph attention network.

[0079] S402 uses a graph attention network to calculate the attention coefficient between each node and all its neighboring nodes based on the feature similarity and dependency strength between nodes. The attention coefficient is used to quantify the influence weight of neighboring nodes on the current node.

[0080] S403, based on the attention coefficient, performs a weighted summation of the temporal feature vectors of all neighboring nodes of each node, so as to initially fuse the representations of neighboring nodes with the representation of the current node.

[0081] S404 employs a multi-layer stacked graph attention network structure. Through iterative propagation of the multi-layer network, each node aggregates the node representations of its multi-hop neighbor nodes layer by layer.

[0082] Each layer of the network captures the feature information of the current node's one-hop neighbor nodes. The node representation includes each node's own temporal features, neighbor node association features, and multi-hop topology information.

[0083] When performing topological association fusion on the temporal feature vectors of nodes, the temporal feature vector of each node is first used as the input data of the graph attention network to ensure that the temporal features of all nodes can be fully captured by the network and participate in the association calculation. The input layer of the graph attention network is adapted to the dimension of the node temporal feature vectors, and can directly receive high-dimensional temporal feature vectors, laying the foundation for subsequent attention coefficient calculation and feature fusion.

[0084] Subsequently, the graph attention network calculates the attention coefficient between each node and all its neighboring nodes based on the feature similarity and dependency strength between nodes. Feature similarity is determined by calculating the cosine similarity or Euclidean distance between the temporal feature vectors of the current node and its neighboring nodes, reflecting the degree of fit between them in the representation of the running node. Dependency strength is obtained based on the weight attribute of the edges between nodes in the graph structure. This weight attribute is set during the graph structure construction phase according to the importance of the dependency (such as the coreness of service calls and the stability of network connections). By weightedly fusing feature similarity and dependency strength, the comprehensive influence weight of each neighboring node on the current node is obtained, i.e., the attention coefficient. This coefficient quantifies the contribution of the neighboring node's representation to the current node's representation, giving higher weights to closely related nodes with similar features.

[0085] Based on the calculated attention coefficients, the temporal feature vectors of all neighboring nodes of each node are weighted and summed. Specifically, the temporal feature vector of each neighboring node is multiplied by its corresponding attention coefficient, and then the summation of all weighted neighboring node representation vectors is performed to obtain the neighboring node representation fusion result. Subsequently, this result is concatenated or weighted and fused with the temporal feature vector of the current node itself to achieve a preliminary fusion of the neighboring node representation and the current node representation, so that the node representation retains its own temporal dimension characteristics while incorporating the association features of directly adjacent nodes.

[0086] To fully explore the correlation information in multi-hop topology, a multi-layered stacked graph attention network structure is used for iterative propagation and fusion. Each layer of the graph attention network performs the aforementioned attention coefficient calculation and neighborhood feature weighted summation operations, and each layer only captures the feature information of the current node's one-hop neighbor nodes. In the first layer, nodes aggregate the representations of one-hop neighbor nodes; upon entering the second layer, each node's input features already include its own temporal features and one-hop neighbor features. At this point, the "one-hop neighbor" captured by the network is actually the original node's two-hop neighbor, and the two-hop neighbor features are aggregated to the current node through the second layer fusion; and so on. Through iterative propagation of multiple layers, each node can aggregate the node representations of multi-hop neighbor nodes layer by layer, and the number of hops is consistent with the number of network layers.

[0087] After topological correlation fusion through a multi-layer graph attention network, the final output is a node representation for each node. This node representation is a deep fusion of temporal features and topological relationships. It not only includes the node's own temporal features and the correlation features of its direct neighboring nodes, but also integrates the indirect correlation features in multi-hop topological relationships. It can comprehensively and three-dimensionally reflect the node's own operating state and its correlation operating features with other nodes in the entire system, providing high-quality feature input for subsequent anomaly detection based on variational autoencoders and isolated forests.

[0088] In some embodiments, such as Figure 5 As shown, variational autoencoders and isolated forests are used to determine anomalies in node representations and generate anomalous events, including:

[0089] S501 inputs the node representation into the variational autoencoder, maps the node representation to a low-dimensional latent variable distribution through the encoding unit of the variational autoencoder, and then outputs the reconstructed node representation corresponding to the node representation through the decoding unit of the variational autoencoder.

[0090] S502, based on the reconstructed node representation and node representation, calculate the reconstruction error or reconstruction probability as the first anomaly score.

[0091] S503: Input the node representation into the isolated forest, and randomly divide the node representation into isolated outlier data points through the isolated forest to obtain the second outlier score.

[0092] S504 adopts a dynamic threshold determination mechanism, combining the first anomaly score and the second anomaly score to determine the anomaly determination result.

[0093] S505, when the anomaly determination result is an anomaly, generates an anomaly event that includes the time window in which the anomaly occurred, the tenants involved, the nodes involved, the anomaly indicators and the degree of deviation of the values, the anomaly score, and a list of associated nodes.

[0094] When performing anomaly detection on node representations that integrate temporal features and topological relationships, the node representation of each node is first input into the variational autoencoder (VAE) and the isolated forest model respectively. Multi-dimensional anomaly scores are obtained through collaborative computation of the two models, thereby improving the accuracy of anomaly identification.

[0095] For the processing flow of a variational autoencoder (VAE), node representations, as high-dimensional feature vectors, are first input into the encoding unit of the VAE. The encoding unit maps the high-dimensional node representations to a low-dimensional latent variable distribution through neural network layers. This latent variable distribution follows a predefined probability distribution (such as a Gaussian distribution), which can compress and abstract the core features of the node representations. Subsequently, the decoding unit reconstructs the node representations based on the low-dimensional latent variable distribution, restoring the high-dimensional feature vectors through reverse mapping. The reconstruction process aims to maximize the restoration of key information from the original node representations. After reconstruction, the difference index between the reconstructed node representations and the original node representations is calculated. Reconstruction error (such as mean squared error or cosine distance) or reconstruction probability can be chosen as the first anomaly score: if reconstruction error is used, a larger error value indicates a more significant deviation of the original node representation from the normal pattern; if reconstruction probability is used, a smaller probability value indicates a lower probability that the node representation belongs to a normal distribution, corresponding to a more obvious anomaly tendency.

[0096] Simultaneously, the same node representation is input into the Isolation Forest model for parallel anomaly detection. The Isolation Forest model constructs multiple decision trees, randomly partitioning the feature space containing the node representation. Each partition is based on randomly selected feature dimensions and a partition threshold, progressively isolating data points deviating from the normal cluster. Since anomalous data points are typically isolated in the feature space, they can be quickly isolated with fewer partitioning steps. Therefore, a second anomaly score is calculated based on the isolation path length: the shorter the isolation path, the higher the second anomaly score, indicating a greater probability that the node representation is an anomalous data point; conversely, the longer the isolation path, the higher the second anomaly score, indicating a lower probability that the node representation is an anomalous data point and a node representation closer to the normal pattern.

[0097] To adapt to the dynamic changes in system operation status, a dynamic threshold determination mechanism is adopted to integrate the first and second anomaly scores to determine the final anomaly determination result. The dynamic threshold is set based on node characterization data from historical normal operation periods. Statistical analysis is used to obtain the distribution characteristics (such as mean, variance, and quantiles) of the first and second anomaly scores, and an initial threshold is determined by combining this with a preset confidence level (such as a 95% confidence interval). During actual operation, the system automatically updates historical data samples and recalculates threshold parameters every preset period (such as 1 hour) or after accumulating a certain amount of new data, ensuring that the threshold can dynamically adapt to changes in data distribution. During anomaly determination, the first and second anomaly scores are first standardized (e.g., mapped to the [0,1] interval), and then different weights are assigned according to model importance (the weights can be determined through verification and optimization using historical anomaly data). The weighted sum of the two is calculated as the comprehensive anomaly score. The comprehensive anomaly score is compared with the dynamic threshold. If the comprehensive anomaly score exceeds the dynamic threshold, the anomaly determination result is anomaly; if it does not exceed the threshold, it is determined as normal.

[0098] When an anomaly is determined, the system automatically generates structured anomaly events to ensure the integrity and traceability of anomaly information. Anomaly events contain multi-dimensional key information: the time window of the anomaly is accurate to the second, with corresponding nodes representing the time period to clearly define the specific time period in which the anomaly occurred; the tenant-related nodes represent the corresponding tenant-level entities, accurately locating the business affected by the anomaly; the involved nodes clearly identify the monitoring entities corresponding to the anomaly (such as host node entities, database entities, etc.), indicating the specific object where the anomaly occurred; the anomaly indicators and their degree of deviation, with anomaly indicator items in the multi-dimensional time series data corresponding to the related nodes, calculating the deviation ratio or absolute difference between the indicator value and the normal threshold range to quantify the severity of the anomaly; the anomaly score includes a first anomaly score, a second anomaly score, and a comprehensive anomaly score, providing a basis for anomaly level determination; the list of related nodes is based on the topological relationships in the operation and maintenance knowledge graph, listing nodes with direct or indirect dependencies on the anomaly node, providing clues for subsequent fault tracing. By integrating the above information, the generated anomaly events can comprehensively and clearly characterize the anomaly features, providing accurate data support for operation and maintenance personnel to quickly locate the root cause of the anomaly and assess the scope of impact.

[0099] In some embodiments, based on operational knowledge graphs and structured causal models, fault impact paths and root causes are traced and located for abnormal events to obtain root cause analysis results, including:

[0100] Obtain a pre-built and real-time updated operations and maintenance knowledge graph.

[0101] The nodes of the operation and maintenance knowledge graph include at least one of business applications, microservices, host nodes, databases, middleware, network devices, configuration items, and change events. The edges of the operation and maintenance knowledge graph include at least one of call relationships, deployment relationships, dependency relationships, network connection relationships, and tenant affiliation relationships. The operation and maintenance knowledge graph is built based on configuration item data, service governance framework data, and historical fault analysis data.

[0102] Based on the node information involved in the abnormal event, the corresponding abnormal node is located in the operation and maintenance knowledge graph;

[0103] Fault impact diffusion analysis is performed along the adjacency relationship of abnormal nodes, and relevant neighbor nodes are screened to form a candidate fault subgraph;

[0104] Root cause reasoning is performed on candidate fault subgraphs based on a structured causal model, and the root cause analysis results are output to indicate the root cause nodes and fault propagation paths.

[0105] When tracing the impact path and root cause of abnormal events, a pre-built and real-time updated operations and maintenance (O&M) knowledge graph is first obtained. This knowledge graph provides structured support for relational queries and path tracing. The nodes of the O&M knowledge graph cover at least one of the following: business applications, microservices, host nodes, databases, middleware, network devices, configuration items (CMDB), and change events. Each node contains a unique identifier, type, and attribute information (such as the hardware configuration of the host node, the deployment address of the microservice, and the parameter values ​​of the configuration item). It comprehensively covers the key entities of system operations and maintenance (applications, servers, databases, etc.) and their dependencies. That is, the edges between the nodes corresponding to entities correspond to various types of relationships, including at least one of the following: call relationships, deployment relationships, dependency relationships, network connection relationships, and tenant affiliation relationships. For example, the call relationship between microservices and databases, the deployment relationship between business applications and host nodes, and the dependency relationship between middleware and configuration items. Each edge records key attributes such as relationship type and relationship strength.

[0106] It should be understood that the construction of this operations and maintenance knowledge graph is based on configuration item data, service governance framework data (such as the service list and interface dependencies of the service registry), and historical fault analysis data. Entity association information is extracted using automated parsing tools to initialize nodes and edges. Simultaneously, a real-time update mechanism is set up to connect with real-time data such as configuration change notifications, service online / offline events, and fault handling records, dynamically adjusting node attributes and edge relationships to ensure the knowledge graph accurately reflects the system's current operational architecture and relational logic. Edge relationships are derived from summaries of the CMDB, service governance framework, and historical fault analysis. For example, information such as service A deployed on host M, upstream call to service B, and dependency on database D are all recorded in the graph as triples.

[0107] In addition, the operations and maintenance knowledge graph integrates monitoring data relationships, such as the relationship between nodes and their generated logs / metrics, and the relationship between alarm events and related components. The operations and maintenance knowledge graph is updated in real time when new monitoring data or configuration changes occur. For example, if a new service is deployed, the corresponding nodes and relationships are automatically added; if a change causes a topology adjustment, the edges are also modified accordingly. As a globally unified semantic model, the operations and maintenance knowledge graph provides background knowledge and relationship network support for root cause analysis.

[0108] Next, based on the node information contained in the abnormal event, the corresponding abnormal node is accurately located in the operation and maintenance knowledge graph. The specific node involved in the abnormal event (such as a host node or a microservice instance) is clearly recorded in the abnormal event. By using the node's unique identifier or attribute matching (such as node name, IP address, service ID, etc.), the node is quickly queried and located in the operation and maintenance knowledge graph. This node is marked as the starting node for fault tracing, thus defining the starting point for subsequent impact path analysis.

[0109] Subsequently, a fault impact diffusion analysis is performed along the adjacency relationships of the abnormal nodes, and relevant neighboring nodes are selected to form a candidate fault subgraph. Centered on the marked abnormal node, all its directly adjacent nodes in the knowledge graph (i.e., nodes directly related by an edge) are traversed, including the deployed applications, dependent network devices, and associated configuration items of the abnormal host node. For each directly adjacent node, the correlation of the node's impact from the fault is assessed by combining abnormal event characteristics (such as abnormal time windows and abnormal indicators) and preprocessed multi-source operation and maintenance data: if a neighboring node exhibits operational status changes related to the abnormal node within the abnormal time window (such as prolonged response time, sudden increase in resource utilization, or increased frequency of log errors), it is identified as a relevant neighboring node. Based on this, the neighboring nodes of the relevant neighboring nodes are further traversed, and the correlation assessment process is repeated to gradually expand the scope of impact screening.

[0110] Through multiple rounds of adjacency traversal and correlation filtering, all nodes and corresponding edges that are directly or indirectly related to the abnormal event are extracted to form a candidate fault subgraph. This candidate fault subgraph can focus on the core entities and relationships that the fault may involve, eliminate interference from irrelevant nodes, and narrow the analysis scope for root cause reasoning.

[0111] Finally, root cause reasoning is performed on the candidate fault subgraph based on the structured causal model, and the root cause analysis results are output. The structured causal model is built based on historical fault handling data, expert experience, and causal discovery algorithms. It includes causal association rules between various entities (such as configuration parameter modification → microservice response delay, database connection limit exceeded → call failure). Each rule clearly defines the cause node, result node, causal association strength, and triggering conditions.

[0112] In the root cause reasoning process, the nodes in the candidate fault subgraph are first matched with the causal rule base of the structured causal model to identify potential causal relationships between nodes. Then, based on the strength of causal associations and evidence from multi-source operational data (such as configuration change records, alarm event sequences, and work order descriptions), the probability score of each node as a root cause of the fault is calculated. The node with the highest probability score is selected as the root cause node by sorting, and the propagation path from the root cause node to the initial abnormal node is traced along the association edges in the candidate fault subgraph to clarify the transmission logic and order of impact of the fault among different nodes. The final root cause analysis results clearly indicate the specific information of the root cause node (such as node type, attributes, and abnormal state) and the complete fault propagation path, providing precise guidance for operations and maintenance personnel to quickly locate the root cause of the fault and formulate a handling plan.

[0113] By explicitly modeling entity relationships (e.g., "Service A depends on component B"), knowledge graphs clearly depict possible failure propagation paths, helping to reveal implicit relationships in traditional monitoring. On one hand, leveraging the causal relationship knowledge provided by knowledge graphs, graph inference algorithms (such as graphical neural networks (GNNs) or probabilistic graphical models / Bayesian networks) are used to calculate the probability of failure propagation on the graph, gradually tracing back along the causal chain to locate possible root cause nodes. On the other hand, by integrating a large amount of historical monitoring data and failure cases from the AIOps field, causal discovery algorithms (such as the IGSP algorithm) are used to learn the strength of causal influence between entities from the data, summarizing implicit failure causal relationships (e.g., discovering that high disk I / O leads to increased database latency). Open-source causal inference frameworks can be used to verify and estimate hypothetical causal relationships, making the root cause analysis process and results more interpretable.

[0114] By combining knowledge graphs with causal reasoning, the impact path of a fault can be automatically reconstructed when an anomaly occurs, tracing it from the apparent anomaly all the way to the initial root cause. For example, when an application service experiences an anomaly in response time, the anomalies of downstream components can be checked by following the dependencies in the knowledge graph. The probability of each candidate cause leading to the anomaly can be calculated using a causal model, ultimately pinpointing the root cause component and the anomaly event that triggered the fault. Compared to traditional rule-based or correlation-based analysis methods, this approach demonstrates stronger adaptability in multi-cloud distributed deployment environments: regardless of how many service nodes the fault chain spans or which cloud platforms it is deployed on, the unified knowledge graph makes cross-regional and cross-platform dependencies readily apparent. Causal reasoning can locate faults across cloud environment boundaries, thus supporting end-to-end root cause analysis in complex multi-cloud architectures. Explicit knowledge representation and causal link reasoning also improve the reliability and interpretability of the location results.

[0115] In some embodiments, fault impact diffusion analysis is performed along the adjacency relationships of anomalous nodes to screen relevant neighbor nodes and form a candidate fault subgraph, including:

[0116] A breadth-first search or depth-first search algorithm is used to traverse the adjacency relationship of the operation and maintenance knowledge graph starting from the abnormal node, so as to perform fault impact diffusion analysis on the adjacency relationship of the abnormal node.

[0117] During the process of traversing the adjacency relationships of the operations and maintenance knowledge graph, the probability of fault propagation corresponding to each adjacency relationship is evaluated by combining the detection results of abnormal events.

[0118] Based on the fault propagation probability corresponding to each adjacency, neighbor nodes with a fault propagation probability greater than a preset threshold are selected and together with abnormal nodes, they form a candidate fault subgraph.

[0119] In some embodiments, when performing fault impact diffusion analysis to form a candidate fault subgraph, a breadth-first search (BFS) or depth-first search (DFS) algorithm is first used. Starting from the located abnormal node, the algorithm traverses along the adjacency relationships in the operation and maintenance knowledge graph to achieve an ordered diffusion analysis of the fault impact range. If the breadth-first search algorithm is selected, it starts from the abnormal node and first traverses all its direct adjacent nodes (1-hop neighbor nodes). After evaluating the direct adjacent nodes, it then traverses the adjacent nodes (2-hop neighbor nodes) of each direct adjacent node in turn, and so on, to ensure that the fault spreads layer by layer according to the distance of the fault propagation, fully covering short-distance associated nodes. If the depth-first search algorithm is selected, it starts from the abnormal node and continues to traverse along a certain adjacency relationship chain to the end of the path, and then backtracks to the untraversed adjacency relationships to continue to delve deeper. This is suitable for scenarios that need to prioritize the discovery of deeply affected nodes under specific associated paths.

[0120] It should be noted that both of the above search algorithms can be flexibly adapted to the topological characteristics of different systems, ensuring a comprehensive and orderly analysis of the fault impact diffusion of the adjacency relationships of abnormal nodes, and avoiding the omission of key related nodes.

[0121] During the traversal of the adjacency relationships in the operations and maintenance knowledge graph, the probability of fault propagation for each adjacency relationship is quantitatively assessed by combining the detection results of abnormal events. The detection results of abnormal events include key information such as abnormal time windows, abnormal indicators, abnormal scores, and association characteristics. This information is combined with the attributes of adjacency relationships (such as relationship type and association strength) to construct a probability assessment model.

[0122] For example, for service call relationship-type adjacencies, combined with the detection results of "abnormal interface response time" in abnormal events, if the calling node and the called node have synchronized performance fluctuations within the abnormal time window, and the adjacency relationship has a large number of historical fault propagation records, then the corresponding fault propagation probability is high. For deployment relationship-type adjacencies, combined with the detection results of "host resource utilization exceeding limits" in abnormal events, if application nodes deployed on the same host all experience resource contention-type anomalies, then the fault propagation probability of the adjacency relationship is significantly increased. By integrating real-time detection data of abnormal events with historical correlation data of adjacencies, the fault propagation probability of each adjacency relationship can be accurately quantified, providing an objective basis for node selection.

[0123] Subsequently, based on the fault propagation probability corresponding to each adjacency, a preset threshold is set (this threshold can be configured based on historical fault data statistics or expert experience, such as 0.6), and neighboring nodes with a fault propagation probability greater than the preset threshold are selected. These nodes are determined to be strongly correlated with the abnormal event and have a high probability of being affected by the fault. They are extracted together with the initial abnormal nodes, while retaining the adjacency relationships between these nodes, forming a candidate fault subgraph. This candidate fault subgraph focuses on core associated nodes with a fault propagation probability that meets the threshold. It eliminates irrelevant nodes with no fault propagation risk or extremely low propagation probability, while fully retaining the key nodes and associated paths that the fault may affect, providing a precise and efficient analytical scope for subsequent root cause inference based on a structured causal model.

[0124] In some embodiments, root cause reasoning is performed on candidate fault subgraphs based on a structured causal model, outputting root cause analysis results that indicate the root cause nodes and fault propagation paths, including:

[0125] Based on historical operation and maintenance data, a causal discovery algorithm is used to learn the causal relationships between entities and form a causal rule base.

[0126] For each node in the candidate fault subgraph, an initial causal correlation degree is assigned to each node based on the causal rule base;

[0127] Using counterfactual reasoning or Bayesian inference, calculate the causal impact score of each node on the anomalous event;

[0128] Each node is sorted from high to low based on the causal impact score, and the node with the highest score is identified as the root cause node of the failure.

[0129] Based on the relationship chains in the operation and maintenance knowledge graph, the failure propagation path from the root cause node to the abnormal node is sorted out to obtain the root cause analysis results.

[0130] In this application embodiment, the focus is on the logic of constructing a causal rule base, assigning causal correlation degree, calculating causal impact score, and determining root cause nodes and propagation paths, ensuring that the description and corresponding technical solution are accurately matched, and without introducing irrelevant technical details.

[0131] When performing root cause reasoning on candidate fault subgraphs, a causal discovery algorithm is first used to learn the causal relationships between various entities in the system based on historical operation and maintenance data, thus constructing a structured causal rule base. Historical operation and maintenance data includes past fault handling records, abnormal correlation data from multi-source operation and maintenance data, and expert experience summaries, providing rich sample support for causal relationship mining. The causal discovery algorithm can be a constraint-based algorithm (such as the PC algorithm or the IGSP algorithm) or a score-based algorithm (such as the Bayesian information criterion scoring algorithm), which analyzes the conditional independence, correlation strength, and temporal order of variables between entities to uncover potential causal relationships.

[0132] For example, by analyzing the temporal and statistical correlations between "configuration parameter modification" and "microservice response time extension" in historical data, a causal relationship between the two can be discovered; by mining the conditional dependency between "database connection pool full" and "interface call failure," corresponding causal rules can be established. Each causal rule clearly includes key information such as the cause node type, result node type, causal correlation strength, triggering conditions, and time delay. All rules are integrated to form a causal rule library, providing a foundation for subsequent causal reasoning.

[0133] Subsequently, for each node in the candidate fault subgraph, an initial causal correlation degree is assigned to each node based on the constructed causal rule base. All nodes in the candidate fault subgraph are traversed, and the causal relationship between each node and other nodes (especially the initial anomalous node) is queried in the causal rule base: if node A has a direct causal rule with the initial anomalous node B, and the causal correlation strength in the rule is 0.85, then the initial causal correlation degree of node A is assigned a value of 0.85; if node C does not have a direct causal relationship with the initial anomalous node B, but there is an indirect causal link (such as C→D→B), then the initial causal correlation degree of node C is calculated based on the product or weighted sum of the causal correlation strengths at each step in the indirect link; if a node has no direct or indirect causal relationship with the initial anomalous node, then the initial causal correlation degree is assigned a value of 0, and it is marked as a low-correlation candidate node. Through the assignment of initial causal correlation degrees, nodes with potential causal relationships with the anomalous event are initially screened.

[0134] Next, counterfactual reasoning or Bayesian inference is used to optimize the initial causal correlation of each node in the candidate fault subgraph, and to calculate the causal impact score of each node on the abnormal event. If counterfactual reasoning is used, a counterfactual scenario is constructed based on a causal inference framework (such as Do-calculus): assuming "removing the abnormal state of a candidate node," the causal impact of the node on the abnormal event is evaluated by simulating the change in the probability of the abnormal event occurring in this scenario. If the probability of the abnormal event occurs significantly after simulation (e.g., a decrease of more than 50%), the corresponding node has a higher causal impact score; otherwise, the score is lower. If Bayesian inference is used, the node state in the candidate fault subgraph is used as the observed variable, and the abnormal event is used as the outcome variable. Based on the prior probability in the causal rule base, combined with real-time evidence from multi-source operation and maintenance data (such as node anomaly indicators, configuration change records, and alarm timing), the posterior probability is updated. The posterior probability value is used as the causal impact score of the node; a higher posterior probability indicates a greater likelihood that the node is the root cause. Both methods can effectively quantify the causal contribution of nodes to abnormal events and improve the accuracy of root cause determination.

[0135] After calculating the causal impact score, all nodes in the candidate fault subgraph are sorted from highest to lowest score. Typically, the top 1-3 nodes are selected as the root cause nodes; the specific number can be adjusted based on system complexity and actual analysis needs. For systems with simple topologies, the top-ranked node can be selected as the sole root cause node; for complex distributed systems, the top 3 nodes can be selected as joint root cause nodes or primary and secondary root cause nodes. During the sorting process, if nodes with the same or similar scores exist, a secondary screening can be performed by combining node type priority (e.g., configuration item nodes and core service nodes have higher priority than ordinary network device nodes) and the frequency of historical root causes to ensure the rationality of the root cause node determination.

[0136] Finally, based on the relationship chains in the operations and maintenance knowledge graph, the complete fault propagation path from the identified root cause node to the initial anomalous node is traced. Starting from the root cause node, the intermediate nodes and related links of the fault propagation are traced along the adjacency relationships preserved in the candidate fault subgraph, clarifying the relationship type (such as service call relationship, dependency relationship), propagation order, and degree of impact of each propagation link. For example, the root cause node is "configuration item X parameter modification", which affects the intermediate node "microservice A" through "configuration dependency relationship", and then affects the intermediate node "database B" through "service call relationship", ultimately causing the initial anomalous node "application C" to malfunction, forming a complete propagation path of "configuration item X → microservice A → database B → application C". The detailed information of the root cause node (such as node name, anomalous status, causal impact score) is integrated with the traced fault propagation path to form the final root cause analysis results, providing clear and practical guidance for operations and maintenance personnel to accurately handle faults.

[0137] In some embodiments, after acquiring multi-source operation and maintenance data, the method further includes:

[0138] Perform missing value imputation and noise filtering on multi-source operation and maintenance data;

[0139] Resample and time-align multi-source operation and maintenance data according to a preset time window;

[0140] The multi-source operation and maintenance data is cleaned separately to remove irrelevant fields and format characters, and the key fields in the multi-source operation and maintenance data are extracted. The key fields include at least one of the following: log level, error code, and exception stack trace.

[0141] Key fields are segmented and vectorized to obtain preprocessed multi-source operation and maintenance data.

[0142] In some embodiments, after acquiring multi-source operation and maintenance data, to eliminate the interference of data quality issues on subsequent anomaly modeling, preprocessing is required according to a standardized process to convert the raw data into standardized, usable data. First, missing value imputation and noise filtering are performed on the multi-source operation and maintenance data. For missing values ​​in the time-series indicator data, an appropriate imputation method is selected based on the data type:

[0143] For continuous metrics (such as CPU utilization and response time), linear interpolation or moving average methods are used to fit and calculate the imputed value based on normal data points before and after the missing value. For discrete data (such as log event frequency and alarm status), the mode imputation method is used to fill in the missing value with the value that appears most frequently within the time interval of the missing value to ensure the continuity of the data.

[0144] Meanwhile, targeted filtering strategies are adopted to address noise interference in the data: for sudden abnormal peaks in the time series data of indicators, noise points that exceed the normal distribution range are identified and removed by the 3σ principle, and then the data curve is smoothed by a low-pass filtering algorithm; for invalid characters and redundant information (such as garbled characters and duplicate records) in application logs and work order data, regular expression matching is used to filter and retain valid data content, thereby improving data purity.

[0145] Subsequently, the multi-source operation and maintenance data is resampled and time-aligned according to a preset time window. The preset time window can be flexibly configured according to the system monitoring accuracy requirements (e.g., 5 seconds, 10 seconds, 30 seconds) to ensure that the time granularity of all data is consistent: for time-series indicator data, downsampling or upsampling is used to adjust the data to the preset time window granularity. During downsampling, the mean, maximum, or median of the data within the time window is taken as the representative value of that time window. During upsampling, interpolation is used to supplement missing time point data. For non-time-series data such as application logs and call chain tracing data, the timestamps are mapped to the corresponding preset time windows, and the data frequency or key features within each time window are statistically analyzed. On this basis, using a unified time axis as a benchmark, all types of multi-source operation and maintenance data are time-aligned to ensure that data within the same time window can be correlated and to avoid subsequent analysis errors caused by time deviations.

[0146] Next, the multi-source operation and maintenance data were cleaned separately to remove irrelevant fields and formatting characters, and extract key fields. Differentiated cleaning strategies were adopted for different types of data: For application logs, HTML tags, special escape characters, redundant spaces, and other formatting characters were filtered out; fields irrelevant to anomaly analysis (such as log thread names and irrelevant debugging information) were removed; and key fields such as log levels (such as ERROR, WARN, INFO), error codes, exception stack traces, timestamps, and core event descriptions were extracted. For alarm events and work order data, a unified field naming convention was adopted (such as unifying alarm level and alarm grade as "alarm grade"), invalid descriptive information was filtered out, and key fields such as fault phenomena, involved modules, alarm types, and processing results were extracted. For configuration change records, redundant descriptions in the change instructions were removed, and core fields such as configuration item names, values ​​before and after the change, change time, and the person making the change were extracted. Through targeted cleaning, the key information in the data was focused on, reducing the complexity of data processing.

[0147] Finally, the extracted key fields are segmented and vectorized to obtain preprocessed multi-source operation and maintenance data. For Chinese key fields (such as log descriptions and fault phenomena), word segmentation algorithms are used to segment them into independent semantic units (e.g., splitting "database connection timeout" into "database", "connection", and "timeout"). For English key fields (such as exception stack traces and interface names), they are segmented into word-level results by spaces or special symbols. After word segmentation, the results are converted into feature vectors using a bag-of-words model or TF-IDF algorithm: for discrete key fields such as log levels and error codes, one-hot encoding is used to convert them into binary feature vectors; for text-based key fields (such as exception descriptions and fault phenomena), the weight of each word is calculated based on the TF-IDF algorithm to construct high-dimensional sparse feature vectors. Through word segmentation and vectorization, unstructured and semi-structured key fields are converted into machine-recognizable structured feature data, providing standardized input for subsequent joint anomaly modeling based on a multi-model fusion architecture.

[0148] In some embodiments, key fields also include the time of anomaly occurrence, duration of anomaly, root cause node information, and fault propagation path information; the method further includes:

[0149] The root cause analysis results and key fields are embedded into a preset prompt template to generate model input data. The prompt template is used to specify the paragraph structure and tone style of the fault diagnosis report.

[0150] The large language model is invoked to process the model input data and generate a structured fault diagnosis report.

[0151] Among them, large language models include cloud-based large language models or pre-trained language models deployed locally. The pre-trained language models are fine-tuned with corpora from the operation and maintenance field.

[0152] Post-processing is performed on the structured fault diagnosis report to obtain a corrected structured fault diagnosis report. The post-processing includes checking the completeness and accuracy of key fields and correcting structured fault diagnosis reports with missing information or incorrect descriptions.

[0153] Revised structured fault diagnosis reports will be released through multiple channels.

[0154] These multiple channels include at least two of the following: the operations and maintenance backend system, email, instant messaging tools, and the fault ticket system.

[0155] In some embodiments, after the root cause analysis of the abnormal event is completed, in order to transform the technical root cause analysis results into a structured report that can be directly referenced by operation and maintenance personnel and to achieve efficient distribution, a fault diagnosis report needs to be generated and published according to the standard process.

[0156] First, the root cause analysis results and key fields are embedded into a pre-defined prompt template to generate model input data. Key fields, in addition to log level, error code, and exception stack trace, include exception occurrence time, exception duration, root cause node information, and fault propagation path information. These fields comprehensively cover the core information required for fault diagnosis. The pre-defined prompt template pre-defines the paragraph structure and tone of the fault diagnosis report. The paragraph structure typically includes modules such as fault overview, exception details, root cause analysis, handling suggestions, and impact scope description. The tone is set to a professional, concise, and precise style specific to operation and maintenance scenarios, avoiding vague expressions.

[0157] For example, a preset prompt template could be:

[0158] Failure time: {time}

[0159] Scope of impact: {tenants / systems}

[0160] Abnormal phenomenon: {symptom description}

[0161] Root cause analysis: {Root cause path and cause}

[0162] Recommendations: {Measures Implemented / Follow-up Recommendations}

[0163] During the embedding process, the details of the root cause nodes and the fault propagation path links in the root cause analysis results, along with information such as the time of occurrence of the anomaly (accurate to the second) and the duration of the anomaly (calculated from the time span from the triggering of the anomaly to the present) in the key fields, are filled one by one according to the field mapping rules of the prompt template to form structured model input data, ensuring the integrity and format standardization of the input information.

[0164] Subsequently, a Large Language Model (LLM) is invoked to process the model input data (including summaries of anomalies, fluctuations in relevant metrics, key log information, identified root cause components, scope of impact, and suggested remedial measures) to generate a structured fault diagnosis report. The LLM can be either a cloud-based LLM or a pre-trained language model deployed locally. The locally deployed pre-trained language model requires fine-tuning with an operations and maintenance (O&M) corpus, which includes historical fault diagnosis reports, O&M professional documents, and fault handling specifications, ensuring the model can accurately understand O&M terminology and generate report content that conforms to industry conventions.

[0165] During the invocation process, model input data is passed to the large language model via API interfaces or localized service calls. Based on the structure and style defined by the prompt template, the model organizes, refines, and logically organizes the input information, summarizing key information from anomaly analysis into easily understandable textual descriptions. For example, the fault overview module summarizes core anomaly information, the anomaly details module details anomaly indicators and deviation levels, the root cause analysis module clearly describes the root cause node representation and propagation path, and the handling suggestions module provides targeted solutions based on operational experience. The generated structured fault diagnosis report must be logically coherent, informationally complete, and professionally worded to meet the needs of operations personnel for quickly locating and handling faults.

[0166] For example, please generate a diagnostic report based on the following fault information:

[0167] Time: 2025-11-24 10:05

[0168] Tenant: Company X

[0169] System: Order Processing Service

[0170] Abnormal phenomenon: The average response time of the order service increased from 200ms to 2s, and a large number of requests timed out; the log showed a "Database connection timeout" error.

[0171] Root cause analysis: The CPU utilization of the DB-01 node was 100%, causing query delays; DB-01 is part of the order service backend, and the surge in CPU usage caused a request backlog.

[0172] Measures have been taken: Database node DB-01 has been automatically restarted, and performance has returned to normal.

[0173] Please provide a fault diagnosis report.

[0174] After generating a structured fault diagnosis report, post-processing can be performed to check the completeness and accuracy of key fields, resulting in a revised structured fault diagnosis report. The completeness check focuses on key fields such as the time of anomaly occurrence, root cause node information, and fault propagation path information, confirming whether any information is missing (e.g., the duration of the anomaly is not clearly defined, or root cause node attributes are missing). The accuracy check verifies the use of technical terminology, data format specifications, and the rationality of logical relationships. For example, it verifies whether the description of the fault propagation path is consistent with the root cause analysis results, whether the anomaly indicator values ​​match the original data, and whether there are any errors in the expression of technical terminology.

[0175] For reports with missing information, supplement the missing content based on the root cause analysis results and original key fields; for reports with incorrect descriptions, correct them in accordance with the operation and maintenance domain standards (such as correcting the description of "root cause node" as "cause node") to ensure the accuracy and reliability of the report information.

[0176] Finally, revised structured fault diagnosis reports are published through multiple channels to ensure that operations and maintenance (O&M) personnel receive fault information promptly. These channels include at least two of the following: the O&M backend system, email, instant messaging tools, and the fault ticketing system. The publication process can be tailored to the severity of the fault: for urgent faults (such as core business interruptions), alerts and report links are simultaneously pushed through instant messaging tools (such as WeChat Work and DingTalk), and a "urgent" priority ticket is created in the fault ticketing system to ensure immediate response from O&M personnel; for ordinary faults, reports are published through the O&M backend system, and email notifications are sent to relevant personnel, while the corresponding business module's O&M logs are updated. After publication, the report's publication status and read feedback are recorded to ensure effective reach of fault information and support rapid collaborative fault handling.

[0177] The above embodiments can reduce the burden on operations and maintenance personnel to manually compile and analyze reports, automating the alarm-to-report process. The generated diagnostic content covers the fault background, symptoms, root cause localization process, and recommended measures, ensuring the completeness and accuracy of the information. Simultaneously, the large language model can generate professional descriptions that conform to the operational context, improving the readability and standardization of anomaly reports. In scenarios with frequent faults or difficulties in cross-team communication, automated diagnostic reports help all roles obtain key information in a timely manner, supporting rapid collaborative handling.

[0178] In some embodiments, after tracing the failure impact path and locating the root cause of abnormal events based on the operation and maintenance knowledge graph and structured causal model to obtain the root cause analysis results, the method further includes:

[0179] Based on the root cause node type and fault type in the root cause analysis results, the corresponding automated handling strategy is matched from the preset handling strategy library. The automated handling strategy includes at least one of service restart, resource expansion, traffic switching, configuration rollback and fault isolation.

[0180] The automated operation and maintenance platform executes the automated handling actions corresponding to the automated handling strategies.

[0181] After completing the automated processing actions, acquire new multi-source operation and maintenance data;

[0182] The effectiveness of automated handling actions was verified based on new multi-source operation and maintenance data.

[0183] In some embodiments, after obtaining the root cause analysis results, in order to quickly respond to the fault and reduce the impact of the fault on the system operation, it is necessary to perform automated processing based on the root cause information and verify the effect.

[0184] First, based on the root cause node type and fault type in the root cause analysis results, corresponding automated handling strategies are matched from the preset handling strategy library. Root cause node types include configuration items, microservices, host nodes, databases, middleware, etc., and fault types cover resource exhaustion, configuration anomalies, service unavailability, network interruption, etc. The preset handling strategy library pre-stores standardized handling solutions corresponding to different combinations of root cause node types and fault types. Automated handling strategies specifically include at least one of the following: service restart, resource expansion, traffic switching, configuration rollback, and fault isolation.

[0185] For example, if the root cause node type is "microservice" and the fault type is "service unresponsive," then the "service restart" strategy is matched; if the root cause node type is "host node" and the fault type is "CPU resource exhaustion," then the "resource expansion" strategy is matched; if the root cause node type is "configuration item" and the fault type is "parameter configuration error," then the "configuration rollback" strategy is matched. Specifically, the matching process is implemented through keyword mapping and rule matching algorithms to ensure that the most suitable automated handling strategy for the current fault is quickly located.

[0186] Subsequently, the automated operation and maintenance (O&M) platform executes the corresponding automated handling actions based on the matched automated handling strategies. The automated O&M platform supports integration with management interfaces of various system components (such as service registry, container orchestration platform, cloud server management API, and database management tools), and can translate automated handling strategies into executable operation instructions. For example, when executing the "service restart" strategy, the platform sends a restart command through the service registry interface to restart the target microservice instance; when executing the "resource expansion" strategy, it requests additional CPU or memory resources through the cloud server management API, or adjusts the number of Pod replicas and resource limits through the container orchestration platform; when executing the "traffic switching" strategy, it routes traffic from the faulty node to the backup node through the management interface of the load balancer or service gateway; when executing the "configuration rollback" strategy, it calls the configuration management system interface to restore erroneous configuration items to a historical, normal version; when executing the "fault isolation" strategy, it cuts off the connection between the faulty node and other nodes through a network firewall or service governance framework to prevent the fault from spreading. During execution, the automated O&M platform monitors the execution status of the handling actions in real time and records execution logs (such as execution time, operation steps, and return results) to ensure that the handling actions proceed as expected.

[0187] After the automated handling actions are completed, multi-source operation and maintenance data from the system operation process are re-collected according to the multi-source operation and maintenance data acquisition specifications, i.e., new multi-source operation and maintenance data. The scope of the new data collection is consistent with the initial multi-source operation and maintenance data, including time-series data of indicators (such as CPU utilization and response time of the root cause node and related nodes), application logs (such as service operation logs and error logs), alarm events (such as whether new alarms are generated or whether existing alarms are cleared), etc. The collection frequency and preprocessing process also follow the initial configuration to ensure that the new data is comparable to the original data and can objectively reflect the system status after the handling actions are performed.

[0188] Finally, based on the new multi-source operation and maintenance data, the effectiveness of the automated handling actions is verified. For example, the verification process revolves around whether the fault has been resolved and whether the system has returned to normal operation, and whether utilization (response time of abnormal nodes) has returned to the normal threshold range; secondly, application logs are checked to confirm whether error logs and exception stack information related to the fault are no longer generated; thirdly, alarm events are checked to verify whether the original alarms have been automatically cleared and no new related alarms have been triggered; finally, combined with system business indicators (such as request success rate and transaction volume), the business has returned to normal after the fault handling is evaluated. If the new data shows that the fault-related indicators have returned to normal, there are no related error logs and alarms, and the business is running stably, then the automated handling actions are deemed to have been effective; if the fault-related problems still exist (such as indicators not meeting the standards, error logs continuing to be generated), then the handling is deemed ineffective, and the root cause needs to be re-analyzed or the handling strategy adjusted, and manual intervention should be triggered if necessary. Through effectiveness verification, it is ensured that the fault is effectively resolved and the stable operation of the system is guaranteed.

[0189] In some embodiments, the method further includes:

[0190] Obtain performance metrics and external feedback for the anomaly analysis process. Performance metrics include false alarm rate, false negative rate, root cause localization accuracy, fault location time, automatic repair success rate, and alarm response time. External feedback includes manual correction feedback and feedback on the effectiveness of handling.

[0191] Construct a reinforcement learning model and define its system state, action space, and reward function. The system state includes performance metrics and current system parameter configuration. The action space includes anomaly detection threshold adjustment, causal inference confidence threshold adjustment, and automated handling strategy activation state adjustment. The reward function is designed based on the improvement in anomaly detection accuracy and the reduction in fault recovery time.

[0192] The reinforcement learning model trained using reinforcement learning algorithms pre-generates multiple automated handling policies. The reinforcement learning algorithms include Q-learning algorithm and deep deterministic policy gradient algorithm.

[0193] In some embodiments, in order to continuously optimize the accuracy of the anomaly analysis process and the efficiency of automated handling, this embodiment learns the system's operating rules through a reinforcement learning model, dynamically adjusts key parameters, and generates optimized automated handling strategies.

[0194] First, we obtain performance metrics and external feedback for the anomaly analysis process to provide data support for model training. Performance metrics cover the core evaluation dimensions of the entire anomaly analysis process, specifically including: false positive rate (the proportion of false anomalies to total detected events), false negative rate (the proportion of undetected real anomalies to total real anomalies), root cause localization accuracy (the proportion of events whose root cause determination matches the actual fault source), fault localization time (the time span from anomaly detection to the completion of root cause analysis), automatic repair success rate (the proportion of automated actions that successfully resolve the fault), and alarm response time (the time from alarm triggering to the initiation of the handling process). These metrics are automatically calculated by statistically analyzing the anomaly analysis system's operation logs and handling records, and can objectively reflect the process's operational effectiveness. External feedback includes manual correction feedback and feedback on the effectiveness of handling: manual correction feedback comes from the records of manual corrections made by operations and maintenance personnel to anomaly detection results and root cause analysis results (such as marking false alarms and correcting root cause nodes), while feedback on the effectiveness of handling is the subjective evaluation of the automated handling results by operations and maintenance personnel (such as completely resolved, partially resolved, or not resolved). It is collected through the feedback entry of the operations and maintenance backend system to supplement the subjective experience and detailed information that cannot be covered by quantitative indicators.

[0195] Subsequently, a reinforcement learning model is constructed, and the three core elements of the model are defined: system state S, action space A, and reward function R. The system state S is defined as the set of performance metrics and current system parameter configurations. Performance metrics may include measures such as false positive rate, false negative rate, average fault location time, and automatic repair success rate. The current system parameter configurations include anomaly detection thresholds (such as the collaborative decision threshold for variational autoencoders and isolated forests), causal inference confidence thresholds (such as the causal influence scoring threshold in root cause inference), and the activation status of automated handling strategies (such as the on / off status of various handling strategies). By integrating this information into a high-dimensional state vector, the model can comprehensively perceive the current system operating state and parameter configuration.

[0196] The action space A is defined as the set of adjustment operations that the model can execute. Specifically, it includes adjusting anomaly detection thresholds (e.g., fluctuating dynamic thresholds within a preset range), adjusting causal inference confidence thresholds (e.g., adjusting the minimum scoring standard for root cause node determination), and adjusting the activation status of automated handling strategies (e.g., enabling new handling strategies and disabling inefficient ones). Each action corresponds to clear parameter modification rules and range limitations to ensure the safety and rationality of action execution. The reward function R is designed based on the improvement in anomaly detection accuracy and the reduction in fault recovery time. The improvement in anomaly detection accuracy is calculated as (accuracy after adjustment - accuracy before adjustment), and the reduction in fault recovery time is calculated as (average recovery time before adjustment - average recovery time after adjustment). A penalty term is also introduced (e.g., deducting rewards when the false positive rate increases or the automatic repair success rate decreases). The final reward value is a weighted sum of positive returns and penalty terms, guiding the model to learn optimization actions that improve analysis accuracy and handling efficiency.

[0197] Finally, a reinforcement learning model was trained using reinforcement learning algorithms, and multiple optimized automated response strategies were pre-generated. The reinforcement learning algorithms included Q-learning and Deep Deterministic Policy Gradient (DDPG) algorithms. The appropriate algorithm was selected based on the dimensionality of the system's state space: if the state and action spaces were discrete, Q-learning was used, storing the value of state-action pairs in a Q-table and updating the Q-value based on temporal difference learning to gradually optimize the action selection strategy; if the state and action spaces were continuous, DDPG was used, outputting continuous actions through an Actor network and evaluating the action value through a Critic network. Combined with an experience replay mechanism and target network updates, the training stability and convergence speed of the model were improved. During training, the acquired performance metrics and external feedback were used as training data. The model interacted with the environment (anomaly analysis system), continuously trying different actions (parameter adjustment), and adjusting the action selection strategy based on the feedback from the reward function, gradually learning the optimized strategy that maximizes the reward value. After training, the model can pre-generate multiple adapted automated handling strategies (such as optimized handling combinations for different root cause types) based on the current system state, and store them in the preset handling strategy library, providing more efficient and accurate strategy support for subsequent fault handling, and realizing self-optimization of anomaly analysis and handling processes.

[0198] In this embodiment, a reinforcement learning-based policy adjustment mechanism can be used to continuously self-optimize based on feedback, gradually improving the sensitivity and accuracy of anomaly analysis. For example, a reinforcement learning agent can be used to automatically adjust these key parameters based on feedback signals collected during operation, achieving a feedback evolution that becomes increasingly intelligent with use. Specifically, the system state is defined as the performance indicator of the anomaly analysis process (e.g., recent false positive rate, false negative rate, fault location time, automatic handling success rate, etc.), and threshold control and policy parameters are used as the action space of the reinforcement learning agent. For example, actions such as increasing or decreasing the anomaly detection sensitivity threshold, relaxing or tightening the causal inference confidence threshold, and triggering automatic handling can be implemented.

[0199] The agent updates its policy based on feedback after each action. The feedback reward is designed to comprehensively measure improvements in metrics such as detection accuracy, alarm effectiveness, and fault recovery time, encouraging the policy to evolve towards reducing false alarms and missed alarms and accelerating recovery. Q-learning algorithms can be used to optimize the selection of discrete actions (such as different threshold levels), or deep reinforcement learning algorithms such as DDPG (Deep Deterministic Policy Gradient) can be used to handle the continuous action space (fine-tuning threshold values). Reinforcement learning allows the system to learn the optimal policy through continuous trial and error, achieving intelligent dynamic adjustment of alarm thresholds.

[0200] For scenarios with fewer discrete actions, the Q-table can maintain value estimates for each state-action relationship and continuously update the Q-value by trying new actions through an ε-greedy strategy. For example, for single-parameter optimization such as CPU utilization alarm thresholds, the parameters can be discretized into several levels, and the optimal level can be found by Q-learning.

[0201] For scenarios involving multiple parameters and continuous actions, deep neural networks are introduced to approximate the policy or value function. For example, when using the DDPG algorithm, an Actor network outputs the adjustment amounts of each adjustable parameter, while a Critic network evaluates the Q-value of the action given a state. During training, historical running data and online sampling data are used to update the network weights, optimizing the policy in the direction of maximizing cumulative reward.

[0202] In the initial training phase, for safety reasons, the model can be trained offline on real-world environments or historical replay data. This means it doesn't directly affect the real system but simulates the impact of different parameters on metrics based on historical records to accelerate convergence. Once the model is relatively stable, it can be gradually deployed to an online environment for minor adjustments to ensure that exploratory actions do not cause significant disturbances to the real-time system.

[0203] It's worth noting that optimizing the anomaly detection threshold can significantly improve alarm accuracy, overcoming the shortcomings of fixed thresholds in adapting to changes. Over time, it can automatically optimize detection sensitivity based on historical performance (e.g., increasing the threshold to reduce false alarms, or lowering the threshold to capture more anomalies), dynamically adjust the causal inference confidence judgment criteria, and improve automatic handling strategies (e.g., more proactively and automatically restarting the service when proven effective, otherwise conservatively prompting manual review). This closed-loop feedback learning mechanism ensures that system performance continuously improves under changing business scenarios, reducing manual intervention and repeated parameter tuning. When new anomaly patterns emerge or the environment changes, it can quickly adjust itself according to reinforcement learning strategies, ensuring that anomaly detection and localization always maintain high accuracy and timeliness.

[0204] In some embodiments, a reinforcement learning model is trained using a reinforcement learning algorithm, including:

[0205] Offline training of reinforcement learning models is performed based on historical operational data to simulate the impact of different model parameter configurations on anomaly analysis performance and accelerate model convergence.

[0206] The offline-trained model is deployed to an online environment to train the offline-trained model online using amplitude adjustment and effect verification mechanisms;

[0207] During the online training of the offline-trained model, the effect of parameter adjustment in each iteration is recorded;

[0208] The reward value, calculated based on the reward function and parameter adjustment effect, is used to continuously update the model parameters of the reinforcement learning model.

[0209] When training a reinforcement learning model using reinforcement learning algorithms, a two-stage training mode of "offline training + online training" is adopted to balance model training efficiency and adaptability to the real environment, ensuring that the model can converge quickly and output optimization strategies that fit the actual scenario. First, the reinforcement learning model is trained offline based on historical running data to accelerate model convergence.

[0210] Historical operational data encompasses past performance metrics, external feedback, system parameter configurations, and corresponding results for anomaly analysis processes. This includes data on false positive rates, root cause localization accuracy, and automated handling success rates for different anomaly types and system load scenarios, as well as records of the impact of various parameter adjustments on analysis performance. During offline training, a virtual training environment is constructed by simulating anomaly analysis performance changes corresponding to different model parameter configurations (such as different combinations of anomaly detection thresholds and causal inference confidence thresholds) using historical data. Historical system states are used as model input, allowing the model to attempt different actions (parameter adjustments) in the virtual environment. Reward values ​​are calculated based on performance changes recorded in historical data, thereby updating model parameters. In this way, the model can quickly learn optimization patterns from historical data, avoiding system performance fluctuations caused by random parameter adjustments in the early stages of online training, significantly shortening model convergence time, and laying a solid foundation for online training.

[0211] After offline training, the model is deployed to a real online environment. A magnitude adjustment and performance verification mechanism are used for online training to ensure the model is adapted to real-world scenarios. The magnitude adjustment mechanism strictly controls the magnitude of parameter adjustments during online training (e.g., anomaly detection threshold adjustments are no more than ±10% each time, and the step size for causal inference confidence threshold adjustments is set to 0.05). This prevents anomaly analysis processes or automated handling strategies from failing due to sudden parameter changes, ensuring system stability. The performance verification mechanism reserves a preset observation period (e.g., 5 minutes, 10 minutes) after each parameter adjustment. During this period, real-time system data is collected to evaluate the actual impact of parameter adjustments on core metrics such as anomaly detection accuracy and fault recovery time, ensuring the effectiveness of the adjustments is quantifiable and traceable. During online training, the model dynamically selects parameter adjustment actions based on real-time system status (current performance metrics, system parameter configuration) and strategies learned during offline training, achieving continuous optimization of the model in the real-world environment.

[0212] During online training, the effects of parameter adjustments in each iteration are meticulously recorded to provide data support for model parameter updates. The records include the adjustments made in this iteration (e.g., increasing the anomaly detection threshold by 8%, disabling automated handling strategy A), the system state parameters before the adjustment, the changes in performance metrics after the adjustment (e.g., a 3% decrease in false alarm rate, a 20-second reduction in fault location time), and the corresponding reward values ​​for the adjustments, forming a complete iterative training log. These records not only visually reflect the actual effects of each adjustment but also provide a basis for subsequent analysis of model optimization patterns and troubleshooting of training anomalies, ensuring the traceability and controllability of the online training process.

[0213] Finally, the model parameters of the reinforcement learning model are continuously updated based on the reward value calculated according to the reward function and the effect of parameter adjustment. After each iteration, based on the preset reward function and the effect of parameter adjustment in this iteration (such as the improvement in anomaly detection accuracy, the reduction in fault recovery time, and the change in false positive rate), the reward value corresponding to this action is calculated: if the adjustment of the model optimizes the indicators of the reinforcement learning model (such as improved accuracy and reduced recovery time), a positive reward is obtained; if the adjustment of the model deteriorates the indicators of the reinforcement learning model (such as increased false positive rate and decreased handling success rate), a negative reward is obtained. The calculated reward value is fed back to the reinforcement learning model. The model adjusts the internal weight parameters and policy selection probability according to the update rules of the reinforcement learning algorithm (such as temporal difference update of Q-learning and Actor-Critic network update of deep deterministic policy gradient), so that the model is more inclined to choose adjustment actions that can obtain higher reward values. Through multiple rounds of iterative training, the model parameters are continuously optimized, gradually forming an optimal policy adapted to the online environment, ensuring that the output automated handling strategy can continuously improve the performance of anomaly analysis and handling processes.

[0214] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0215] Corresponding to the anomaly analysis method for multi-source operation and maintenance data in the above embodiment, Figure 6 This is a schematic diagram of the structure of a multi-source operation and maintenance data anomaly analysis device provided in an embodiment of this application. This device can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be... Figure 7 The electronic device shown.

[0216] Reference Figure 6 The anomaly analysis device for multi-source operation and maintenance data includes, in performing any one of the anomaly analysis methods for multi-source operation and maintenance data, the following:

[0217] The acquisition unit 601 is used to acquire multi-source operation and maintenance data, which includes at least two of the following: time series data of indicators, application logs, call chain tracing data, configuration change records, alarm events, and work orders.

[0218] The processing unit 602 is used to perform joint anomaly modeling on the preprocessed multi-source operation and maintenance data based on a multi-model fusion architecture, so as to identify abnormal events in the multi-source operation and maintenance data.

[0219] Analysis unit 603 is used to trace the path of failure impact and locate the root cause of abnormal events based on the operation and maintenance knowledge graph and the structured causal model, and obtain the root cause analysis results.

[0220] The root cause analysis results are used to indicate the root cause nodes and propagation paths of faults in abnormal events.

[0221] It is understood that the embodiments of the anomaly analysis device for multi-source operation and maintenance data and any implementation thereof correspond to the embodiments of the anomaly analysis method for multi-source operation and maintenance data and any implementation thereof. The technical effects corresponding to the embodiments of the anomaly analysis device for multi-source operation and maintenance data and any implementation thereof can be found in the above-mentioned embodiments of the anomaly analysis method for multi-source operation and maintenance data and any implementation thereof, and will not be repeated here.

[0222] It should be noted that the anomaly analysis device for multi-source operation and maintenance data provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0223] The functional units and modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of this application.

[0224] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0225] This application also provides an electronic device, which includes one or more processors and a memory;

[0226] The memory is coupled to one or more processors. The memory is used to store computer program code, which includes computer instructions. One or more processors call the computer instructions to cause the electronic device to execute the aforementioned method for anomaly analysis of multi-source maintenance data.

[0227] Figure 7This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 700 can be a mobile phone, smart screen, tablet computer, wearable electronic device, in-vehicle electronic device, augmented reality (AR) device, virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), projector, or a server, storage device, base station, or other communication device, or a smart car, etc. This application embodiment does not impose any limitations on the specific type of electronic device.

[0228] The memory 701 can be used to store computer software programs 702 and modules. The processor 703 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 701. The memory 701 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device (such as audio data, telephone directory, etc.). In addition, the memory 701 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0229] The processor 703 may include one or more processors such as a central processing unit (CPU), an application processor (AP), and a baseband processor. The processor can serve as the nerve center and command center of the wireless router. The processor 703 can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. The memory 701 can be used to store executable program code, including instructions. The processor 703 executes various functional applications and data processing of the network device by running the instructions stored in the memory. The memory 701 may include a program storage area and a data storage area, such as storing data for audio signals to be played. For example, the memory may be Double Data Rate Synchronous Dynamic Random Access Memory (DDR) or Flash memory.

[0230] This application also provides a computer-readable storage medium storing computer instructions; when the computer-readable storage medium is used on an electronic device, it causes the electronic device to execute the aforementioned method for anomaly analysis of multi-source operation and maintenance data.

[0231] The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or can include one or more data storage devices such as servers or data centers that can be integrated with media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media, or semiconductor media (e.g., solid-state disks (SSDs)).

[0232] This application also provides a computer program product containing computer instructions, which, when run on an electronic device, enables the electronic device to execute the aforementioned method for anomaly analysis of multi-source operation and maintenance data.

[0233] The computer storage medium and computer program product provided in the embodiments of this application are used to execute the methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects corresponding to the methods provided above, and will not be repeated here.

[0234] In the above embodiments, implementation can also be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line, DSL) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc., and the storage medium can also include combinations of the above types of memory.

[0235] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0236] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments claimed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0237] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0238] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0239] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for anomaly analysis of multi-source operation and maintenance data, characterized in that, include: Acquire multi-source operation and maintenance data, which includes at least two of the following: time-series indicator data, application logs, call chain tracing data, configuration change records, alarm events, and work orders; Based on a multi-model fusion architecture, joint anomaly modeling is performed on the preprocessed multi-source operation and maintenance data to identify abnormal events in the multi-source operation and maintenance data. Based on the operation and maintenance knowledge graph and the structured causal model, the abnormal event is traced to find the source of the fault impact path and the root cause, and the root cause analysis results are obtained. The root cause analysis results are used to indicate the root cause node of the fault and the fault propagation path in the abnormal event.

2. The method according to claim 1, characterized in that, The method of performing joint anomaly modeling on preprocessed multi-source operation and maintenance data based on a multi-model fusion architecture to identify abnormal events in the multi-source operation and maintenance data includes: Construct a graph structure of monitoring objects, where nodes represent monitoring entities at different granularities. These monitoring entities include tenant-level entities, application-level entities, host node entities, database entities, middleware entities, and network device entities. Edges between nodes are established based on the dependencies of the monitoring objects, including service call relationships, deployment relationships, network connection relationships, and resource dependencies. A time-series coding model is used to extract features from the multi-dimensional time-series data corresponding to each node of the graph structure to obtain node time-series feature vectors. The multi-dimensional time-series data includes the indicator time-series data and log event frequency sequence data in the preprocessed multi-source operation and maintenance data. The node temporal feature vectors are fused by topological association using a graph attention network to obtain a node representation that combines temporal features and topological relationships. Anomaly determination is performed on the node representation using variational autoencoder and isolated forest to generate the anomaly event. The anomaly event includes at least one of the following: anomaly time window, involved tenant, involved node, anomaly index, and degree of deviation. The types of the anomaly event include tenant-level anomaly events and node-level anomaly events.

3. The method according to claim 2, characterized in that, The step involves using a time-series coding model to extract features from the multi-dimensional time-series data corresponding to each node of the graph structure, resulting in a node time-series feature vector, including: A long short-term memory network with shared parameters is deployed for each node, and the long short-term memory network is used as the temporal coding model; The multidimensional time series data of each node within a preset time period are input into the corresponding long short-term memory network. By using the hidden layers of the Long Short-Term Memory network, the temporal correlation of the multidimensional time series data of each node within a preset time period is modeled, and the temporal feature vector of the node is output.

4. The method according to claim 2, characterized in that, The step of fusing the temporal feature vectors of the nodes through a graph attention network to obtain a node representation that fuses temporal features and topological relationships includes: The temporal feature vector of each node is input into the graph attention network; The graph attention network calculates the attention coefficient between each node and all its neighboring nodes based on the feature similarity and dependency strength between nodes. The attention coefficient is used to quantify the influence weight of neighboring nodes on the current node. The temporal feature vectors of all neighboring nodes of each node are weighted and summed based on the attention coefficient to achieve a preliminary fusion of the neighboring node representation and the current node representation. The graph attention network structure is stacked in multiple layers. Through iterative propagation of the multiple layers, each node aggregates the node representations of its multi-hop neighbors layer by layer. Each layer of the network captures the feature information of the current node's 1-hop neighbors. The node representation includes the node's own temporal features, neighborhood node association features, and multi-hop topology information.

5. The method according to claim 2, characterized in that, The step of using variational autoencoders and isolated forests to determine anomalies in the node representations and generating the anomalous events includes: The node representation is input into the variational autoencoder, so that the encoding unit of the variational autoencoder maps the node representation to a low-dimensional latent variable distribution, and the decoding unit of the variational autoencoder outputs the reconstructed node representation corresponding to the node representation. Based on the reconstructed node representation and the node representation, the reconstruction error or reconstruction probability is calculated as the first anomaly score; The node representation is input into the isolated forest, and the isolated forest randomly divides the node representation to isolate abnormal data points, thereby obtaining a second anomaly score; A dynamic threshold determination mechanism is adopted, and the anomaly determination result is determined by combining the first anomaly score and the second anomaly score; If the anomaly determination result is abnormal, an anomaly event is generated, which includes the time window in which the anomaly occurred, the tenants involved, the nodes involved, the anomaly indicators and the degree of deviation of the values, the anomaly score, and a list of associated nodes.

6. The method according to claim 1, characterized in that, The method, based on operational knowledge graphs and structured causal models, traces the failure impact path and locates the root causes of the abnormal events, obtaining root cause analysis results, including: Obtain a pre-built and real-time updated operation and maintenance knowledge graph, wherein the nodes of the operation and maintenance knowledge graph include at least one of business applications, microservices, host nodes, databases, middleware, network devices, configuration items, and change events, and the edges of the operation and maintenance knowledge graph include at least one of call relationships, deployment relationships, dependency relationships, network connection relationships, and tenant affiliation relationships. The operation and maintenance knowledge graph is built based on configuration item data, service governance framework data, and historical fault analysis data. Based on the node information involved in the abnormal event, the corresponding abnormal node is located in the operation and maintenance knowledge graph; Fault impact diffusion analysis is performed along the adjacency relationship of the abnormal nodes, and relevant neighbor nodes are screened to form a candidate fault subgraph; Based on the structured causal model, root cause reasoning is performed on the candidate fault subgraph, and root cause analysis results are output to indicate the root cause nodes and fault propagation paths.

7. The method according to claim 6, characterized in that, The fault impact diffusion analysis is performed along the adjacency relationship of the abnormal node, and relevant neighbor nodes are screened to form a candidate fault subgraph, including: A breadth-first search or depth-first search algorithm is used to traverse the adjacency relationship of the operation and maintenance knowledge graph starting from the abnormal node, so as to perform fault impact diffusion analysis on the adjacency relationship of the abnormal node. During the process of traversing the adjacency relationships of the operation and maintenance knowledge graph, the probability of fault propagation corresponding to each adjacency relationship is evaluated by combining the detection results of the abnormal events. Based on the fault propagation probability corresponding to each adjacency relationship, neighbor nodes with a fault propagation probability greater than a preset threshold are selected and together with the abnormal nodes, they form a candidate fault subgraph.

8. The method according to claim 6, characterized in that, The root cause reasoning of the candidate fault subgraph based on the structured causal model, outputting root cause analysis results indicating the root cause nodes and fault propagation paths, includes: Based on historical operation and maintenance data, a causal discovery algorithm is used to learn the causal relationships between entities and form a causal rule base. For each node in the candidate fault subgraph, an initial causal correlation degree is assigned to each node based on the causal rule base; The causal impact score of each node on the abnormal event is calculated using counterfactual reasoning or Bayesian inference. Based on the causal impact score, each node is sorted from high to low, and the node with the highest ranking is determined as the root cause node of the failure. Based on the relationship chain in the operation and maintenance knowledge graph, the fault propagation path from the root cause node to the abnormal node is sorted out to obtain the root cause analysis results.

9. The method according to any one of claims 1 to 8, characterized in that, After acquiring multi-source operation and maintenance data, the method further includes: The multi-source operation and maintenance data are processed by imputing missing values ​​and filtering noise. The multi-source operation and maintenance data is resampled and time-aligned according to a preset time window. The multi-source operation and maintenance data are cleaned to remove irrelevant fields and format characters, and key fields are extracted from the multi-source operation and maintenance data. The key fields include at least one of log level, error code, and exception stack. The key fields are segmented and vectorized to obtain the preprocessed multi-source operation and maintenance data.

10. The method according to claim 9, characterized in that, The key fields also include the time of the anomaly occurrence, the duration of the anomaly, the root cause information of the fault, and the fault propagation path information; the method also includes: The root cause analysis results and the key fields are embedded into a preset prompt template to generate model input data, wherein the prompt template is used to specify the paragraph structure and tone style of the fault diagnosis report; The large language model is invoked to process the input data of the model and generate a structured fault diagnosis report. The large language model includes a cloud-based large language model or a pre-trained language model deployed locally. The pre-trained language model is fine-tuned with corpus data from the operation and maintenance field. The structured fault diagnosis report is post-processed to obtain a corrected structured fault diagnosis report. The post-processing includes checking the completeness and accuracy of the key fields and correcting structured fault diagnosis reports with missing information or incorrect descriptions. The revised structured fault diagnosis report shall be published through multiple channels, including at least two of the following: the operation and maintenance backend system, email, instant messaging tools, and fault ticket system.

11. The method according to any one of claims 1 to 8, characterized in that, After tracing the failure impact path and locating the root cause of the abnormal event based on the operation and maintenance knowledge graph and structured causal model, and obtaining the root cause analysis results, the method further includes: Based on the root cause node type and fault type in the root cause analysis results, a corresponding automated handling strategy is matched from the preset handling strategy library. The automated handling strategy includes at least one of service restart, resource expansion, traffic switching, configuration rollback and fault isolation. The automated handling actions corresponding to the automated handling strategy are executed through the automated operation and maintenance platform; After the automated processing actions are completed, new multi-source operation and maintenance data is acquired; The effectiveness of the automated handling actions is verified based on the new multi-source operation and maintenance data.

12. The method according to claim 11, characterized in that, The method further includes: The effectiveness metrics and external feedback of the anomaly analysis process are obtained. The effectiveness metrics include the false alarm rate, false negative rate, root cause localization accuracy, fault location time, automatic repair success rate, and alarm response time. The external feedback includes manual correction feedback and feedback on the effectiveness of the handling. Construct a reinforcement learning model and define the system state, action space, and reward function of the reinforcement learning model. The system state includes the performance index and the current system parameter configuration. The action space includes the adjustment of the anomaly detection threshold, the adjustment of the causal inference confidence threshold, and the adjustment of the activation state of the automated handling strategy. The reward function is designed based on the improvement of anomaly detection accuracy and the reduction of fault recovery time. The reinforcement learning model trained using reinforcement learning algorithms pre-generates multiple automated handling strategies, wherein the reinforcement learning algorithms include Q-learning algorithm and deep deterministic policy gradient algorithm.

13. The method according to claim 12, characterized in that, The step of training the reinforcement learning model using a reinforcement learning algorithm includes: The reinforcement learning model is trained offline based on historical operating data to simulate the impact of different model parameter configurations on anomaly analysis performance and accelerate model convergence. The offline-trained model is deployed to an online environment to perform online training on the offline-trained model using amplitude adjustment and effect verification mechanisms; During the online training of the offline-trained model, the effect of parameter adjustment in each iteration is recorded; The model parameters of the reinforcement learning model are continuously updated based on the reward value calculated according to the reward function and the parameter adjustment effect.

14. A multi-source monitoring anomaly analysis device, characterized in that, A method for performing anomaly analysis of multi-source operation and maintenance data as described in any one of claims 1 to 13, comprising: The acquisition unit is used to acquire multi-source operation and maintenance data, which includes at least two of the following: indicator time series data, application logs, call chain tracing data, configuration change records, alarm events, and work orders. The processing unit is used to perform joint anomaly modeling on the preprocessed multi-source operation and maintenance data based on a multi-model fusion architecture, so as to identify abnormal events in the multi-source operation and maintenance data. The analysis unit is used to trace the fault impact path and locate the root cause of the abnormal event based on the operation and maintenance knowledge graph and the structured causal model, and obtain the root cause analysis results. The root cause analysis results are used to indicate the fault root cause node and fault propagation path in the abnormal event.

15. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it causes the electronic device to implement the method as described in any one of claims 1 to 13.

16. A computer program product, characterized in that, Includes a computer program, which, when run, causes the method as described in any one of claims 1 to 13 to be performed.

17. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 13.

Citation Information

Cited By

  • Mining equipment fault causal analysis method and equipment based on mapping knowledge domain, and medium

    CN122066411A

  • AI training exception diagnosis method, electronic device, and storage medium

    CN122132218A

  • A Knowledge Graph-Based Causal Analysis Method for Mining Equipment Failures

    CN122134329A

  • A single-node fault operation and maintenance method and system based on multi-algorithm fusion

    CN122173329A

  • A power distribution room operation and maintenance full-link knowledge graph cross-modal alignment method

    CN122334433A