Intelligent operation and maintenance alarm generation method based on unsupervised learning
By constructing a dynamic topology graph and a spatiotemporal neural network to fuse multimodal data and dynamically adjusting thresholds, the accuracy and real-time issues of intelligent operation and maintenance alarms in existing technologies are solved, enabling efficient alarms and root cause localization for distributed systems.
Patent Information
- Application Number
- CN202511394990.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-10-31
AI Technical Summary
Existing intelligent operation and maintenance alarm technologies struggle to capture complex anomaly patterns across services and time periods when dealing with distributed systems. Furthermore, their reliance on manually labeled data and fixed thresholds leads to high false alarm and false negative rates, making it impossible to respond to system status changes in real time.
A dynamic topology graph is constructed, and a spatiotemporal graph neural network model is used to fuse multimodal data. Alarms are generated through unsupervised learning, thresholds are dynamically adjusted, and root cause localization is performed. A cross-modal attention module and an online model adaptive mechanism are utilized.
It improves the accuracy of alarms and the precision of root cause location, has the ability to generalize to unknown faults, reduces false alarm rate and false negative rate, and adapts to dynamic environmental changes.
Smart Images

Figure CN120880926A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent operation and maintenance alarm technology under machine learning, and in particular to an intelligent operation and maintenance alarm generation method based on unsupervised learning. Background Technology
[0002] With the rapid development of information technology, especially with the widespread application of microservices, containerization, and distributed architectures, the complexity of IT systems has increased dramatically, posing serious challenges to their daily operation and maintenance. To improve system reliability and operational efficiency, intelligent operation and maintenance (O&M) technologies have emerged. Alarm prediction, as a core component of intelligent O&M technology, predicts potential anomalies and failures through the analysis and modeling of system monitoring data (such as time-series metrics, logs, and tracing data), and issues alerts before problems occur. This helps O&M teams intervene proactively, reduce the risk of business interruption, and improve system availability and continuity.
[0003] Currently, mainstream intelligent alarm technologies fall into three categories. The first category is based on machine learning models, which learn fault patterns from massive amounts of historical operational data. This can be further subdivided into supervised learning relying on manual annotation and unsupervised learning without labels. The second category is based on expert knowledge bases, which use logical reasoning by constructing a knowledge graph containing fault patterns, handling rules, and repair suggestions. The third category is a hybrid approach, combining machine learning with a knowledge base. For example, it uses a pre-trained large model for initial prediction, and then corrects the results through knowledge base reasoning to improve accuracy. For instance, CN120596341A discloses an operational alarm prediction method, device, equipment, storage medium, and product. This solution employs a large model trained through supervised and unsupervised iterative training, combined with an automatically generated operational knowledge base (containing fault pattern rules, anomaly handling rules, and repair suggestions), and generates the final alarm prediction result through a Bayesian posterior probability fusion strategy. This method, to some extent, addresses the rigidity of rules in traditional expert systems and improves adaptability.
[0004] However, even with this approach, significant limitations remain. First, supervised or hybrid learning-based solutions heavily rely on large-scale, high-quality manually labeled data. In operational scenarios, obtaining this data is often costly and struggles to comprehensively cover emerging or unknown fault types. Furthermore, while knowledge base reasoning can automatically generate rules, it suffers from inherent rigidity, failing to respond in real-time to rapidly changing system states and the emergence of new fault modes. Second, these methods fail to fully exploit and utilize the dynamic topological dependencies between service units within the system when processing operational data. They typically treat service metrics or logs as isolated data streams, ignoring the cascading propagation of faults along service call chains in distributed architectures. This lack of modeling the system's "spatiotemporal" characteristics makes it difficult for models to capture complex anomaly patterns across services and time periods, thus affecting the accuracy of root cause localization and overall robustness of alerts. In addition, existing multimodal data fusion mechanisms are usually shallow and difficult to reveal deep semantic relationships between different data sources. Furthermore, the setting of alarm thresholds often relies on fixed values and simple statistics, lacking adaptive capabilities, which leads to excessively high false alarm and false negative rates, especially under dynamic environments (such as system updates).
[0005] Therefore, in view of the above problems, there is an urgent need for an intelligent operation and maintenance alarm generation method to improve the accuracy, real-time performance and generalization ability of alarms. Summary of the Invention
[0006] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.
[0007] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides an intelligent operation and maintenance alarm generation method based on unsupervised learning to solve the problems mentioned in the background art.
[0008] To address the aforementioned technical problems, this invention provides the following technical solution: an intelligent operation and maintenance alarm generation method based on unsupervised learning, comprising:
[0009] Based on operational data, a dynamic topology graph representing each service node in the system and their interrelationships is constructed.
[0010] For each service node in the dynamic topology graph, generate a node state vector representing its current state;
[0011] The structure of the dynamic topology graph and the node state vectors of all service nodes are input into a pre-trained graph neural network model.
[0012] The anomaly score for each service node is calculated using the graph neural network model.
[0013] An operation and maintenance alarm is generated when the abnormal score of any service node exceeds the preset alarm threshold.
[0014] As a preferred embodiment of the intelligent operation and maintenance alarm generation method based on unsupervised learning described in this invention, the operation and maintenance data includes at least indicator data, log data, and link tracing data, and the dynamic topology map is constructed based on the link tracing data.
[0015] As a preferred embodiment of the intelligent operation and maintenance alarm generation method based on unsupervised learning described in this invention, wherein: for each service node in the dynamic topology graph, a node state vector representing its current state is generated, including:
[0016] The node status vector is generated by integrating the metric data and log data corresponding to each service node.
[0017] As a preferred embodiment of the intelligent operation and maintenance alarm generation method based on unsupervised learning described in this invention, the specific method of fusion includes:
[0018] Based on a time series model, time-series features are extracted from the indicator data;
[0019] Based on a natural language processing model, semantic features are extracted from the log data;
[0020] A cross-modal attention module is applied, which uses the temporal features as queries and the semantic features as keys and values, calculates attention weights and weights the semantic features, and combines the weighted semantic features with the temporal features to generate the node state vector.
[0021] As a preferred embodiment of the intelligent operation and maintenance alarm generation method based on unsupervised learning described in this invention, the graph neural network model is a spatiotemporal graph neural network model, which processes the input node state vector by alternately executing spatial graph convolution operations and temporal convolution operations.
[0022] As a preferred embodiment of the intelligent operation and maintenance alarm generation method based on unsupervised learning described in this invention, the graph neural network model is pre-trained through a meta-learning framework to obtain a set of meta-model parameters. When a new service node joins the system, the meta-model parameters are fine-tuned using some initial data from the new service node.
[0023] As a preferred embodiment of the intelligent operation and maintenance alarm generation method based on unsupervised learning described in this invention, the anomaly score is calculated based on the reconstruction error between the prediction result of the graph neural network model of the current node state vector and the actual node state vector.
[0024] In a preferred embodiment of the intelligent operation and maintenance alarm generation method based on unsupervised learning described in this invention, the alarm threshold is dynamically determined, and the dynamic determination step includes:
[0025] Obtain historical outlier score data and select a high quantile point;
[0026] The generalized Pareto distribution model was used to fit the outlier scores that exceeded the high quantile points;
[0027] Based on the fitted generalized Pareto distribution model and the preset alarm rate, the current alarm threshold is calculated.
[0028] As a preferred embodiment of the intelligent operation and maintenance alarm generation method based on unsupervised learning described in this invention, after generating the operation and maintenance alarm, a root cause localization step is further included, the step comprising:
[0029] Among the multiple service nodes that trigger the alarm, the node that exceeds the alarm threshold earliest in time and has the highest anomaly score will be identified as the root cause node.
[0030] The fault propagation path is determined by calculating the gradient of the anomaly score of an abnormal service node relative to the state vector of its upstream neighbor nodes.
[0031] As a preferred embodiment of the intelligent operation and maintenance alarm generation method based on unsupervised learning described in this invention, the method further includes an online model adaptation step, which includes:
[0032] Continuously monitor the overall anomaly score distribution of the system;
[0033] When a predetermined significant change in the score distribution is detected, the graph neural network model is incrementally learned online using recent data.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] 1. By constructing a dynamic topology graph and using a spatiotemporal graph neural network model to model the system state as a whole, it is possible to capture the spatial dependency relationships and temporal evolution characteristics between service nodes in the system from a global perspective. This overcomes the limitations of existing technologies that treat each service as an isolated data stream and ignore the characteristics of fault propagation along the service call chain. As a result, it significantly improves the ability to identify complex, cross-service cascading fault modes in distributed systems, and fundamentally improves the accuracy of alarms and the precision of root cause location.
[0036] 2. A multimodal data deep fusion mechanism based on a cross-modal attention module was designed. This mechanism can dynamically learn the intrinsic relationship between the temporal features in the indicator data and the semantic features in the log data, and perform adaptive weighting according to the strength of the relationship. It effectively overcomes the information loss problem caused by simple splicing or shallow fusion in the existing technology, and generates a node state vector that accurately reflects the real running status of the service node, providing a feature basis for unsupervised anomaly detection.
[0037] 3. By adopting unsupervised learning, the reliance on expensive and hard-to-obtain manually labeled data is eliminated, and the ability to generalize to unknown and emerging fault types is achieved. At the same time, a dynamic alarm threshold determination method based on extreme value theory is introduced, which enables the alarm threshold to be adaptively adjusted according to the real-time distribution of system anomaly scores. This solves the problem of high false alarm and false negative rates of traditional fixed thresholds when facing system state fluctuations. In addition, the online model adaptive mechanism ensures the high robustness of the present invention in dynamically changing IT system environments. Attached Figure Description
[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0039] Figure 1 This is a flowchart illustrating the overall process of an intelligent operation and maintenance alarm generation method based on unsupervised learning, as described in one embodiment of the present invention. Detailed Implementation
[0040] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0041] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0042] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0043] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.
[0044] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0045] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0046] Example 1
[0047] Reference Figure 1 This is the first embodiment of the present invention, which provides an intelligent operation and maintenance alarm generation method based on unsupervised learning, including:
[0048] S1. Based on operation and maintenance data, construct a dynamic topology diagram representing each service node in the system and their interrelationships;
[0049] It should be noted that the task of this step is to abstract the logically interconnected but data-dispersed IT systems into a structured graph data structure that can be processed by a mathematical model.
[0050] Specifically, multiple operational data streams from the observability platform in the IT system are accessed in real time through a streaming processing engine (Apache Flink or Spark Streaming);
[0051] Specifically, the operation and maintenance data stream includes at least metric data, log data, and link tracing data;
[0052] It should be noted that the construction of dynamic topology graphs mainly relies on link tracing data, because this data type records the call relationships and request paths between services;
[0053] Furthermore, link tracing data conforming to the OpenTelemetry standard is used to identify all independent service units constituting the system from the link tracing data, and these service units are defined as nodes in the dynamic topology graph.
[0054] It should be noted that the system used in the present invention is a distributed system. The complete processing of a user request will generate a call chain. The call chain consists of multiple spans. Each span represents an operation within a service or a remote procedure call between services. Each span carries key metadata (i.e., service.name (service name) and instance.id (instance ID)).
[0055] Furthermore, after obtaining the key metadata, it is initially cleaned and standardized. For example, the service name convention for different metadata reports is unified (such as normalizing frontend-service and frontend to the same logical node), and known, non-business-related traffic, such as link tracing data generated by health checks or internal monitoring heartbeats, is filtered out to reduce noise in the topology graph.
[0056] Furthermore, by processing all link tracing data within a specific time window (e.g., the past 5 minutes), the service names of the metadata across all spans are extracted. By deduplicating these service names, a unique list of services is obtained. Then, each unique service name in the system is abstracted into a service node. All service nodes then constitute the node set of the dynamic topology graph. Where n is the total number of active service nodes within the current time window;
[0057] It should be noted that after determining the nodes of the dynamic topology graph, it is also necessary to construct the edges of the graph based on the actual interactions between services. The construction of the edges is also based on the deep analysis of the link tracing data, especially the parent-child relationship between spans. That is, in a call chain, if span A is the parent span of span B, it means that the operation represented by span A triggered the operation represented by span B.
[0058] Furthermore, for asynchronous communication via message queues (such as Kafka and RabbitMQ), the parent-child span relationship may not exist directly. To address this issue, this invention preferably employs a context propagation mechanism to infer such dependencies. That is, when the producer service sends a message to the message queue, it injects the current tracing data into the message header. When the consumer service receives the message, it extracts the tracing data and starts a new, logically related span. In other words, by associating producer and consumer spans with the same tracing data but belonging to different call chain segments in the streaming engine, the asynchronous dependency relationship between services can be accurately inferred, and the edges of the corresponding dynamic topology graph can be constructed.
[0059] Specifically, by iterating through each span B within the time window, it checks if there is a parent span A. If so, it retrieves the service names of spans A and B respectively. If these two service names are different (i.e., service nodes), it checks if the parent span A exists. This indicates the service node. To service node A call was initiated, and based on this, a directed edge was constructed in the dynamic topology graph from one service node to another. By repeating the above operation of processing all parent-child span relationships within the time window, a complete service call relationship can be constructed, forming the edge set E of the dynamic topology graph;
[0060] Furthermore, simply establishing binary connections is insufficient. To allow the dynamic topology graph to carry richer information, the present invention assigns a multi-dimensional attribute vector to each directed edge.
[0061] Specifically, this attribute vector is calculated using the information retrieved by all service nodes within the aggregation window that point to the service nodes. The calculation includes the following:
[0062] Call count: This counts the total number of calls from service nodes to other service nodes within the window, reflecting the frequency of interactions between services.
[0063] Average / quantile latency is a statistic that calculates the response time of all service node calls to service nodes, such as the average, P95 or P99 quantile, to reflect the performance of dependent services.
[0064] Error rate is the percentage of calls with an ERROR status code (i.e., service node error call) out of the total number of calls, used to reflect the health status of inter-service interactions.
[0065] It should be noted that the topology of IT systems, especially those based on microservices and cloud-native technologies, is not static. Instead, service instances often change dynamically due to factors such as automatic scaling, version releases, and failover. To ensure that the dynamic topology map can accurately reflect the current operating structure of the system, this invention adopts an update mechanism based on a sliding time window to solve this problem.
[0066] Specifically, a time window length and a sliding step size are set, and all link tracing data within the time period of the most recent time window length are acquired periodically (every sliding step size), and the node identification and edge construction process described above is repeated to generate a new dynamic topology graph.
[0067] In addition, to avoid unstable "spurt" edges caused by instantaneous and accidental calls, this invention introduces an edge confirmation threshold. When the number of calls between two services exceeds a certain preset threshold (e.g., 10 times) within a time window, or when the call relationship occurs in multiple consecutive time windows, the edge will be confirmed as stably existing in the dynamic topology graph, thereby enhancing the robustness of the dynamic topology graph.
[0068] Specifically, the final output dynamic topology graph at each time step (time t) is... It is expressed as follows:
[0069]
[0070] in, It is the set of service nodes within that time window. It is the set of directed edges between nodes. It is a tensor that stores the multidimensional attribute vectors of all edges;
[0071] S2. For each service node in the dynamic topology graph, generate a node state vector representing its current state.
[0072] It should be noted that after the dynamic topology graph is constructed, a high-dimensional, information-dense numerical representation, namely the node state vector, needs to be generated for each service node in the topology graph. This vector aims to quantify the running state of the service node within the time window.
[0073] Specifically, the node state vector is generated by fusing the metric data and log data corresponding to the service node;
[0074] Furthermore, for each service node in the dynamic topology graph, its metric data and log data within the time window length are first retrieved and aligned from the operation and maintenance data stream.
[0075] Specifically, key performance indicators (KPIs) related to the service nodes are acquired as indicator data. This indicator data includes, but is not limited to, CPU utilization, memory consumption, network throughput (input / output), service request rate (QPS / RPM), request error rate, and P95 / P99 response latency. The above indicator data is then organized into a multivariate time series matrix. Where H represents the number of indicator data. The number of sampling points within the time window length;
[0076] Specifically, all log entries generated by the service node within the time window are obtained, and the original semi-structured log text is parsed.
[0077] Specifically, in one implementation of this invention, the Drain online log parsing algorithm is used to map each log entry to a fixed log template and a set of dynamically changing parameters, transforming unstructured text data into a structured event sequence. Then, all log events within the time window are sorted by timestamp to form a log event sequence. p is the total number of log entries. This refers to the log events in the sequence (usually log template IDs obtained after log parsing).
[0078] In addition, in order to overcome the shortcomings of superficial multimodal data fusion in the prior art (such as simple splicing), the present invention adopts a fusion framework based on attention mechanism. This framework is mainly used to ensure that data features of different modalities can interact dynamically and adaptively according to their inherent correlation.
[0079] Furthermore, a time series model is used to extract the dynamic behavior patterns and trends, i.e., time series features, from the indicator data;
[0080] It should be noted that, in order to effectively capture the mutual influence and time dependence between indicators, this invention adopts an architecture based on a gated recurrent unit (GRU).
[0081] Specifically, a multivariate time series matrix is input into a multilayer GRU network, which is configured to perform recursive computation along the time dimension. Furthermore, the hidden state at the last time step is regarded as a dynamic summary of the index data within the entire time window, and this hidden state is the extracted time series feature vector. ,in, This is the dimension of the time-series feature vector, which encapsulates the core dynamic information of the service node in terms of performance metrics.
[0082] Furthermore, a natural language processing model is used to extract deep semantic information, i.e. semantic features, from the preprocessed log event sequence;
[0083] It should be noted that, in order to deeply understand the specific terms and anomaly patterns contained in the operation and maintenance logs, the present invention adopts a Transformer model that has been pre-trained on a massive corpus of professional terms in the field of information technology operation and maintenance. This model is executed in parallel with the above-mentioned time-series feature extraction operation.
[0084] Specifically, each log template ID in the log event sequence is treated as a token, and the entire log event sequence is input into a pre-trained transformer model. The model's multi-head self-attention mechanism captures the contextual relationships between log events (i.e., which warning logs typically precede a given error log). Simultaneously, the model generates a context-dependent embedding vector for each log event in the sequence, and this embedding vector is used as input for subsequent fusion operations. This embedding vector is represented as:
[0085]
[0086] in, s represents the deep semantic information of the log event. Represented as an embedding vector, The dimension of the embedded vector;
[0087] Specifically, in one feasible implementation of the present invention, the transformer model is a BERT model. This model is pre-trained on a hybrid corpus containing public operation and maintenance log datasets (such as Spirit and Thunderbird) and internal historical operation and maintenance logs through a Masked Language Model (MLM) task, enabling the model to deeply understand the specific semantics in the operation and maintenance scenario.
[0088] It should be noted that the combination of all the embedding vectors constitutes the semantic feature sequence of the log data;
[0089] It should be noted that, in order to achieve the intelligent fusion of the above-mentioned temporal features and semantic features, the present invention adopts a cross-modal attention module. The core idea of this module is to use the macro-state (temporal features) extracted from the indicator data to "focus" on the part (semantic features) in the semantic feature sequence of the log data that is most relevant to the current state, and generate a "focused" semantic feature sequence representation of the log data based on this.
[0090] Specifically, the time-series feature vector is mapped through a linear transformation layer to generate a query vector. ,in, A learnable weight matrix is used to linearly transform the temporal feature vector to the query space. Then, the semantic feature sequence of the log data is passed through two different linear transformation layers to generate the key matrix. Sum matrix ,in, , Two independent, learnable weight matrices are used to linearly transform the semantic feature sequence of log data to the key space and value space, respectively. S is the semantic feature sequence of log data, and T represents the transpose operation. Attention weights are obtained by calculating the dot product similarity between the query vector and all key vectors, and then scaling and Softmax normalizing. The calculated attention weights are applied to the value matrix and weighted summation is performed to obtain a dynamically aggregated, context-aware log feature vector. Finally, the original temporal feature vector is concatenated with the attention-weighted semantic feature vector, and optionally dimensionality reduction and nonlinear fusion are performed through a fully connected layer to generate a node state vector representing the current state of the service node.
[0091] Specifically, the calculated attention weights are expressed by the formula:
[0092]
[0093] in, Represented as an attention weight vector, it indicates the degree of attention that time-series features pay to log events; The dimension of the key vector is used primarily to prevent the gradient from vanishing due to an excessively large dot product result.
[0094] Specifically, the log feature vector can be represented as:
[0095]
[0096] in, The log feature vector, i.e. the weighted semantic features, is a weighted sum of all log value vectors;
[0097] Specifically, the node state vector of the current state of the service node is represented as:
[0098]
[0099] in, This is represented as a node state vector. This is represented as the dimension of the node's state vector;
[0100] It should be noted that through the above fusion operation, the node state vector not only contains the node's own multimodal operation information, but also integrates the inherent logic between different data sources, providing a basic input for the spatiotemporal modeling of the subsequent graph neural network model;
[0101] S3. Input the structure of the dynamic topology graph and the node state vectors of all service nodes into a pre-trained graph neural network model.
[0102] It should be noted that the goal of this step is to use a deep learning model to perform end-to-end modeling of the spatiotemporal dynamics of the entire IT system. Furthermore, in order to overcome the limitations of traditional graph neural networks (GNNs) which can only handle static graph structures and recurrent neural networks (RNNs) which can only handle isolated time series, this invention preferably uses a pre-trained spatio-temporal graph neural network (STGNN) model to simultaneously learn the spatial dependencies on the graph structure (how faults propagate between services) and the dynamic evolution patterns in the time dimension (how service states change over time).
[0103] Furthermore, the core architecture of this spatiotemporal graph neural network consists of multiple spatiotemporal convolutional blocks stacked together. Within each spatiotemporal convolutional block, spatiotemporal features are jointly captured by alternately executing spatial graph convolution operations and temporal convolution operations.
[0104] Specifically, in one implementable process of the present invention, the structure of a spatiotemporal convolutional block can be defined as a three-layer structure of "temporal convolutional layer - spatial graph convolutional layer - temporal convolutional layer", and the entire spatiotemporal graph neural network model is composed of 6 such spatiotemporal convolutional blocks stacked together, and residual connections and layer normalization are used to ensure the stability and depth of the model during training.
[0105] Specifically, this spatial graph convolution operation is used to aggregate information of neighboring nodes in the spatial dimension. At each time step t, for any service node in the dynamic topology graph, the spatial graph convolution will weight and aggregate the state vectors of its neighboring nodes according to the connection relationship of the dynamic topology graph, and combine them with the node's own state vector to generate a new node representation that incorporates neighborhood context information. This operation can be implemented using mechanisms such as Graph Convolutional Network (GCN) or Graph Attention Network (GAT).
[0106] Specifically, in one feasible implementation of the present invention, a single-layer operation using GCN can be represented as follows:
[0107]
[0108] in, This is represented as the output node feature matrix. Each row of this matrix is a new feature vector of the service node after passing through a spatial graph convolution layer, incorporating neighborhood information. The dimension of this matrix is represented by the total number of service nodes in the dynamic topology graph multiplied by the dimension of the output features; similarly, The input node feature matrix is the matrix composed of the node state vectors of the current state mentioned above. Represented as a non-linear activation function, this function is used to introduce non-linearity, enabling the model to learn more complex patterns. Common activation functions include ReLU, LeakyReLU, Sigmoid, or Tanh. The activation function mentioned above is selected based on the computational efficiency of the current model. This is represented as an adjacency matrix with added self-loops. That is, an identity matrix I is added to the original adjacency matrix A. It is worth noting that since the original adjacency matrix only describes the connection relationship between nodes, adding the identity matrix I (i.e., adding 1 on the main diagonal) is equivalent to adding an edge (self-loop) pointing to itself for each service node. This is done so that when aggregating neighbor information, the original feature information of the service node itself can be preserved at the same time to prevent it from being diluted in multi-level transmission. It is represented as a degree matrix with added self-loops. This degree matrix is a diagonal matrix whose diagonal elements are the sum of all elements of the original adjacency matrix A in the current row (i.e., the degree of the service node plus the degree of its self-loop (1)). The learnable space weight matrix is the standard weight matrix in a neural network layer. Its parameters are learned during model training using the backpropagation algorithm. Its function is to map input features to output features.
[0109] It should be noted that in the single-layer operation of GCN mentioned above, taking the negative half power of the degree matrix is equivalent to taking the square root of the reciprocal of each diagonal element in the degree matrix. This operation is also known as the symmetric normalized Laplacian matrix. Its effect is to effectively balance the influence of the features of neighboring nodes and prevent high-degree service nodes (i.e., "center" service nodes with many connections) from occupying too much weight in feature propagation, thereby causing gradient explosion or vanishing, and ensuring the stability of the model during training.
[0110] Specifically, this temporal convolution operation is mainly used to capture the evolutionary pattern of each service node's state over time; after the spatial graph convolution operation, for each service node, its state over a past period of time is... Feature sequences after spatial aggregation It is fed into a temporal convolution module, where, This represents the current service node n within the time window. Similarly, the output node feature matrix is obtained from the following. This represents the current service node n within the time window. The output node feature matrix is then calculated, and so on. This is represented as the output node feature matrix of the current service node n at the previous time step t-1. This is represented as the output node feature matrix of the current service node n at the current time step t; As shown above, this represents a past period, i.e., the length of a historical time window, and satisfies... To ensure the effectiveness of this convolution operation, it is usually set to... ;
[0111] Specifically, the temporal convolution module can be implemented using a gated recurrent unit (GRU) or a temporal convolutional network (TCN). Regardless of whether a gated recurrent unit or a temporal convolutional network is used, the convolution module will learn the characteristics of the service node state changes over time and output a vector that summarizes historical time information to capture the inertia and trend of individual service performance and state.
[0112] Furthermore, by cascading the two spatial graph convolution operations and the temporal convolution operation in a spatiotemporal convolution block and stacking multiple such spatiotemporal convolution blocks, the spatiotemporal graph neural network model can learn deep, high-order spatiotemporal coupling dependencies.
[0113] Furthermore, since the graph neural network model used in the present invention is pre-trained, and the pre-training process is based on a meta-learning framework, the goal of this pre-training is to solve the cold start problem of new services in operation and maintenance scenarios. That is, when a new service node is added to the system for the first time, there is a lack of sufficient historical data to train a dedicated anomaly detection model. Therefore, a meta-learning framework is needed to enable the model to quickly adapt to new tasks.
[0114] Specifically, the pre-training process for this model is as follows:
[0115] From historical operation and maintenance data, a large number of service nodes of different types that have been running stably for a long time are selected as the meta-training set. "Learning the normal behavior pattern of a specific service node" is defined as a meta-task. Then the meta-training set consists of hundreds or thousands of such meta-tasks.
[0116] Furthermore, this invention employs the Model-Agnostic Meta-Learning (MAML) algorithm as the framework for meta-learning. The goal of this algorithm is not to learn a set of model parameters that perform best on any meta-task, but to learn a special set of meta-model parameters. That is, starting from these meta-model parameters, only one or a few gradient descent updates are needed on a small number of samples of a new meta-task to quickly converge to the optimal parameters of the new task.
[0117] First, by randomly sampling a batch of tasks from the meta-task set in the inner loop iteration of each meta-training, and for each task, copying the model parameters and training on a small amount of support set data for that task, a set of task-specific temporary parameters is obtained.
[0118] Then, in each outer loop iteration of the meta-training, the loss is evaluated on the query set of the task using this set of temporary parameters. Based on the evaluated loss, the gradient of the temporary parameters with respect to the meta-model parameters is calculated, and the meta-model parameters are updated.
[0119] Finally, by repeating the above inner and outer loop process on a large number of different tasks, the meta-model parameters can gradually learn a "universal" initialization state, which contains the common knowledge of the normal behavior patterns of various service nodes.
[0120] Specifically, in one implementable process of the present invention, the meta-task is constructed as follows:
[0121] For each service node, a task is to sample data for 24 consecutive hours from historical data. In each task, the data from the first 12 hours is used as the support set and the data from the last 12 hours is used as the query set. The learning rate of the inner loop of the agnostic meta-learning algorithm is set to 0.01 and the learning rate of the outer loop is set to 0.001. The gradient updates of the inner and outer loops use the default step size, which is 1.
[0122] Furthermore, through this meta-learning framework, when a new service node joins the system, there is no need to train the model from scratch. The pre-trained meta-model parameters can be used directly as the initial weights, and then fine-tuned using some initial data of the new service node (e.g., the operation and maintenance data of the first few hours after going live).
[0123] It should be noted that, due to the inherent advantages of meta-parameters, the model can converge quickly with only a small amount of data and very few training steps, thereby rapidly building a high-precision anomaly detection capability for new services.
[0124] Furthermore, to adapt the outputs of steps S1 and S2 to the pre-trained spatiotemporal graph neural network model, the present invention organizes the output results into a temporal graph sequence, and the input to the model is represented as a four-dimensional tensor with dimensions of . Where B represents the batch size. The input is represented by the length of the historical time step, N is represented by the total number of service nodes in the system, and D is represented by the dimension of the state vector of each node.
[0125] It should be noted that, at each time step t, the structure of the dynamic topology graph (i.e., the original adjacency matrix) is also used as input to guide the neural network model to perform correct spatial information aggregation.
[0126] S4. Calculate the anomaly score for each service node using a graph neural network model;
[0127] It should be noted that, since the pre-trained spatiotemporal graph neural network model has already received information from the system in the past... The task of this step is to use the historical state data of each time step to calculate a quantified anomaly score for each service node in the current time step using the model. This process is based on the unsupervised learning paradigm, which identifies anomalies by comparing the deviation between the model's prediction results and the actual observations. The core idea of the unsupervised learning paradigm adopted in this invention is that a model trained on massive amounts of normal operating data can accurately learn the spatiotemporal evolution law of the system under normal operating conditions. Therefore, the model has a strong predictive ability for the normal state in the near future. Conversely, when a service node or its associated neighboring nodes exhibit abnormal behavior, its state evolution will deviate from the learned normal pattern, resulting in a significant deviation between the model's prediction results and the actual observations. The magnitude of this deviation, i.e., the reconstruction error (specifically the prediction error in this invention), directly reflects the degree of anomaly of the node.
[0128] Specifically, since the pre-trained spatiotemporal graph neural network model in this invention acts as a time series graph predictor, its task is to determine the input time window. Based on the historical node state vector sequence and dynamic topology graph structure up to time step t, the system predicts the node state vectors of all service nodes in the system at the next time step t+1. Therefore, the model's prediction output can be denoted as... The actual node state vector generated at time t+1 through step S2 will be denoted as... ;
[0129] Furthermore, an anomaly score is calculated individually for each service node, which is achieved by quantifying the difference between the predicted vector and the true vector;
[0130] It should be noted that, in order to effectively measure the deviation between these two high-dimensional vectors, this invention employs the L2 norm, which calculates the deviation between the two vectors in terms of their relative dimensions. The linear distance in the dimensional space is used to comprehensively reflect the differences in all feature dimensions, and this norm is more sensitive to larger error values, making it very suitable for capturing significant abnormal deviations;
[0131] Specifically, for any service node, the formula for calculating its anomaly score at time point t+1 is as follows:
[0132]
[0133] in, Indicates service node The anomaly score at time t+1 is a non-negative scalar value. The larger the scalar value, the greater the deviation between the actual operating state of the service node at time t+1 and the state predicted based on the historical normal pattern, that is, the higher the probability of the node being abnormal. It is represented as the L2 norm, i.e., the Euclidean distance; and These are respectively represented as the predicted node state vector and the actual node state vector of the current service node n at time step t+1;
[0134] Furthermore, by repeatedly performing the above anomaly score calculation steps on all N service nodes in the system, an anomaly score vector will be output at each time step t+1. Each element in the vector precisely quantifies the degree of anomaly of the corresponding service node at the current moment.
[0135] S5. When the abnormal score of any service node exceeds the preset alarm threshold, an operation and maintenance alarm is generated.
[0136] It should be noted that this step aims to establish a dynamic and adaptive alarm threshold determination mechanism to overcome the limitations of traditional fixed threshold methods in complex and ever-changing IT system environments.
[0137] Specifically, at each evaluation point in time, the abnormal scores of all N service nodes in the system are monitored in real time. For each service node, if the abnormal score of that service node... If the anomaly score at time t+1 is greater than the current alarm threshold, the system will trigger an alarm generation program for the corresponding service node.
[0138] Specifically, the alarm threshold is the critical value at which an anomaly score is determined to be "significantly abnormal" under the current system state, and this value changes over time as the current system state changes.
[0139] Furthermore, static thresholds often exhibit poor robustness in the face of system load fluctuations, version updates, or changes in business models, easily leading to alarm storms (thresholds too low) or missed critical faults (thresholds too high). To overcome this limitation, this invention introduces a Peaks-Over-Threshold (POT) method based on Extreme Value Theory (EVT) to dynamically calculate alarm thresholds. This method is statistically capable of accurately modeling the "tail" of the distribution, i.e., rare but significant extremely high anomaly scores.
[0140] Specifically, the steps for determining the dynamic alarm threshold are as follows:
[0141] The system collects a series of abnormal score data generated by all service nodes within a sliding historical time window. Then, it calculates a higher quantile point u from the score data and uses this quantile point u as a basic threshold to filter out "potential" extreme abnormal scores. It only focuses on those scores that exceed u, because they are most likely to represent real system anomalies.
[0142] According to the Pickands-Balkema-de Haan theorem (PBDH theorem) in extreme value theory, for a sufficiently high base threshold u, the conditional probability distribution of the portion of the random variable that exceeds u (i.e., y = Score - u) can be approximated by the generalized Pareto distribution (GPD).
[0143] Specifically, the cumulative distribution function of the generalized Pareto distribution Represented as:
[0144]
[0145] Where y represents the excess value, i.e., y = Score - u, which is the abnormal score value that exceeds the basic threshold u; Represented as a shape parameter, it is used to control the "thickness" of the tail of the conditional probability distribution. This is because the operational anomaly score usually belongs to a heavy-tailed distribution; Represented as a scale parameter, and This is used to control the extent to which the conditional probability distribution is expanded;
[0146] It should be noted that by using all historical excess values as input and employing statistical methods such as Maximum Likelihood Estimation (MLE), the optimal generalized Pareto distribution can be fitted. and ;
[0147] After fitting the optimal parameters of the generalized Pareto distribution, the abnormal score can be calculated under a very small probability (i.e., the preset alarm rate q), and this value is the final dynamic alarm threshold.
[0148] However, it needs to be explained that this preset alarm value is a business-related hyperparameter, and its initial value can be q=0.001, which means that we expect that only one in a thousand cases will trigger an alarm in all time steps;
[0149] Specifically, the formula for calculating the dynamic alarm threshold is derived from the quantile function of the generalized Pareto distribution (i.e., the inverse function of the generalized Pareto distribution):
[0150]
[0151] in, For dynamic alarm thresholds, and These are the fitted scale parameters and shape parameters, respectively. This represents the total number of historical outlier score data used to select u. This represents the number of data points in the historical anomaly score data that exceed the basic threshold u.
[0152] Furthermore, once the abnormal score of a service node exceeds the dynamic alarm threshold, the system will immediately generate a structured alarm message, which will include at least the following key fields: alarm timestamp, service node name, abnormal score, current alarm threshold, alarm level (which can be set according to the degree to which the score exceeds the threshold (e.g., severe, warning)) and associated context (i.e., attaching the key indicator value and log summary at the time the alarm was triggered).
[0153] Furthermore, the generated alarms will then be pushed to an integrated alarm management platform (such as Prometheus, PagerDuty, etc.) to notify operations and maintenance personnel to intervene. At the same time, the service node information and anomaly score that triggered the alarm will be provided.
[0154] Furthermore, after generating an operational alarm, to help operations personnel quickly and accurately understand the root cause and propagation path of the fault, rather than simply dealing with the surface alarm phenomenon, the present invention also includes a root cause localization step. This step aims to accurately identify the initial source of the fault from multiple service nodes that may trigger alarms simultaneously, and clearly depict the propagation trajectory of the fault in the system topology. This step is as follows:
[0155] When multiple service nodes have abnormal scores that exceed the dynamic alarm threshold successively or simultaneously within a short period of time (e.g., within an alarm aggregation window), the system will activate the root cause node identification logic, which follows two core principles: time priority and anomaly severity.
[0156] Specifically, time priority means that faults in distributed systems usually propagate in a cascading manner. Therefore, the service node that first becomes abnormal is most likely to be the source of the fault. The system will record the timestamp of each service node when it first exceeds the alarm threshold.
[0157] Specifically, the severity of the anomaly refers to the abnormal behavior of the root cause node. This abnormal behavior is usually the most severe, so its state deviates from the normal pattern to the greatest extent. Therefore, among the service nodes that trigger the alarm earliest, the service node with the highest anomaly score is most likely to be the root cause node.
[0158] It should be noted that, combining the above two points, the rule for determining the root cause node can be expressed as follows: among the set of multiple service nodes that trigger alarms, the node that exceeds the alarm threshold earliest in time and has the highest abnormal score near that time point is determined as the root cause node.
[0159] Furthermore, a subset of nodes with the earliest timestamps is selected, and the node with the highest anomaly score in this subset is chosen as the root cause.
[0160] Specifically, select the subset of nodes with the earliest timestamps. Represented as:
[0161]
[0162] in, This represents a set of multiple service nodes that triggered the alarm. This indicates the timestamp when the service node first triggered the alarm. As a minimum value operation, the expression will iterate through... Each service node in the set (here, denoted as ) (used as temporary variables during iteration) to obtain their respective alarm timestamps. And find the earliest (smallest) timestamp among them, for example, if There are three service alarms with timestamps of 10:30:10, 10:30:07, and 10:30:05. The result of this expression is 10:30:05.
[0163] Specifically, the node with the highest anomaly score in this subset is selected as the root cause, represented as follows:
[0164]
[0165] in, Represented as root cause node, Indicates only in Find the node that maximizes the objective function value from the set (i.e., the set of nodes that first issued an alert). This represents the anomaly score of the service node at the time of its alarm.
[0166] Furthermore, once the root cause node is identified, it is necessary to further analyze the causal relationship between nodes to determine how the fault propagated from the root cause node to other affected nodes. This process utilizes the differentiability of the pre-trained spatiotemporal graph neural network model to quantify the state influence between nodes by calculating gradients.
[0167] Specifically, for a confirmed abnormal service node (downstream node), the system traces all its upstream neighbor nodes. By calculating the gradient of the downstream node's anomaly score relative to the state vector of each of its upstream neighbors at the previous time step, the system measures the impact of the upstream neighbors' state changes on the downstream node's anomaly. The more influential the upstream node, the more likely it is to be the direct cause of the fault propagating to the current node. The formula for calculating this influence score is as follows:
[0168]
[0169] in, This represents the influence score of an upstream node on a downstream node, and is a non-negative scalar. The abnormal score of the downstream node at time t+1 is calculated similarly to the aforementioned abnormal score calculation. It is the node state vector of the upstream node at time t; The partial derivative operator is represented by . Since the spatiotemporal graph neural network model is end-to-end differentiable, the gradient of this formula can be efficiently calculated using the backpropagation algorithm in an automatic differentiation framework (such as PyTorch or TensorFlow).
[0170] It should be noted that by calculating the influence scores of all its upstream neighbors for each downstream node, a weighted directed graph can be constructed to represent the probabilistic path of fault propagation; and by starting from the root cause node and traversing along the path with the highest influence score, the complete propagation chain from the source of the fault to all affected nodes can be clearly outlined.
[0171] It should be noted that, in order to ensure that the solution of this invention maintains high accuracy and robustness in dynamically changing IT system environments (e.g., business logic updates, service architecture adjustments, seasonal changes in traffic patterns, etc.), the solution of this invention also includes an online model adaptation step. This step mainly enables the model to autonomously detect environmental changes and perform incremental learning, thereby avoiding performance degradation caused by "model drift". This step specifically includes:
[0172] Because the distributed system does not view each anomaly score in isolation, but continuously monitors the overall distribution of all anomaly scores generated by all service nodes within the most recent sliding time window (e.g., the past 24 hours) from a macro perspective, a healthy and stable system will typically exhibit a stable, long-tailed distribution of anomaly scores. To quantify changes in the overall distribution, the system maintains a reference distribution, which is the empirical probability distribution of anomaly scores collected within a stable operating cycle after the model's last update or initial deployment. At the same time, it calculates the anomaly score distribution within the current time window in real time.
[0173] When the system state undergoes a fundamental change (for example, a version release introduces a new normal behavior pattern), the old model may misjudge this new pattern as abnormal, resulting in a systematic shift in the mean, variance, or shape of the overall abnormal score distribution. The present invention automatically detects this change by using a statistical method to measure distribution differences.
[0174] Specifically, in one feasible implementation of the present invention, the Kullback-Leibler (KL) divergence is used to measure the difference between the current anomaly score distribution P and the reference anomaly score distribution Q. This KL divergence represents the information loss caused when the reference distribution is used to approximate the current distribution, and its calculation formula is as follows:
[0175]
[0176] in, This represents the KL divergence from the reference distribution to the current distribution. It is the probability that the anomaly score is x within the current time window; It is the probability of an anomaly score of x within the reference time window; Given the set of all possible values for the outlier score, in practice, P and Q can be constructed under discrete probability distributions by binning the scores;
[0177] It should be noted that the system will preset a significant change threshold. When the calculated KL divergence is greater than this significant change threshold, it is determined that the overall abnormal score distribution of the system has undergone a preset significant change, indicating that the current model's cognition has deviated significantly from the actual operating state of the system, and a model update operation needs to be triggered.
[0178] Specifically, once a significant change is detected, the system will automatically trigger incremental online learning of the spatiotemporal graph neural network model. This process aims to adapt the model to the "new normal" rather than a complete retraining from scratch, in order to save computing resources and achieve rapid adaptation.
[0179] Specifically, the system collects all operational data (including dynamic topology, node state vectors, etc.) within the time window that caused the distribution change. This recent operational data is assumed to represent the system's new normal behavior pattern. Subsequently, this recent operational data is used to perform several additional training rounds on the pre-trained and fine-tuned graph neural network model. During the training process, a small learning rate is used (for example, the learning rate is set to 1e-5, which is about one-tenth of the initial fine-tuning learning rate) to ensure that the model does not completely forget the old knowledge learned from historical data while learning new patterns (i.e., to mitigate the "catastrophic forgetting" problem). At the same time, after completing incremental learning, the system will continue to generate alarms using the updated model and recalculate and set the reference anomaly score distribution using the data from the latest stable period to prepare for the next round of adaptive monitoring.
[0180] It should be noted that, through the "monitoring-detection-learning" mechanism, the present invention achieves continuous self-evolution of the model and a high degree of adaptability to dynamic environments.
[0181] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0182] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0183] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0184] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0185] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0186] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for generating intelligent operation and maintenance alarms based on unsupervised learning, characterized in that, include: Based on operational data, a dynamic topology graph representing each service node in the system and their interrelationships is constructed. For each service node in the dynamic topology graph, generate a node state vector representing its current state; The structure of the dynamic topology graph and the node state vectors of all service nodes are input into a pre-trained graph neural network model. The anomaly score for each service node is calculated using the graph neural network model. An operation and maintenance alarm is generated when the abnormal score of any service node exceeds the preset alarm threshold.
2. The intelligent operation and maintenance alarm generation method based on unsupervised learning as described in claim 1, characterized in that, The operation and maintenance data includes at least indicator data, log data, and link tracing data, and the dynamic topology map is constructed based on the link tracing data.
3. The intelligent operation and maintenance alarm generation method based on unsupervised learning as described in claim 1 or 2, characterized in that, For each service node in the dynamic topology graph, generate a node state vector representing its current state, including: The node status vector is generated by integrating the metric data and log data corresponding to each service node.
4. The intelligent operation and maintenance alarm generation method based on unsupervised learning as described in claim 3, characterized in that, The specific methods of fusion include: Based on a time series model, time-series features are extracted from the indicator data; Based on a natural language processing model, semantic features are extracted from the log data; A cross-modal attention module is applied, which uses the temporal features as queries and the semantic features as keys and values, calculates attention weights and weights the semantic features, and combines the weighted semantic features with the temporal features to generate the node state vector.
5. The intelligent operation and maintenance alarm generation method based on unsupervised learning as described in claim 1, characterized in that, The graph neural network model is a spatiotemporal graph neural network model, which processes the input node state vector by alternately executing spatial graph convolution operations and temporal convolution operations.
6. The intelligent operation and maintenance alarm generation method based on unsupervised learning as described in claim 1, characterized in that, The graph neural network model is pre-trained using a meta-learning framework to obtain a set of meta-model parameters. When a new service node joins the system, the meta-model parameters are fine-tuned using some of the initial data from the new service node.
7. The intelligent operation and maintenance alarm generation method based on unsupervised learning as described in claim 1, characterized in that, The anomaly score is calculated based on the reconstruction error between the graph neural network model's prediction of the current node's state vector and the actual node's state vector.
8. The intelligent operation and maintenance alarm generation method based on unsupervised learning as described in claim 1, characterized in that, The alarm threshold is dynamically determined, and the steps for this dynamic determination include: Obtain historical outlier score data and select a high quantile point; The generalized Pareto distribution model was used to fit the outlier scores that exceeded the high quantile points; Based on the fitted generalized Pareto distribution model and the preset alarm rate, the current alarm threshold is calculated.
9. The intelligent operation and maintenance alarm generation method based on unsupervised learning as described in claim 1, characterized in that, After generating the operation and maintenance alarm, a root cause localization step is also included, which includes: Among the multiple service nodes that trigger the alarm, the node that exceeds the alarm threshold earliest in time and has the highest anomaly score will be identified as the root cause node. The fault propagation path is determined by calculating the gradient of the anomaly score of an abnormal service node relative to the state vector of its upstream neighbor nodes.
10. The intelligent operation and maintenance alarm generation method based on unsupervised learning as described in claim 1, characterized in that, The method further includes an online model adaptation step, which includes: Continuously monitor the overall anomaly score distribution of the system; When a predetermined significant change in the score distribution is detected, the graph neural network model is incrementally learned online using recent data.
Citation Information
Patent Citations
Operation and maintenance alarm prediction method and device, equipment, storage medium and product
CN120596341A
Network traffic abnormity monitoring method and device based on BiLSTM-Att network
CN119232490A
Task chain self-healing transmission prediction method and system based on parameter fingerprint holographic perception
CN119623775A
Multi-layer coupling power system risk propagation analysis method based on space-time diagram neural network
CN119903394A
Network abnormal state identification method and device
CN120547219A
Cited By
Server cluster operation and maintenance method based on multi-source heterogeneous data fusion and dynamic knowledge graph
CN121705077A
Server cluster operation and maintenance method based on multi-source heterogeneous data fusion and dynamic knowledge graph
CN121705077B
Micro-service fault root cause positioning method, system and equipment
CN122332173A