An alarm root cause positioning and repairing method, device and equipment of a cloud platform
By constructing a knowledge graph and a multi-dimensional scoring mechanism, the problem of information silos in alarm analysis in cloud computing environments has been solved, enabling automated root cause localization and self-healing in intelligent operation and maintenance, and improving fault response speed and repair efficiency.
Patent Information
- Application Number
- CN202511205046.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing technologies struggle to intelligently analyze alarm context, automatically locate root causes, and achieve self-healing in cloud computing environments, resulting in low operational efficiency, high false alarm rates, reliance on human experience, and an inability to cope with dynamically changing fault modes.
By constructing a knowledge graph, based on cloud platform operation and maintenance data, multimodal data is collected and contextual information is generated. Noise reduction processing is performed, a multi-dimensional scoring mechanism is used to screen root cause nodes, and a repair plan is generated for automated repair.
It enables intelligent analysis of alarm context, automated root cause location, reduction of invalid alarm information, improved fault response efficiency and repair timeliness, reduced risk of manual intervention, and increased fault recovery speed.
Smart Images

Figure CN120750739B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of monitoring and maintenance technology, and in particular to a method, apparatus and equipment for locating and repairing alarm root causes in cloud platforms. Background Technology
[0002] With the widespread adoption of cloud computing, microservice architecture, and containerization technologies, the complexity and dynamism of modern IT (Information Technology) systems are growing exponentially. Traditional operation and maintenance methods, such as static threshold-based alerts and manual rule correlation, are no longer sufficient to address the following challenges: single-point failures often trigger cascading effects, generating massive amounts of redundant alerts that overwhelm critical signals, leading to delays in the operation and maintenance team's response; alert information is isolated and lacks contextual correlation with infrastructure, application topology, and business dependencies, making root cause localization reliant on expert experience; traditional operation and maintenance models lack predictive analytics and automated self-healing capabilities, resulting in prolonged business downtime.
[0003] There is an urgent need to shift from traditional operations and maintenance (O&M) to intelligent operations and maintenance (AIOps). Several mainstream solutions exist, but they still have limitations: (1) Rule-based alarm correlation system: It relies on pre-written rules, which are logically clear but rigid and have high maintenance costs. It cannot adapt to the dynamic changes in the cloud environment and cannot handle unforeseen failure modes; (2) Anomaly detection system based on statistics and machine learning: It can detect anomalies by analyzing the time series of indicators and can discover unknown problems, but it lacks an understanding of the system topology and dependencies, has weak root cause analysis capabilities, and has a high false alarm rate; (3) Correlation analysis based on static CMDB (Configuration Management Database) / topology: It relies on static and untimely updated configuration information, which only reflects the "ought" relationship at the time of design and cannot capture the "actual" interaction dependencies of dynamic changes between services at runtime; (4) Separate operation and maintenance data platform: It stores indicators, logs, traces and other data in isolation. Operation and maintenance personnel need to manually jump between data silos. The analysis process is highly dependent on manual labor and has low efficiency.
[0004] In summary, how to provide an intelligent operation and maintenance method that can intelligently analyze alarm context, automatically locate root causes, and achieve self-healing is a problem that needs to be solved. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a method, apparatus, and device for alarm root cause localization and repair in a cloud platform, capable of intelligent operation and maintenance that intelligently analyzes alarm context, automatically locates root causes, and achieves closed-loop self-healing. The specific solution is as follows:
[0006] Firstly, this application discloses a method for locating and resolving alarm root causes on a cloud platform, including:
[0007] After receiving the original alarm information from the cloud platform, the corresponding alarm node is located from the preset knowledge graph, and the context information of each alarm node is generated based on the knowledge graph; the knowledge graph is built based on the operation and maintenance data of the cloud platform.
[0008] The original alarm information is denoised based on context information to eliminate invalid alarm information, and the denoised target alarm information is obtained.
[0009] For each target alarm message, a corresponding candidate subgraph is determined from the knowledge graph, and each node in the candidate subgraph is scored from at least two dimensions. The target nodes selected in descending order of total score are used as root cause nodes. The candidate subgraph includes the target alarm node corresponding to the target alarm message and the adjacent nodes of the target alarm node.
[0010] Generate a root cause analysis report for the root cause node and obtain a remediation plan that matches the root cause analysis report, so as to remediate the root cause node based on the remediation plan.
[0011] Optional, the knowledge graph construction process includes:
[0012] Based on preset data acquisition adapters and preset acquisition probes, multimodal operation and maintenance data is continuously collected from the cloud platform's operating environment; among which, multimodal operation and maintenance data includes infrastructure configuration data, container event data, application business log data, and kernel-level data;
[0013] Entity objects and entity relationships are identified from multimodal operation and maintenance data, and a knowledge graph is constructed with entity objects as nodes and entity relationships as edges.
[0014] Optionally, entity objects include physical entities, logical entities, and business entities; entity relationships include explicit dependencies obtained from target configuration information and implicit dependencies obtained based on data analysis methods.
[0015] Optionally, contextual information for each alarm node is generated based on the knowledge graph, including:
[0016] Using each alarm node as the central node, a graph traversal operation is performed in the knowledge graph to obtain target information that satisfies a preset relationship with the central node;
[0017] The target information is assembled to generate context information for each alarm node.
[0018] Optionally, a graph traversal operation is performed in the knowledge graph to obtain target information that satisfies a preset relationship with the central node, including:
[0019] Perform graph traversal operations in the knowledge graph to obtain underlying resource information that satisfies the preset vertical dependency relationship with the central node;
[0020] Perform graph traversal operations in the knowledge graph to obtain service information that satisfies a preset horizontal dependency relationship with the central node; the service information includes the first service information that the central node depends on and the second service information that depends on the central node.
[0021] Perform graph traversal operations in the knowledge graph to obtain information on the changes of the central node and its adjacent nodes within the alarm time window.
[0022] Optionally, noise reduction processing is performed on each original alarm message based on context information to eliminate invalid alarm messages, including:
[0023] If, based on context information, it is determined that there is a parent-child node relationship between the alarm nodes corresponding to each original alarm message, then the original alarm message corresponding to the child node is suppressed.
[0024] Aggregate duplicate and similar alarm messages from each original alarm message;
[0025] If any of the original alarm messages contain scenario alarm messages under a preset standard business scenario, then the scenario alarm messages will be suppressed.
[0026] Optionally, duplicate and similar alarm messages in each original alarm message can be aggregated, including:
[0027] Determine the alarm source of each original alarm message. If the same alarm source generates duplicate alarm messages within a preset time window, merge the duplicate alarm messages into one alarm message.
[0028] The similarity of each original alarm message is calculated using a preset text similarity algorithm, and similar alarm messages with a similarity greater than a preset threshold are obtained to aggregate the similar alarm messages.
[0029] Optionally, if any of the original alarm messages contain scenario alarm messages under a preset business scenario, then the scenario alarm messages are suppressed, including:
[0030] If any of the original alarm messages contain instantaneous alarm messages generated under the preset system self-healing scenario, then the instantaneous alarm messages will be suppressed.
[0031] If any of the original alarm messages contain expected alarm messages within a preset range of influence, then the expected alarm messages will be suppressed.
[0032] If any of the original alarm messages contain an experimental alarm message corresponding to the preset chaos engineering experiment, then the experimental alarm message will be suppressed.
[0033] Optionally, for each target alarm message, a corresponding candidate subgraph is determined from the knowledge graph, including:
[0034] Identify the target alarm node corresponding to each target alarm information from the knowledge graph;
[0035] Starting with the target alarm node, perform a search operation in the knowledge graph to obtain adjacent nodes within a preset number of hops;
[0036] The range formed by the target alarm node and its adjacent nodes in the knowledge graph is used as a candidate subgraph.
[0037] Optionally, each node in the candidate subgraph is scored from at least two dimensions to select the target nodes as root cause nodes in descending order of total score, including:
[0038] Each node in the candidate subgraph is scored from at least two dimensions to obtain a score value for each dimension;
[0039] The score values corresponding to each dimension are weighted and summed to obtain the total score value for each node;
[0040] Select a predetermined number of target nodes based on the total score from highest to lowest, and use these target nodes as root cause nodes.
[0041] Optionally, each node in the candidate subgraph is scored from at least two dimensions to obtain a score value for each dimension, including:
[0042] The prior probability of failure for each node is calculated based on historical data, and the first score value is determined based on the prior probability of failure.
[0043] Acquire the time-series data of each node within the alarm time window, analyze the time-series data to determine the target time point when the node performance indicators become abnormal, and then determine the second score value based on the order between the target time point and the alarm time point corresponding to the target alarm node.
[0044] Assess whether each node is associated with a target historical risk change event before the alarm time point, and determine the third score based on the association; wherein, the target historical risk change event is a change operation that occurs before the alarm time point and belongs to a preset risk type, and the association includes the time interval between the change time point of the target historical risk change event and the alarm time point being less than a preset interval threshold.
[0045] Perform anomaly analysis on the log data corresponding to each node, and determine the fourth score value based on the anomaly analysis results.
[0046] Optionally, a root cause analysis report can be generated for the root cause nodes, including:
[0047] Obtain scoring evidence for each dimension to determine the corresponding score value;
[0048] Each scoring evidence and root cause node is input into the first pre-trained model to output a root cause analysis report.
[0049] Optionally, obtain remediation plans that match the root cause analysis report, including:
[0050] Identify the pre-established contingency plan knowledge base and check whether there are remediation plans in the contingency plan knowledge base that match the root cause analysis report;
[0051] If it exists, the remediation plan will be retrieved directly from the contingency plan knowledge base;
[0052] If not, the root cause analysis report and remediation goals are input into the second pre-trained model to generate a remediation plan based on the training data.
[0053] Optionally, after repairing the root cause node based on the repair plan, the following may also be included:
[0054] Once the root cause node repair is detected to be complete, if the repair plan is generated by the second pre-trained model, the repair plan will be stored in the plan knowledge base and used as training data for the second pre-trained model.
[0055] Optionally, the process of repairing root cause nodes based on the repair plan also includes:
[0056] Continuously monitor the business status of root cause nodes;
[0057] Once the root cause node is determined to have returned to normal based on the business status, the repair operation is stopped, and the root cause node repair is considered complete.
[0058] Optionally, the root cause node can be repaired based on the remediation plan, including:
[0059] The remediation plan is tested in a pre-set sandbox environment. After the pre-test is passed, the execution mode is determined according to the risk level of the remediation plan. The execution modes include fully automatic execution mode, semi-automatic execution mode and fully manual execution mode.
[0060] Based on the execution mode, the pre-set automated tools are invoked to execute the repair plan in order to repair the root cause node.
[0061] Secondly, this application discloses a cloud platform alarm root cause localization and repair device, comprising:
[0062] The context information generation module is used to locate the corresponding alarm node from the preset knowledge graph after receiving the original alarm information from the cloud platform, and generate the context information of each alarm node based on the knowledge graph; wherein, the knowledge graph is built based on the operation and maintenance data of the cloud platform;
[0063] The noise reduction module is used to perform noise reduction processing on each original alarm information based on context information to eliminate invalid alarm information and obtain the target alarm information after noise reduction.
[0064] The root cause node determination module is used to determine the corresponding candidate subgraph from the knowledge graph for each target alarm information, and to score each node in the candidate subgraph from at least two dimensions, so as to select the target nodes in descending order of total score as root cause nodes; wherein, the candidate subgraph includes the target alarm node corresponding to the target alarm information and the adjacent nodes of the target alarm node;
[0065] The repair module is used to generate root cause analysis reports for root cause nodes and obtain repair plans that match the root cause analysis reports, so as to repair the root cause nodes based on the repair plans.
[0066] Thirdly, this application discloses an electronic device, including:
[0067] Memory, used to store computer programs;
[0068] A processor for executing computer programs to implement the steps of the aforementioned disclosed cloud platform alarm root cause localization and remediation methods.
[0069] Fourthly, this application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the steps of the aforementioned cloud platform alarm root cause localization and repair method.
[0070] Fifthly, this application discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the aforementioned cloud platform alarm root cause localization and repair method.
[0071] As can be seen, after receiving the original alarm information from the cloud platform, this application locates the corresponding alarm nodes from a preset knowledge graph and generates context information for each alarm node based on the knowledge graph. The knowledge graph is built based on the cloud platform's operation and maintenance data. Noise reduction processing is performed on each original alarm information based on the context information to eliminate invalid alarm information, resulting in denoised target alarm information. For each target alarm information, a corresponding candidate subgraph is determined from the knowledge graph, and each node in the candidate subgraph is scored from at least two dimensions. Target nodes selected in descending order of total score are then used as root cause nodes. The candidate subgraph includes the target alarm node corresponding to the target alarm information and its adjacent nodes. A root cause analysis report is generated for the root cause node, and a remediation plan matching the root cause analysis report is obtained. The root cause node is then remediated based on the remediation plan.
[0072] Beneficial Effects: The knowledge graph in this application is built based on cloud platform operation and maintenance data. By locating alarm nodes and generating contextual information through the knowledge graph, isolated raw alarm information is bound to deeply related data such as system topology and dependencies. This avoids the information silo problem in traditional alarm analysis and provides comprehensive correlation evidence for subsequent noise reduction and root cause localization, improving the accuracy of analysis results from a fundamental level. Furthermore, this application performs noise reduction based on contextual information, effectively identifying and eliminating invalid alarm information, reducing the total number of alarms faced by operation and maintenance personnel, eliminating the need to sift through massive amounts of redundant information to select key content, and significantly improving fault response efficiency. For each target alarm information, this application limits the scope of root cause investigation to the target alarm node and its adjacent nodes through candidate subgraphs, and uses a multi-dimensional scoring mechanism to screen root cause nodes. This multi-dimensional scoring method reduces the one-sidedness of single-dimensional judgment, making root cause localization more accurate, while reducing reliance on the experience of senior operation and maintenance personnel. Finally, this application generates a root cause analysis report and matches it with a remediation plan, directly transforming the root cause location results into executable remediation actions. This completes the repair of the root cause node, improving the timeliness and automation of fault repair, and reducing the lag and operational risks of manual intervention. Through the above solution, intelligent operation and maintenance is achieved, enabling intelligent analysis of alarm context, automated root cause location, and closed-loop self-healing, thereby improving fault response speed and reducing fault recovery time. Attached Figure Description
[0073] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0074] Figure 1 This application discloses a flowchart of a cloud platform alarm root cause localization and repair method.
[0075] Figure 2 This application discloses a flowchart of a specific cloud platform alarm root cause localization and repair method.
[0076] Figure 3 This is a diagram illustrating the overall architecture of a cloud platform alarm root cause localization and repair method disclosed in this application.
[0077] Figure 4 This is a schematic diagram of the structure of a cloud platform alarm root cause localization and repair device disclosed in this application;
[0078] Figure 5 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0079] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0080] There are several mainstream solutions, but they still have limitations: (1) Alarm association system based on rules: It relies on pre-written rules, which are logically clear but rigid and have high maintenance costs. It cannot adapt to the dynamic changes in the cloud environment and cannot handle unforeseen fault modes; (2) Anomaly detection system based on statistics and machine learning: It can detect anomalies by analyzing the time series of indicators and can discover unknown problems, but it lacks an understanding of the system topology and dependency relationship, has weak root cause analysis capabilities, and has a high false alarm rate; (3) Association analysis based on static CMDB (Configuration Management Database) / topology: It relies on static and untimely updated configuration information, which only reflects the "ought" relationship at the time of design and cannot capture the "actual" interaction dependency of dynamic changes between services at runtime; (4) Separate operation and maintenance data platform: It stores indicators, logs, tracking and other data in isolation. Operation and maintenance personnel need to manually jump between data silos. The analysis process is highly dependent on manual labor and has low efficiency.
[0081] Therefore, this application discloses a method, apparatus, and device for alarm root cause localization and repair on a cloud platform, which can realize intelligent operation and maintenance by intelligently analyzing alarm context, automatically locating root causes, and achieving closed-loop self-healing.
[0082] See Figure 1As shown in the embodiment of this application, a method for locating and repairing alarm root causes on a cloud platform is disclosed. The method includes:
[0083] Step S11: After receiving the original alarm information from the cloud platform, locate the corresponding alarm node from the preset knowledge graph, and generate the context information of each alarm node based on the knowledge graph; wherein, the knowledge graph is built based on the operation and maintenance data of the cloud platform.
[0084] In this embodiment, after receiving each original alarm message from the cloud platform, the original alarm message needs to be parsed in order to locate the corresponding alarm node from the preset knowledge graph, and generate the context information of each alarm node based on the knowledge graph.
[0085] It should be noted that the knowledge graph is built based on the operation and maintenance data of the cloud platform. The construction process of the knowledge graph specifically includes: continuously collecting multimodal operation and maintenance data from the cloud platform's operating environment based on a preset data acquisition adapter and preset acquisition probes; among which, multimodal operation and maintenance data includes infrastructure configuration data, container event data, application business log data, and kernel-level data; identifying entity objects and entity relationships from the multimodal operation and maintenance data, and constructing a knowledge graph with entity objects as nodes and entity relationships as edges.
[0086] It is understood that this embodiment aims to continuously collect comprehensive, multimodal operational data from private and hybrid cloud environments in a non-intrusive or low-intrusive manner by deploying a modular and scalable data acquisition adapter and acquisition probes. The multimodal operational data specifically includes, but is not limited to, infrastructure configuration data, container event data, application service log data, and kernel-level data. Furthermore, this embodiment uses different adapters for different types of operational data; for example, the infrastructure and virtualization adapter is used to collect infrastructure configuration data, the cloud-native and container adapter is used to collect container event data, and the application and service monitoring adapter is used to collect application service log data. For kernel-level data, this embodiment uses a kernel-level event acquisition probe for collection.
[0087] In specific implementations, the infrastructure and virtualization adapters are deeply integrated with cloud operating systems (such as OpenStack), virtualization management platforms (such as vCenter), and standard APIs (Application Programming Interfaces) for physical device management interfaces. This allows them to obtain detailed configuration information (CI, Configuration Item), performance metrics (such as CPU steal time, disk I / O latency, and network packet loss rate), state change events, and physical and logical topological relationships between virtual machines, physical hosts, virtual switches, data storage volumes, network devices, etc.
[0088] Cloud-native container adapters capture real-time lifecycle events (such as creation, destruction, scaling, and migration) and their dynamic relationships of various resources, including Pods, Services, Deployments, ConfigMaps, and Ingress, by listening to the API Server event stream of container orchestration engines (such as Kubernetes) and the data plane of service meshes (such as Istio). Simultaneously, they collect container-level performance metrics, comparisons of resource requests / limits with actual usage, and key service metrics such as inter-service traffic topology, latency, and success rate.
[0089] Application business monitoring adapters pull or receive application performance metrics (such as application response time, error rate, throughput), distributed call chain data, and structured or unstructured application and system logs from application performance monitoring (APM) tools (such as SkyWalking, Zipkin), distributed tracing systems, and centralized log management platforms (such as ELK Stack, Loki).
[0090] Kernel-level event capture probes are low-overhead, high-precision probes based on eBPF (extended Berkeley Packet Filter) deployed within critical physical host servers or core business virtual machines. This probe securely captures fine-grained events directly at the operating system kernel level without requiring application code modification or service restarts. Specific data collected includes:
[0091] Network communication: By attaching to system calls such as connect, accept, send / recv, etc., the actual network data packet sending and receiving between processes can be accurately captured, and a real-time service communication topology accurate to the process / port level can be constructed. This is crucial for discovering unknown dependencies, identifying "shadow IT", and understanding service relationships in encrypted traffic.
[0092] File and disk I / O (Input / Output): By tracking calls such as open, read, and write, the access patterns and I / O performance of critical configuration files and data files can be monitored, and performance problems caused by disk contention or configuration errors can be accurately located.
[0093] Process activity: Monitor events such as process creation, destruction, and CPU scheduling to diagnose and detect abnormal process behavior.
[0094] This kernel-level data is crucial for accurately depicting the real-world interaction dependencies between application components during runtime and for identifying the root causes of performance bottlenecks. It overcomes the limitation of traditional monitoring, which can only see "points" (metrics) but not the "lines" (real interactions). In other words, the kernel-level probe in this embodiment is implemented based on extended Berkeley packet filtering technology. It is used to capture fine-grained events such as process-level network connections, file accesses, and system calls to discover the actual interaction dependencies between services during runtime, thus serving as implicit relationships in the knowledge graph.
[0095] Furthermore, this embodiment identifies entity objects and entity relationships from the collected multimodal operation and maintenance data, and constructs a knowledge graph using entity objects as nodes and entity relationships as edges. Entity objects include physical entities, logical entities, and business entities; entity relationships include explicit dependencies obtained from target configuration information and implicit dependencies obtained based on data analysis methods. That is, the knowledge graph, in the form of nodes and edges, uniformly models physical entities, logical entities, business entities, and their explicit and implicit dependencies in the cloud platform environment.
[0096] It's important to note that the collected multimodal operation and maintenance data is sent to the processing engine via a high-throughput message queue, where it undergoes unified modeling and fusion. The processing engine identifies and standardizes all collected objects, assigning a globally unique identifier (GUID) to each unique entity, such as a physical server, virtual machine, container, application service, business process, or even a cost center. These entities are then abstracted into nodes in a knowledge graph. Node attributes not only include static configuration and dynamic state but also innovatively incorporate business and cost dimensions. Business attributes refer to the business system to which the entity belongs, the service level agreement (SLA), the business owner, the product engineer, and the product line; cost attributes refer to resource unit price, estimated operating costs, and the cost center to which the entity belongs.
[0097] In addition to identifying entity objects, it is also necessary to identify entity relationships. Understandably, the processing engine also needs to be able to automatically extract entity relationships from multimodal operational data. These relationships are clearly distinguished into different types, and implicit dependencies can be discovered. Entity relationships include explicit dependencies obtained from target configuration information, i.e., deterministic relationships that can be directly obtained from APIs or target configuration information, such as installedOn, runsOn, connectsTo, usesStorage, belongsTo, exposedAs, etc. Entity relationships also include implicit dependencies obtained based on data analysis methods, i.e., non-deterministic relationships inferred through data analysis, such as discovering that service A frequently calls service B based on eBPF traffic statistics, or discovering that changes to configuration item X are strongly correlated with a performance degradation in service Y based on association rules.
[0098] In addition, it can also include business and change relationships, such as support (Business) and was changed (Changed). By modeling change events of CI / CD (Continuous Integration / Continuous Deployment) pipelines and configuration management repositories as relationships and associating them with the resource nodes of the change, the data of development and operation are connected.
[0099] It should also be noted that the entire knowledge graph is stored in a high-performance graph database (such as Neo4j or JanusGraph). The processing engine employs an event-driven incremental update mechanism. Any event from the data source (such as a Pod deletion event from the Kubernetes API) triggers a workflow that performs near real-time creation, modification, or deletion operations on the corresponding nodes and relationships in the graph database, ensuring that the knowledge graph remains synchronized with the real world within seconds, achieving real-time updates.
[0100] In this way, through the above collection process, not only is "explicit" data from traditional sources such as cloud platform APIs and APM integrated, but also "implicit" data collected through eBPF kernel probes that reflect the real interactions during system operation, such as process-level network topology, is innovatively integrated, thereby constructing a high-fidelity cloud platform digital twin that is highly synchronized with the physical and logical world, namely the aforementioned knowledge graph.
[0101] Furthermore, once a raw alarm message is obtained (such as "Pod `cart-service-abcde` CPU utilization reaches 95%), by parsing the raw alarm and extracting the main identifier, such as hostname, IP, and Pod name, the corresponding Pod node, which is the alarm node, can be accurately located in the knowledge graph through multi-dimensional indexing.
[0102] In a specific implementation, the above-mentioned generation of context information for each alarm node based on a knowledge graph includes: using each alarm node as a central node, performing a graph traversal operation in the knowledge graph to obtain target information that satisfies a preset relationship with the central node; and assembling the target information to generate the context information for each alarm node. That is, in this embodiment, using each alarm node as a central node, a series of parallel, multi-hop graph traversal queries are performed in the knowledge graph to obtain target information that satisfies a preset relationship with the central node, and then this target information is assembled to generate the context information for each alarm node.
[0103] In a specific implementation, a graph traversal operation is performed in the knowledge graph to obtain target information that satisfies a preset relationship with the central node. This includes: performing a graph traversal operation in the knowledge graph to obtain underlying resource information that satisfies a preset vertical dependency relationship with the central node; performing a graph traversal operation in the knowledge graph to obtain service information that satisfies a preset horizontal dependency relationship with the central node; the service information includes first service information that the central node depends on and second service information that depends on the central node; and performing a graph traversal operation in the knowledge graph to obtain change information of the central node and its adjacent nodes within the alarm time window. That is, this embodiment mainly obtains target information that satisfies a preset relationship with the central node from three dimensions.
[0104] The first dimension is vertical dependency chain tracing, which forms a complete "technology stack" snapshot by tracing all the underlying supporting resources of the central node, such as Pod:cart-service-abcde->running on->Node:worker-03->running on->VM:ubuntu-vm-123->running on->Host:blade-server-08->connected to->Switch:core-sw-01,port:ge-0 / 0 / 1.
[0105] The second dimension is horizontal impact domain analysis, which involves expanding the scope of all dependent or dependent peer services of the central node and relating them to business processes. For example, Pod:cart-service-abcde <- call -Service:frontend-service -> support -> Business Process: user order placement process. This allows for immediate prediction of the business scope that a failure might affect.
[0106] The third dimension is the time dimension correlation, which involves querying other change information of the central node and its adjacent nodes within the alarm time window. Specifically, this includes (1) recent change records, such as whether there was an image update of the Deployment associated with the Pod 5 minutes ago? Has the related ConfigMap been changed? (2) related alarms and events, such as whether the host machine worker-03 also reported a memory pressure alarm 2 minutes ago? Is the cluster performing node maintenance? (3) performance indicator snapshots, such as detailed time-series data of key indicators such as CPU, memory, network IO, and application latency of the Pod, its host machine, and the frontend-service that calls it before and after the alarm occurs. (4) related log fragments, such as automatically retrieving and extracting log summaries from Loki that match the Pod before and after the alarm time point and contain keywords such as "error", "exception", and "timeout".
[0107] Furthermore, all extracted context information can be structured and organized into a standardized JSON object. This object not only contains the technical context but also a Business Impact Statement (SLA), automatically generated by the system based on business attributes and dependencies within the graph. For example: the user order placement process may be disrupted; this business has the highest SLA, the responsible party is Zhang San, and the initial estimate is a 5% impact on transaction success rate, potentially leading to financial losses. The SLA faces default risk.
[0108] As can be seen, when an alarm occurs, this application starts with the alarm subject and performs multi-directional, multi-hop traversal in the knowledge graph to automatically capture its complete context, including its technology stack, dependent services, recent changes, relevant indicator logs, and upper-layer business support relationships. In other words, by locating alarm nodes and generating contextual information through the knowledge graph, isolated raw alarm information is bound to deeply related data such as system topology and dependency relationships. This avoids the information silo problem in traditional alarm analysis and provides comprehensive correlation evidence for subsequent noise reduction and root cause localization, fundamentally improving the accuracy of the analysis results.
[0109] Step S12: Based on the context information, perform noise reduction processing on each original alarm information to eliminate invalid alarm information and obtain the noise-reduced target alarm information.
[0110] In this embodiment, in the face of alarm storms, noise reduction processing can be performed based on context information, which can effectively identify and eliminate invalid alarm information, thereby compressing thousands of original alarms into a few core target alarm events, reducing the total number of alarms faced by operation and maintenance personnel, eliminating the need for them to sift through massive amounts of redundant information to select key content, and significantly improving fault response efficiency.
[0111] In a specific implementation, noise reduction processing is performed on each original alarm message based on context information to eliminate invalid alarm messages. This includes: if a parent-child node relationship exists between the alarm nodes corresponding to each original alarm message based on context information, then the original alarm messages corresponding to the child nodes are suppressed; duplicate and similar alarm messages in each original alarm message are aggregated; if scene alarm messages under a preset target service scenario exist in each original alarm message, then the scene alarm messages are suppressed. That is, this embodiment can perform noise reduction processing on each original alarm message from multiple aspects based on context information to obtain the noise-reduced target alarm message.
[0112] In one specific embodiment, if there is a parent-child relationship between alarm nodes, the original alarm information of the child nodes is suppressed. For example, strong causal relationships such as runsOn and dependsOn in the knowledge graph are used. If physical machine A crashes and alarms, all unreachable alarms of virtual machines B, C, and D on it will be automatically determined as derived symptoms, thus being suppressed and merged under the alarm of A, forming an alarm tree rooted at A.
[0113] In one specific embodiment, the aggregation of duplicate and similar alarm information in each original alarm message includes: determining the alarm source of each original alarm message; if the same alarm source generates duplicate alarm information within a preset time window, then the duplicate alarm information is merged into one alarm message; calculating the similarity of each original alarm message using a preset text similarity algorithm, and obtaining similar alarm information with a similarity greater than a preset threshold, so as to aggregate the similar alarm information. That is, this embodiment can also aggregate duplicate and similar alarm information, for example, performing time window aggregation on duplicate alarms generated by the same alarm source due to state fluctuations within a short period of time, so as to merge them into one alarm message; at the same time, using a text similarity algorithm to aggregate alarm information from different alarm sources but describing similar problems, such as multiple web server instances simultaneously reporting database connection timeouts.
[0114] In one specific embodiment, if the original alarm information contains scenario alarm information under a preset business scenario, the scenario alarm information is suppressed. This includes: if the original alarm information contains instantaneous alarm information generated under a preset system self-healing scenario, the instantaneous alarm information is suppressed; if the original alarm information contains expected alarm information within a preset influence range, the expected alarm information is suppressed; if the original alarm information contains experimental alarm information corresponding to a preset chaos engineering experiment, the experimental alarm information is suppressed. That is, this embodiment can achieve more humanized and accurate alarm management based on semantic suppression of operation and maintenance and business scenarios by incorporating a deep understanding of operation and maintenance and business scenarios. Specifically, the system can identify system self-healing events such as virtual machine high availability, database master-slave switching, and K8s Pod elastic scaling. Instantaneous and expected alarm information generated during these processes will be automatically identified and suppressed, as this is a normal manifestation of the system's self-healing or adaptive process. Furthermore, alarms within a preset impact range can be suppressed. This is understandable, as the system is integrated with LCI / CD and the change management system. Therefore, any alarms occurring within an approved change window and related to the change object and its preset impact range (such as canary instances in a canary release) will be automatically marked as "change-related" and automatically muted or downgraded according to the policy. Simultaneously, if the system is connected to a chaos engineering platform, alarms caused by fault injection during chaos engineering experiments will also be automatically identified and isolated, marked as "chaos engineering experiments," to avoid interfering with normal operations and maintenance. In addition, the system can identify benign high-load patterns, such as the generation of monthly reports for nighttime data ETL (extraction, transformation, loading), through learning or predefined rules. It can also perform known problem association suppression. Through integration with a ticket system (such as Jira), if the characteristics of a new alarm highly match a known, ongoing problem, the system will not create a new event but will attach this alarm as new evidence to an existing ticket.
[0115] As can be seen, this application achieves high-precision compression and suppression of alarm storms by comprehensively utilizing the causal topology chain of knowledge graphs, the time lead / lag relationship of indicator sequences, the semantic similarity of log content, and the intelligent identification of operation and maintenance scenarios such as change windows and high availability switching.
[0116] Step S13: For each target alarm information, determine the corresponding candidate subgraph from the knowledge graph, and score each node in the candidate subgraph from at least two dimensions, so that the target node selected in descending order of total score is the root cause node; wherein, the candidate subgraph includes the target alarm node corresponding to the target alarm information and the neighboring nodes of the target alarm node.
[0117] In this embodiment, for each target alarm message obtained after noise reduction, a corresponding candidate subgraph needs to be determined from the knowledge graph. This candidate subgraph limits the scope of root cause investigation to the target alarm node and its adjacent nodes. Furthermore, a multi-dimensional scoring mechanism is used to filter root cause nodes. This multi-dimensional scoring method reduces the bias of single-dimensional judgment, making root cause location more accurate, while reducing reliance on the experience of senior operations and maintenance personnel.
[0118] Step S14: Generate a root cause analysis report for the root cause node and obtain a remediation plan that matches the root cause analysis report, so as to remediate the root cause node based on the remediation plan.
[0119] In this embodiment, after identifying the root cause node, a root cause analysis report is generated and a remediation plan is matched. The root cause location results are directly converted into executable remediation actions, thereby completing the remediation of the root cause node. This improves the timeliness and automation level of fault repair and reduces the lag and operational risks of manual intervention. Through the above scheme, intelligent operation and maintenance is achieved, which intelligently analyzes alarm context, automatically locates the root cause, and achieves closed-loop self-healing, improving fault response speed and reducing fault recovery time.
[0120] As can be seen, the knowledge graph in this application is built based on cloud platform operation and maintenance data. By locating alarm nodes and generating contextual information through the knowledge graph, isolated raw alarm information is bound to deeply related data such as system topology and dependencies. This avoids the information silo problem in traditional alarm analysis and provides comprehensive correlation evidence for subsequent noise reduction and root cause localization, improving the accuracy of analysis results from a fundamental level. Furthermore, this application performs noise reduction based on contextual information, which can effectively identify and eliminate invalid alarm information, reducing the total number of alarms faced by operation and maintenance personnel. This eliminates the need for them to sift through massive amounts of redundant information to select key content, significantly improving fault response efficiency. For each target alarm information, this application limits the scope of root cause investigation to the target alarm node and its adjacent nodes through candidate subgraphs, and uses a multi-dimensional scoring mechanism to screen root cause nodes. The multi-dimensional scoring method reduces the one-sidedness of single-dimensional judgment, making root cause localization more accurate, while reducing reliance on the experience of senior operation and maintenance personnel. Finally, this application generates a root cause analysis report and matches it with a remediation plan, directly transforming the root cause location results into executable remediation actions. This completes the repair of the root cause node, improving the timeliness and automation of fault repair, and reducing the lag and operational risks of manual intervention. Through the above solution, intelligent operation and maintenance is achieved, enabling intelligent analysis of alarm context, automated root cause location, and closed-loop self-healing, thereby improving fault response speed and reducing fault recovery time.
[0121] See Figure 2 As shown in the illustration, this application discloses a specific method for locating and repairing alarm root causes on a cloud platform. Compared to the previous embodiment, this embodiment further explains and optimizes the technical solution. Specifically, it includes:
[0122] Step S21: After receiving the original alarm information from the cloud platform, locate the corresponding alarm node from the preset knowledge graph, and generate the context information of each alarm node based on the knowledge graph; wherein, the knowledge graph is built based on the operation and maintenance data of the cloud platform.
[0123] Step S22: Based on the context information, perform noise reduction processing on each original alarm information to eliminate invalid alarm information and obtain the noise-reduced target alarm information.
[0124] Step S23: Determine the target alarm node corresponding to each target alarm information from the knowledge graph, and perform a search operation in the knowledge graph with the target alarm node as the starting node to obtain the adjacent nodes within the preset number of hops. Then, the range formed by the target alarm node and the adjacent nodes in the knowledge graph is used as a candidate subgraph.
[0125] In this embodiment, after obtaining the noise-reduced target alarm information, the root cause is further determined based on the target alarm information. First, for each alarm information, the target alarm node corresponding to each target alarm information needs to be determined from the knowledge graph. Then, with the target alarm node as the starting node, a search operation is performed in the knowledge graph to obtain adjacent nodes within a preset hop range, thereby using the range formed by the target alarm node and adjacent nodes in the knowledge graph as a candidate subgraph.
[0126] In a specific implementation, the target alarm node and its upstream dependent nodes within one to three hops can be defined as a candidate subgraph, and all nodes within the candidate subgraph become potential root cause candidates.
[0127] Step S24: Score each node in the candidate subgraph from at least two dimensions to obtain the score value corresponding to each dimension, and perform a weighted summation operation on the score values corresponding to each dimension to obtain the total score value of each node. Select a preset number of target nodes in descending order of total score value, and use the target nodes as root cause nodes.
[0128] In this embodiment, each node in the candidate subgraph needs to be independently scored from at least two dimensions to obtain the score value corresponding to each dimension and form a quantitative chain of evidence.
[0129] In a specific implementation, each node in the candidate subgraph is scored from at least two dimensions to obtain a score value for each dimension. This includes: calculating the prior probability of failure for each node based on historical data and determining a first score value based on the prior probability; acquiring time-series data of each node within the alarm time window and analyzing the time-series data to determine the target time point when the node's performance indicators become abnormal, then determining a second score value based on the chronological order between the target time point and the alarm time point corresponding to the target alarm node; evaluating whether each node has a correlation with a target historical risk change event before the alarm time point, and determining a third score value based on the correlation; wherein, the target historical risk change event is a change operation that occurs before the alarm time point and belongs to a preset risk type, and the correlation includes the time interval between the change occurrence time point of the target historical risk change event and the alarm time point being less than a preset interval threshold; and performing anomaly analysis on the log data corresponding to each node, and determining a fourth score value based on the anomaly analysis results. In other words, this embodiment mainly scores each node in the candidate subgraph from four dimensions to obtain a score value for each dimension.
[0130] The first score can be understood as a combination of historical failure and causal probability score. This embodiment uses historical event data and methods such as Bayesian networks to statistically determine the prior probability of failure for different types of resources, as well as the conditional probability of failure propagation along different types of relationships. The first score is then determined by combining the prior probability and conditional probability of the failure. It is understood that components that have historically been prone to failure will receive a higher score.
[0131] The second score can be understood as the leading score for time-series anomalies. In this embodiment, time-series data of each node in all candidate subgraphs within the alarm time window is obtained. Specifically, time-series data of multiple key performance indicators within a period before the alarm are obtained. Then, through advanced time-series analysis algorithms and cross-correlation analysis, the target time point when the node performance indicators of each node become abnormal is accurately calculated. The second score is then determined based on the chronological order between the target time point and the alarm time point corresponding to the target alarm node. In this process, the node whose performance indicators first show a significant inflection point is identified. This node is likely to receive a very high score and is one of the strongest signals for locating the root cause. That is, the closer the target time point of the inflection point is to the alarm time point, the higher the second score will be.
[0132] The third score can be understood as a recent change correlation score. This embodiment needs to assess whether each node has a correlation with a target historical risk change event before the alarm time. The target historical risk change event refers to a change operation that occurred before the alarm time and belongs to a preset risk type. Specifically, the correlation can be that the time interval between the change occurrence time of the target historical risk change event and the alarm time is less than a preset interval threshold. It can be understood that the recent change correlation score is mainly used to assess the potential contribution of a node's (such as a service or configuration item) change behavior to the current alarm. The core logic is: the closer the change is to the alarm time and the higher the risk, the greater the probability of causing a failure. Therefore, the higher the risk level of the target historical risk change event and the closer the change occurrence time is to the alarm time, the greater its correlation score. That is, this application assesses the third score by checking whether there is a target historical risk change event in the context of the node. Among them, the "suspicion" of a change event is determined by its risk level (such as the risk of modifying kernel parameters is higher than that of modifying tags), time proximity, and the degree of overlap between the change scope and the failure scope.
[0133] The fourth score can be understood as a score based on log and traced anomaly evidence. Specifically, this embodiment utilizes a domain-adjusted NLP (Natural Language Processing Model) model and template clustering algorithms to perform deep analysis on the log data associated with each node, identifying anomalous event templates with sudden frequency changes and rare content, and calculating their semantic relevance to the fault phenomenon. Based on the anomaly analysis results, the fourth score is determined. Furthermore, the distributed call chain can be analyzed to find spans exhibiting abnormal delays or error states; the corresponding service instances receive anomaly scores.
[0134] Furthermore, after obtaining the score value for each dimension, a weighted summation is performed on the scores for each dimension to obtain the total score value for each node. The total score value can also be understood as the root cause index, and the weighting weights can be dynamically optimized through machine learning. Then, a predetermined number of target nodes are selected in descending order of the total score value, and these target nodes are designated as root cause nodes. For example, the three target nodes with the highest root cause indices can be selected as root cause nodes. It is evident that this application uses multi-dimensional evidence, such as historical failure probability, temporal anomaly leadership, change correlation, and log / tracing anomalies, to quantitatively score and select a high-probability root cause candidate set. This multi-dimensional scoring method reduces the bias of single-dimensional judgments, making root cause localization more accurate, while reducing reliance on the experience of senior operations personnel.
[0135] Step S25: Generate a root cause analysis report for the root cause node, determine the pre-established contingency plan knowledge base, and check whether there is a remediation plan in the contingency plan knowledge base that matches the root cause analysis report.
[0136] In this embodiment, an analysis report is generated for the root cause node, and the system first checks whether there is a remediation plan matching the root cause analysis report in the pre-established contingency plan knowledge base. That is, for the identified root cause, a match is first made in the contingency plan knowledge base.
[0137] In a specific implementation, generating a root cause analysis report for the root cause node includes: obtaining scoring evidence for each dimension to determine the corresponding score value; inputting each scoring evidence and the root cause node into a first pre-trained model to output a root cause analysis report using the first pre-trained model.
[0138] In other words, in this embodiment, it is necessary to obtain the scoring evidence used to determine the corresponding score value for each dimension, such as time series diagrams, log fragments, change records, topology paths, etc., and input all scoring evidence and corresponding root cause nodes into the first pre-trained model to output a root cause analysis report. The first pre-trained model is a Large Language Model (LLM) fine-tuned using massive amounts of operational documents and an internal knowledge base. The LLM generates a structured root cause analysis report in natural language, clearly demonstrating its inference process. For example, the following root cause analysis report was obtained: "There is a 95% confidence that the root cause of the failure is [Virtual Machine VM-007]. The main evidence is as follows: 1) Timing evidence: The memory usage metric of this virtual machine spiked abnormally 3 minutes and 12 seconds earlier than the response latency alarms of all downstream applications. 2) Change evidence: 8 minutes before the alarm occurred, there was a high-risk kernel parameter change record (change ID: 12345) for this virtual machine. 3) Log evidence: The critical Out of memory: Kill process error message was found in its system log. 4) Topological evidence: All affected application services and their running Pods were ultimately traced back to running on this virtual machine."
[0139] In addition, the first pre-trained model can also support interactive "conversational diagnosis." For example, operations and maintenance personnel can use a chatbot-like interface to conduct "conversational" diagnosis of the event, asking the LLM further questions such as "What other changes have been made to VM-007 recently?", "Compare the memory usage patterns of VM-007 with another normal VM in the same cluster", and "What is the purpose of this kernel parameter change? What side effects might it have?". The LLM will query the knowledge graph and underlying data sources in real time based on the MCP protocol, providing intelligent answers and even executing read-only diagnostic commands, greatly improving the depth and efficiency of troubleshooting.
[0140] Step S26: If it exists, the remediation plan is directly obtained from the plan knowledge base; if it does not exist, the root cause analysis report and remediation target are input into the second pre-trained model to generate a remediation plan based on the training data, and the root cause node is remediated based on the remediation plan.
[0141] In this embodiment, if a remediation plan matching the root cause analysis report exists in the remediation plan knowledge base, the remediation plan can be directly retrieved from the knowledge base. If no remediation plan matching the root cause analysis report exists in the knowledge base, the root cause analysis report and the remediation objective (such as restoring the availability of the cart-service) are input into the second pre-trained model to generate a remediation plan based on the training data. The second pre-trained model is an LLM with code generation capabilities running in a secure sandbox. This LLM can dynamically generate a sequence of remediation steps, such as a set of kubectl commands or a Python script, based on its training on massive amounts of operational scripts, API documentation, and internal best practices.
[0142] Furthermore, after repairing the root cause node based on the repair plan, the process also includes: when the root cause node repair is detected as complete, if the repair plan is generated by the second pre-trained model, the repair plan is stored in the plan knowledge base and used as training data for the second pre-trained model. That is, if the repair plan generated by the second pre-trained model successfully repairs the root cause node, the repair plan will be stored in the plan knowledge base and used as training data for the second pre-trained model. It should also be noted that when maintenance personnel manually handle a new type of fault, the system will prompt them to record the solution, which will then be standardized and structured with the assistance of LLM, ultimately forming a new automated plan. Simultaneously, dynamic plans successfully generated by LLM and verified as effective will also be automatically added to the knowledge base, expanding its scope.
[0143] Furthermore, each event that is manually confirmed or successfully self-healed serves as high-quality labeled data. Through online learning and reinforcement learning from human feedback (RLHF), the probabilistic model, weight parameters, and behavior of the root cause inference engine are trained and fine-tuned, making its future inferences and decisions more accurate and closer to the level of experts.
[0144] This embodiment also supports adaptive adjustment of the dynamic baseline. If the repair operation is to expand the capacity, the system will automatically learn and adjust the future performance baseline of the resource to avoid normal high load after expansion being misjudged as abnormal.
[0145] The process of repairing the root cause node based on the remediation plan includes: continuously monitoring the business status of the root cause node; and stopping the repair operation once the business status confirms that the root cause node has returned to normal, thus determining that the root cause node repair is complete. In other words, after executing the remediation plan, the system enters an observation and verification period, continuously monitoring the business status, health status, and related business indicators of the root cause node to verify the repair effect with objective data, ensuring that the problem is truly resolved. Once it is confirmed that the root cause node has returned to normal, the repair operation stops, and the root cause node repair is determined to be complete.
[0146] In a specific implementation, root cause node repair based on the remediation plan includes: conducting a pre-test of the remediation plan in a preset sandbox environment; and determining the execution mode based on the risk level of the remediation plan after the pre-test passes. The execution modes include fully automatic, semi-automatic, and fully manual execution. Based on the execution mode, preset automated tools are invoked to execute the remediation plan to repair the root cause node. It is understood that the remediation plan will first undergo a pre-test in a preset sandbox environment to assess its potential impact and success rate. After the pre-test passes, the execution mode is determined based on the risk level of the remediation plan, and then the preset automated tools are invoked to execute the remediation plan to repair the root cause node. Specifically, the execution mode may include three modes: fully automatic execution, execution requiring manual approval (with a one-click approval button), or providing only a suggestion.
[0147] Furthermore, the system can not only analyze past failures but also continuously analyze time-series data in the knowledge graph to predict potential future risks. Using time-series prediction models (such as LSTM, Long Short-Term Memory networks), the system can predict that storage volume data-vol-01 will be full within 48 hours, or that the P99 latency of application order-service will exceed the SLA threshold in 3 hours. For these predicted problems, the system can proactively trigger low-risk intervention plans in advance, such as automatically cleaning up old logs, sending alerts to the team and suggesting capacity expansion, or even automatically performing a preventative service restart during off-peak traffic periods.
[0148] For more detailed processing procedures of steps S21 and S22, please refer to the relevant content disclosed in the foregoing embodiments, which will not be repeated here.
[0149] As can be seen, this application discloses a method that uses knowledge graphs as its core, deeply integrates multimodal data analysis, predictability analysis, and generative artificial intelligence, to achieve a shift from passive alarm response to a proactive, predictive, and self-healing intelligent operation and maintenance paradigm. Through the solution of this application, while ensuring the high availability and stability of cloud platform services, it is possible to reduce the mean time to detect (MTTD) and mean time to recover (MTTR) by orders of magnitude, ultimately improving business resilience and operational cost-effectiveness.
[0150] The solution proposed in this application can be deployed in an enterprise's data center or cloud environment. See [link / reference]. Figure 3 As shown, it consists of the following core functional modules: a data acquisition layer, a knowledge graph layer, an intelligent processing layer, an AI analysis engine, a self-healing execution layer, and a learning and evolution layer. The following section uses a typical microservice application failure scenario to illustrate the above solution in detail.
[0151] Scenario: An e-commerce application deployed on a Kubernetes (K8s) cluster experiences a sudden performance degradation in its "cart-service," resulting in high latency and partial failures when users add items.
[0152] Step 1: Ubiquitous data collection and real-time knowledge graph construction.
[0153] The various adapters and probes pre-deployed by the system continuously operate. The Kubernetes adapter, by subscribing to the API Server, has already recorded resources related to cart-service, such as Pods, Services, and Deployments, as nodes and their relationships (belongs to, exposed as) into the graph database. The APM adapter obtains from SkyWalking that cart-service supports the "user shopping" business process and has call relationships with its upstream frontend-service. More importantly, the eBPF probe deployed on the Kubernetes worker nodes captures frequent network communications between the cart-service Pod process and a MySQL-Pod at the kernel level, thus establishing a high-fidelity, runtime-connected implicit relationship in the knowledge graph, verified by runtime traffic, even though this dependency is not explicitly declared in any configuration file. Simultaneously, performance metrics (CPU, memory, network I / O), logs, and change records (via Jenkins / GitLab associations) of all components are continuously collected and associated with the corresponding nodes in the graph.
[0154] Step 2: Alarm generation, context enrichment, and intelligent noise reduction.
[0155] 1. Alarm Generation: The APM system first detects that the P99 response latency of the cart-service exceeds the SLA threshold, and the error rate spikes, issuing the first alarm. Subsequently, due to the slowdown of the cart-service, the upstream frontend-service also begins to time out, issuing a chain of alarms. Kubernetes may also issue a Pod restart alarm due to a failed health check of the cart-service. A small "alarm storm" has formed.
[0156] 2. Context Enrichment: Each alarm entering the system is immediately enriched by the intelligent analysis core. Taking the cart-service delayed alarm as an example, the system instantly pulls its context through graph traversal: it runs on Node-05, which is hosted on physical machine Host-10; it is called by frontend-service, supporting user shopping business; it depends on MySQL-Pod; and the StatefulSet associated with MySQL-Pod underwent a change 5 minutes ago.
[0157] 3. Intelligent Noise Reduction: The noise reduction module begins operation. Utilizing the topological relationships in the knowledge graph, it identifies that the frontend-service alarm is a downstream symptom of the cart-service alarm, thus suppressing it and merging it under the cart-service event. K8s Pod restart alarms are also identified as manifestations of the same issue. Ultimately, multiple original alarms are intelligently compressed into a single target alarm event centered around the cart-service.
[0158] Step 3: Multi-strategy root cause inference and interpretability analysis.
[0159] The root cause analysis engine initiates diagnostics for this event:
[0160] 1. Generate candidate set: The engine identifies the cart-service and its one-hop and two-hop neighbors (including MySQL-Pod, Node-05, Host-10, etc.) in the graph as candidate subgraphs.
[0161] 2. Multi-dimensional scoring: Time series analysis: The engine retrieved the time series of key performance indicators for all suspected nodes. Through cross-correlation analysis, it was found that the disk I / OWait metric of the node where MySQL-Pod resides had an abnormal inflection point 1 minute and 30 seconds earlier than the application latency inflection point of cart-service. This time-leading evidence assigned a very high suspicion score to MySQL-Pod and its underlying storage.
[0162] Change of association: The engine noticed that the StatefulSet associated with MySQL-Pod had a change record before the alert, which further increased its suspicion.
[0163] Log analysis: The engine automatically retrieves the logs of the MySQL-Pod and uses NLP models to discover a large number of abnormal log patterns related to slow queries and lock waits.
[0164] 3. Generative AI-Enhanced Adjudication and Reporting: Based on all scores, the engine identified MySQL-Pod as the most probable root cause. This structured evidence (time-series comparison charts, change details, log snippets) was submitted to the Large Language Model (LLM). The LLM generated a natural language analysis report: "Root Cause Diagnosis: The root cause of the cart-service performance degradation is a severe performance bottleneck in its dependent MySQL database instance (MySQL-Pod) with a 98% confidence level. Key evidence includes: 1) The database node's disk I / O wait time was 90 seconds earlier than the application latency; 2) The database logs were filled with slow query and lock wait errors; 3) The database underwent a configuration change 5 minutes before the failure."
[0165] Step 4: Closed-loop self-healing and knowledge evolution.
[0166] 1. Dynamic Contingency Plan Generation and Execution: Based on the root cause of "MySQL performance bottleneck," the system searches for matching remediation plans in the contingency plan knowledge base. If no existing plan is found, the LLM (Local Management Module) can be invoked. Based on its training in MySQL tuning knowledge, it dynamically generates an operational suggestion or script, such as: "It is recommended to execute the OPTIMIZETABLE command to optimize hot tables and analyze slow query logs to add indexes." After the contingency plan undergoes sandbox testing and one-click approval by operations personnel, it is executed through automated tools (such as Ansible).
[0167] 2. Effect Verification and Knowledge Accumulation: After execution, the system continuously monitors the latency and error rate of the cart-service and the I / O metrics of the MySQL-Pod. Once business operations are confirmed to have returned to normal, the event is closed. This successful "diagnosis-repair" case, along with all its context and evidence chain, is automatically stored in the operations and maintenance knowledge base. This new knowledge can be used for:
[0168] Improve the contingency plan knowledge base: The solution has been standardized into a new automated contingency plan.
[0169] Model retraining: Using high-quality labeled data, the root cause inference model and LLM are fine-tuned through the RLHF mechanism, so that they can diagnose more quickly and make more accurate decisions when facing similar problems in the future.
[0170] Through the above implementation methods, this application constructs a complete, efficient, and self-evolving intelligent operation and maintenance system that integrates data collection, intelligent analysis, and closed-loop self-healing.
[0171] This solution has the following effects:
[0172] 1. Order-of-magnitude improvement in fault response speed, significantly reducing MTTD and MTTR: Through automated, real-time knowledge graph construction, alarm noise reduction, and root cause inference, the manual troubleshooting process that previously took hours or even days is compressed to minutes or even seconds. Closed-loop self-healing capability further minimizes fault recovery time, thereby minimizing business downtime and ensuring high availability and continuity of services.
[0173] 2. Achieving a fundamental shift from "passive firefighting" to "proactive prediction and defense": This application not only diagnoses existing faults but also predicts potential capacity risks, performance bottlenecks, and SLA default risks through in-depth analysis of time-series data in the knowledge graph, triggering proactive intervention measures. This approach shifts operations and maintenance from a reactive, post-event remediation model to a proactive, pre-emptive risk avoidance model, fundamentally improving system robustness.
[0174] 3. Breaking down data and knowledge silos to provide a holistic and unified business insight: The knowledge graph disclosed in this application, for the first time, integrates previously isolated data dimensions such as infrastructure, cloud-native resources, applications, business processes, and even costs into a unified, real-time cognitive view. This enables the business impact of any technical event to be quantified and assessed instantly and accurately, helping decision-makers focus on the core issues that are truly important to the business and achieving a deep alignment between technology and business.
[0175] 4. Reduce reliance on experts and achieve the accumulation, inheritance, and widespread access to O&M knowledge: This application uses generative AI to transform the diagnostic logic and analysis process of top O&M experts into interpretable, interactive natural language reports and conversational diagnostic capabilities. This significantly lowers the barrier to advanced fault diagnosis, enabling ordinary O&M engineers to possess expert-level insights. Simultaneously, the system learns from each incident handling, distilling solutions into reusable automated contingency plans, building a continuously evolving organizational-level O&M knowledge asset.
[0176] 5. Enhance the level of automation and intelligence in operation and maintenance, and achieve significant cost reduction and efficiency improvement: Through intelligent noise reduction, automated root cause localization, and dynamic contingency plan generation and execution, this application frees operation and maintenance personnel from a large number of repetitive and tedious daily alarm handling and fault diagnosis work, enabling them to focus on higher-value architecture optimization and business innovation work, thereby significantly reducing labor costs and operating expenses while improving the quality of operation and maintenance.
[0177] See Figure 4As shown in the figure, this application discloses an alarm root cause location and repair device for a cloud platform, the device comprising:
[0178] The context information generation module 11 is used to locate the corresponding alarm node from the preset knowledge graph after receiving the original alarm information from the cloud platform, and generate the context information of each alarm node based on the knowledge graph; wherein, the knowledge graph is built based on the operation and maintenance data of the cloud platform.
[0179] The noise reduction module 12 is used to perform noise reduction processing on each original alarm information based on context information to eliminate invalid alarm information and obtain the target alarm information after noise reduction.
[0180] The root cause node determination module 13 is used to determine the corresponding candidate subgraph from the knowledge graph for each target alarm information, and to score each node in the candidate subgraph from at least two dimensions, so as to select the target nodes in descending order of total score as root cause nodes; wherein, the candidate subgraph includes the target alarm node corresponding to the target alarm information and the adjacent nodes of the target alarm node.
[0181] Repair module 14 is used to generate a root cause analysis report for the root cause node and obtain a repair plan that matches the root cause analysis report, so as to repair the root cause node based on the repair plan.
[0182] As can be seen, the knowledge graph in this application is built based on cloud platform operation and maintenance data. By locating alarm nodes and generating contextual information through the knowledge graph, isolated raw alarm information is bound to deeply related data such as system topology and dependencies. This avoids the information silo problem in traditional alarm analysis and provides comprehensive correlation evidence for subsequent noise reduction and root cause localization, improving the accuracy of analysis results from a fundamental level. Furthermore, this application performs noise reduction based on contextual information, which can effectively identify and eliminate invalid alarm information, reducing the total number of alarms faced by operation and maintenance personnel. This eliminates the need for them to sift through massive amounts of redundant information to select key content, significantly improving fault response efficiency. For each target alarm information, this application limits the scope of root cause investigation to the target alarm node and its adjacent nodes through candidate subgraphs, and uses a multi-dimensional scoring mechanism to screen root cause nodes. The multi-dimensional scoring method also reduces the one-sidedness of single-dimensional judgment, making root cause localization more accurate, while reducing reliance on the experience of senior operation and maintenance personnel. Finally, this application generates a root cause analysis report and matches it with a remediation plan, directly transforming the root cause location results into executable remediation actions. This completes the repair of the root cause node, improving the timeliness and automation of fault repair, and reducing the lag and operational risks of manual intervention. Through the above solution, intelligent operation and maintenance is achieved, enabling intelligent analysis of alarm context, automated root cause location, and closed-loop self-healing, thereby improving fault response speed and reducing fault recovery time.
[0183] Since the embodiments of the device part correspond to the embodiments described above, please refer to the embodiments described in the method part for the embodiments of the device part, and will not be repeated here.
[0184] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the cloud platform alarm root cause localization and repair method disclosed in any of the foregoing embodiments.
[0185] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0186] The processor 21 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 21 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 21 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0187] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored on it include operating system 221, computer program 222 and data 223, etc., and the storage method can be temporary storage or permanent storage.
[0188] The operating system 221 manages and controls the various hardware devices and computer programs 222 on the electronic device 20 to enable the processor 21 to perform calculations and processing on the massive amounts of data 223 in the memory 22. The operating system can be Windows, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the alarm root cause localization and repair method of the cloud platform executed by the electronic device 20 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the electronic device from external devices, as well as data collected by its own input / output interface 25.
[0189] Furthermore, this application also discloses a computer-readable storage medium storing a computer program. When the computer program is loaded and executed by a processor, it implements the alarm root cause localization and repair method steps of the cloud platform disclosed in any of the foregoing embodiments.
[0190] This invention also discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the alarm root cause localization and repair method for the cloud platform disclosed in any of the foregoing embodiments.
[0191] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0192] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0193] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art.
[0194] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0195] The above provides a detailed description of a cloud platform alarm root cause localization and repair method, apparatus, and device provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for locating and repairing the root cause of alarms on a cloud platform, characterized in that, include: After receiving the original alarm information from the cloud platform, the corresponding alarm node is located from the preset knowledge graph, and the context information of each alarm node is generated based on the knowledge graph; wherein, the knowledge graph is built based on the operation and maintenance data of the cloud platform; Based on the context information, each of the original alarm messages is denoised to eliminate invalid alarm messages, thereby obtaining the denoised target alarm message; For each target alarm information, a corresponding candidate subgraph is determined from the knowledge graph, and each node in the candidate subgraph is scored from at least two dimensions, so that the target nodes selected in descending order of total score are used as root cause nodes; wherein, the candidate subgraph includes the target alarm node corresponding to the target alarm information and the neighboring nodes of the target alarm node; A root cause analysis report is generated for the root cause node, and a remediation plan matching the root cause analysis report is obtained, so as to remediate the root cause node based on the remediation plan; The step of scoring each node in the candidate subgraph from at least two dimensions, and selecting the target nodes as root cause nodes in descending order of total score, includes: Each node in the candidate subgraph is scored from at least two dimensions to obtain a score value for each dimension; The score values corresponding to each dimension are weighted and summed to obtain the total score value for each node; A predetermined number of target nodes are selected according to the total score in descending order, and these target nodes are used as root cause nodes. The step of scoring each node in the candidate subgraph from at least two dimensions to obtain a score value for each dimension includes: The prior probability of failure for each node is calculated based on historical data, and a first score value is determined based on the prior probability of failure. Acquire the time-series data of each node within the alarm time window, and analyze the time-series data to determine the target time point when the node performance index becomes abnormal. Then, determine the second score value according to the order between the target time point and the alarm time point corresponding to the target alarm node. Assess whether each node is associated with a target historical risk change event before the alarm time point, and determine a third score based on the association; wherein, the target historical risk change event is a change operation that occurs before the alarm time point and belongs to a preset risk type, and the association includes the time interval between the change occurrence time of the target historical risk change event and the alarm time point being less than a preset interval threshold; Perform anomaly analysis on the log data corresponding to each node, and determine the fourth score value based on the anomaly analysis results.
2. The cloud platform alarm root cause location and repair method according to claim 1, characterized in that, The construction process of the knowledge graph includes: Multimodal operation and maintenance data is continuously collected from the cloud platform's operating environment based on a preset data acquisition adapter and preset acquisition probes; wherein, the multimodal operation and maintenance data includes infrastructure configuration data, container event data, application business log data, and kernel-level data; Entity objects and entity relationships are identified from the multimodal operation and maintenance data, and a knowledge graph is constructed using the entity objects as nodes and the entity relationships as edges.
3. The cloud platform alarm root cause location and repair method according to claim 2, characterized in that, The entity objects include physical entities, logical entities, and business entities; the entity relationships include explicit dependencies obtained from the target configuration information and implicit dependencies obtained based on data analysis methods.
4. The cloud platform alarm root cause localization and repair method according to claim 1, characterized in that, The generation of context information for each alarm node based on the knowledge graph includes: Using each alarm node as a central node, a graph traversal operation is performed in the knowledge graph to obtain target information that satisfies a preset relationship with the central node. The target information is assembled to generate context information for each alarm node.
5. The cloud platform alarm root cause localization and repair method according to claim 4, characterized in that, The step of performing a graph traversal operation in the knowledge graph to obtain target information that satisfies a preset relationship with the central node includes: Perform graph traversal operations in the knowledge graph to obtain underlying resource information that satisfies a preset vertical dependency relationship with the central node; A graph traversal operation is performed in the knowledge graph to obtain service information that satisfies a preset horizontal dependency relationship with the central node; the service information includes first service information that the central node depends on and second service information that depends on the central node. A graph traversal operation is performed in the knowledge graph to obtain the change information of the central node and its adjacent nodes within the alarm time window.
6. The cloud platform alarm root cause location and repair method according to claim 1, characterized in that, The step of performing noise reduction processing on each of the original alarm messages based on the context information to eliminate invalid alarm messages includes: If, based on the context information, it is determined that there is a parent-child node relationship between the alarm nodes corresponding to each of the original alarm messages, then the original alarm messages corresponding to the child nodes are suppressed. The duplicate and similar alarm information in each of the original alarm information is aggregated; If any of the original alarm messages contain scenario alarm messages under a preset standard service scenario, then the scenario alarm messages are suppressed.
7. The cloud platform alarm root cause location and repair method according to claim 6, characterized in that, The aggregation of duplicate and similar alarm information in each of the original alarm information includes: The alarm source of each original alarm message is determined. If the same alarm source generates duplicate alarm messages within a preset time window, the duplicate alarm messages are merged into one alarm message. The similarity of each of the original alarm messages is calculated using a preset text similarity algorithm, and similar alarm messages with similarity greater than a preset threshold are obtained to aggregate the similar alarm messages.
8. The cloud platform alarm root cause location and repair method according to claim 6, characterized in that, If any of the original alarm messages contain scenario alarm messages under a preset standard service scenario, then the scenario alarm messages are suppressed, including: If any of the original alarm messages contain instantaneous alarm messages generated under a preset system self-healing scenario, then the instantaneous alarm messages are suppressed. If any of the original alarm messages contain expected alarm messages within a preset range of influence, then the expected alarm messages are suppressed. If any of the original alarm messages contain an experimental alarm message corresponding to a preset chaos engineering experiment, then the experimental alarm message is suppressed.
9. The method for locating and repairing alarm roots on a cloud platform according to claim 1, characterized in that, For each of the target alarm messages, determining the corresponding candidate subgraph from the knowledge graph includes: Determine the target alarm node corresponding to each target alarm information from the knowledge graph; Starting with the target alarm node, a search operation is performed in the knowledge graph to obtain adjacent nodes within a preset number of hops. The range formed by the target alarm node and the adjacent nodes in the knowledge graph is used as a candidate subgraph.
10. The cloud platform alarm root cause localization and repair method according to claim 1, characterized in that, The generation of a root cause analysis report for the root cause nodes includes: Obtain scoring evidence for each dimension to determine the corresponding score value; Each of the scoring evidences and the root cause nodes are input into the first pre-trained model to output a root cause analysis report.
11. The method for locating and repairing alarm roots on a cloud platform according to claim 1, characterized in that, The process of obtaining a remediation plan that matches the root cause analysis report includes: Identify the pre-established contingency plan knowledge base and check whether there is a remediation plan in the contingency plan knowledge base that matches the root cause analysis report; If it exists, the repair plan will be retrieved directly from the plan knowledge base; If not, the root cause analysis report and remediation target are input into the second pre-trained model to generate a remediation plan based on the training data.
12. The cloud platform alarm root cause localization and repair method according to claim 11, characterized in that, After repairing the root cause node based on the repair plan, the process further includes: Once the root cause node repair is detected to be complete, if the repair plan is a plan generated by the second pre-trained model, the repair plan is stored in the plan knowledge base and used as training data for the second pre-trained model.
13. The cloud platform alarm root cause localization and repair method according to claim 12, characterized in that, The process of repairing the root cause node based on the repair plan also includes: Continuously monitor the business status of the root cause node; Once the root cause node is determined to have returned to normal based on the business status, the repair operation is stopped, and the root cause node repair is deemed complete.
14. The method for locating and repairing alarm roots on a cloud platform according to any one of claims 1 to 13, characterized in that, The repair of the root cause node based on the repair plan includes: The remediation plan is tested in a pre-set sandbox environment. After the pre-test is passed, the execution mode is determined according to the risk level of the remediation plan. The execution modes include fully automatic execution mode, semi-automatic execution mode, and fully manual execution mode. Based on the execution mode, the preset automated tools are invoked to execute the repair plan in order to repair the root cause node.
15. A cloud platform alarm root cause location and repair device, characterized in that, include: The context information generation module is used to locate the corresponding alarm node from a preset knowledge graph after receiving the original alarm information from the cloud platform, and generate context information for each alarm node based on the knowledge graph; wherein, the knowledge graph is built based on the operation and maintenance data of the cloud platform; The noise reduction module is used to perform noise reduction processing on each of the original alarm messages based on the context information to eliminate invalid alarm messages and obtain the target alarm message after noise reduction processing. The root cause node determination module is used to determine the corresponding candidate subgraph from the knowledge graph for each target alarm information, and to score each node in the candidate subgraph from at least two dimensions, so as to select the target nodes in descending order of total score as root cause nodes; wherein, the candidate subgraph includes the target alarm node corresponding to the target alarm information and the adjacent nodes of the target alarm node; The repair module is used to generate a root cause analysis report for the root cause node and obtain a repair plan that matches the root cause analysis report, so as to repair the root cause node based on the repair plan. Specifically, the root cause node determination module is used to score each node in the candidate subgraph from at least two dimensions to obtain a score value corresponding to each dimension; perform a weighted summation operation on the score values corresponding to each dimension to obtain a total score value for each node; and filter out a preset number of target nodes in descending order of the total score value, and use the target nodes as root cause nodes. The root cause node determination module is further configured to: statistically analyze the prior probability of failure for each node based on historical data, and determine a first score based on the prior probability of failure; acquire time-series data of each node within the alarm time window, and analyze the time-series data to determine the target time point when the node's performance indicators become abnormal, and then determine a second score based on the chronological order between the target time point and the alarm time point corresponding to the target alarm node; evaluate whether each node has a correlation with a target historical risk change event before the alarm time point, and determine a third score based on the correlation; wherein, the target historical risk change event is a change operation that occurs before the alarm time point and belongs to a preset risk type, and the correlation includes the time interval between the change occurrence time point of the target historical risk change event and the alarm time point being less than a preset interval threshold; perform anomaly analysis on the log data corresponding to each node, and determine a fourth score based on the anomaly analysis results.
16. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the alarm root cause localization and repair method for a cloud platform as described in any one of claims 1 to 14.
17. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when executed by a processor, the computer programs implement the steps of the alarm root cause localization and repair method for the cloud platform as described in any one of claims 1 to 14.
18. A computer program product comprising a computer program / instructions, characterized in that, When executed by a processor, the computer program / instruction performs the steps of alarm root cause localization and repair for the cloud platform as described in any one of claims 1 to 14.
Citation Information
Patent Citations
Monitoring fault analysis method fused with multi-modal knowledge base
CN120407272A