A large model-based microservice fault diagnosis and self-healing system

By building a microservice fault diagnosis and self-healing system based on a large model, the problem of time-consuming and labor-intensive traditional microservice fault diagnosis is solved, and rapid and accurate fault location and automated repair are achieved, thereby improving the availability and reliability of the system.

CN120610872BActive Publication Date: 2025-11-18INSPUR SOFTWARE TECH CO LTD

Patent Information

Application Number
CN202511113248.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-18
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Traditional microservice fault diagnosis relies on human experience, which is time-consuming and labor-intensive, and is prone to omissions or misjudgments in large-scale clusters, reducing system availability and reliability.

Method used

Construct a microservice fault diagnosis and self-healing system based on a large model, including multimodal data acquisition, deep analysis and self-healing execution layer. It uses the semantic understanding capability of the large model to automatically locate the root cause of the fault and realize automated repair, combined with knowledge graph and self-healing closed-loop control mechanism.

Benefits of technology

It enables rapid and accurate fault location and automated recovery in microservice architecture, reducing operation and maintenance costs and improving system availability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610872B_ABST
    Figure CN120610872B_ABST
Patent Text Reader

Abstract

The application provides a large model-based microservice fault diagnosis and self-recovery system, and belongs to the technical field of software engineering and artificial intelligence.The application collects multi-dimensional running data of each service instance under a microservice architecture through a data collection layer, and pre-processes the collected data to adapt to the input requirements of a large model.Subsequently, a large model analysis layer analyzes the pre-processed data based on a pre-trained fault diagnosis large model, accurately locates fault causes, and generates fault diagnosis results and corresponding self-recovery strategies.Finally, a self-recovery execution layer automatically executes repair operations to realize automatic fault recovery of microservices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of software engineering and artificial intelligence, and in particular to a microservice fault diagnosis and self-healing system based on a large model. Background Technology

[0002] As the complexity of internet and enterprise applications continues to increase, microservice architecture has been widely adopted due to its advantages such as flexibility, scalability, and maintainability. However, its distributed nature also brings many challenges in areas such as fault diagnosis and recovery. Traditional microservice fault diagnosis mainly relies on manual experience. When a microservice system fails, operations and maintenance personnel analyze log files, monitoring metrics, and other information to gradually troubleshoot the source of the fault. This approach is not only time-consuming and labor-intensive, but also prone to omissions or misjudgments when dealing with large-scale microservice clusters. The repair process also needs to be performed manually, reducing the availability and reliability of the system.

[0003] With the development of artificial intelligence technology, especially the breakthroughs of large models in natural language processing and data analysis, new opportunities have been provided for its application in the field of microservice fault diagnosis. Summary of the Invention

[0004] To address the above technical problems, this invention provides a microservice fault diagnosis and self-healing system based on a large model. This system can quickly and accurately locate the root causes of faults in a microservice architecture and achieve automated fault recovery, thereby improving system availability and reliability. Specific objectives include:

[0005] Build an automated operation and maintenance management system to achieve multi-dimensional real-time perception and cross-level correlation analysis of microservice faults, thereby reducing operation and maintenance costs and complexity.

[0006] By leveraging the semantic understanding capabilities of large models, we can construct efficient fault diagnosis and repair models, extract deep-level features from massive heterogeneous data, and improve the accuracy of fault diagnosis.

[0007] Establish a self-healing closed-loop control mechanism that includes diagnosis, decision-making, and verification to ensure the stable operation of the microservice architecture.

[0008] We continuously optimize the diagnostic accuracy and repair effectiveness of this system through online learning.

[0009] The technical solution of this invention is:

[0010] A microservice fault diagnosis and self-healing system based on a large model, including:

[0011] The multimodal data acquisition layer is used to collect runtime data from each service instance in the microservice architecture and to preprocess the collected data.

[0012] The large model analysis layer, connected to the data acquisition layer, uses a fault diagnosis and repair model to perform in-depth analysis on the preprocessed data, locate the root cause of microservice anomalies, and generate corresponding diagnostic results.

[0013] The self-healing execution layer, connected to the large model analysis layer, is used to execute self-healing strategies and automatically repair microservice faults.

[0014] Furthermore,

[0015] The data acquisition module establishes a real-time connection with the microservice instance through tools such as log collectors, performance monitoring agents, and network traffic sniffers. The types of data collected include log information, performance metrics, network communication data, and service configuration information.

[0016] Data cleaning algorithms are used to remove invalid and noisy data. Numerical data is standardized by normalization. Text data is preprocessed by word segmentation, stemming, and part-of-speech tagging, and key features are extracted from the data.

[0017] Furthermore,

[0018] The large model analysis module uses a pre-trained large model based on the Transformer architecture, which is trained on a large-scale microservice running dataset. It can automatically extract high-level semantic features of the data, generate vector representations of the microservice running status, and use classifiers or regression models for anomaly detection and root cause localization.

[0019] Based on the root cause diagnosis results and the fault handling rule base, combined with the microservice architecture topology information, a self-healing strategy is generated, which includes operations such as restarting service instances, redeploying, and adjusting resource configuration.

[0020] Furthermore,

[0021] The large model analysis layer includes anomaly model identification, causal reasoning engine, and dynamic knowledge base. Anomaly model identification detects anomalies and locates them. The causal reasoning engine analyzes the causes of the faults and provides repair suggestions. The dynamic knowledge base stores historical fault cases, repair solutions, and repair effect evaluation data, supporting the large model to continuously learn and optimize its diagnostic accuracy and repair effectiveness.

[0022] Furthermore,

[0023] The self-healing execution module is integrated with the orchestration tools, configuration management tools and container runtime environment of the microservice architecture. It can automatically perform self-healing operations, monitor the system status in real time during the execution process, verify the self-healing results, and record fault diagnosis and self-healing process information.

[0024] Furthermore,

[0025] The self-healing execution layer includes a security sandbox, a policy executor, and an effect verification module. The security sandbox supports the pre-execution of repair strategies in an isolated environment to prevent secondary failures during repair. The policy executor is responsible for performing repair operations. The effect verification module compares service health indicators through A / B testing and stores the evaluation results in the dynamic knowledge base of the large model analysis layer.

[0026] Furthermore,

[0027] The workflow is as follows:

[0028] Step 1: Collect service logs, performance metrics, network communication data, and configuration data in real time;

[0029] Step 2: Multimodal data alignment and feature fusion;

[0030] Step 3: Generate candidate root cause hypotheses from the large model;

[0031] Step 4: Causal credibility assessment based on knowledge graph;

[0032] Step 5: Generate a repair strategy and simulate its verification;

[0033] Step Six: Execute the repair command issued through the Kubernetes Operator;

[0034] Step 7: Collect feedback data to update the model.

[0035] The beneficial effects of this invention are:

[0036] This system introduces a fault diagnosis model optimized by deep learning technology and builds an automated operation and maintenance management system. It realizes automatic fault recovery and stable operation of microservice architecture, reduces operation and maintenance costs and complexity, and significantly improves the reliability and stability of microservice architecture. It is suitable for large-scale microservice clusters. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the workflow of the present invention;

[0038] Figure 2 This is a schematic diagram of the overall technical architecture of the present invention. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0040] This invention constructs an AI-Driven Microservice Fault Diagnosis and Autonomic Healing System (AMFD-AHS) based on a large model. The system collects multi-dimensional operational data from each service instance within the microservice architecture through a data acquisition layer, and preprocesses the collected data to adapt to the input requirements of the large model. Subsequently, the large model analysis layer performs in-depth analysis on the preprocessed data based on a pre-trained fault diagnosis model, accurately locating the root cause of the fault and generating fault diagnosis results and corresponding self-healing strategies. Finally, the self-healing execution layer automatically executes repair operations, achieving automatic fault recovery of the microservice.

[0041] The system includes the following key modules:

[0042] (1) Multimodal data acquisition layer

[0043] This module mainly includes log data collection, performance indicator collection, and network communication data extraction, and preprocesses the collected data to meet the input requirements of the large model analysis layer.

[0044] (2) Large Model Analysis Layer

[0045] This module mainly includes anomaly model identification, causal reasoning engine, and dynamic knowledge base. Anomaly model identification mainly performs anomaly detection to quickly find the location of anomalies. The causal reasoning engine is responsible for analyzing the cause of the fault and providing repair suggestions. The dynamic knowledge base stores historical fault cases, repair solutions, and repair effect evaluation data, and supports large models to continuously learn and optimize the diagnostic accuracy and repair effectiveness of large models.

[0046] (3) Self-healing execution layer

[0047] This module mainly includes a security sandbox, a policy executor, and an effect verification module. The security sandbox supports the pre-execution of repair policies in an isolated environment to prevent secondary failures during repair. The policy executor is responsible for performing repair operations. The effect verification module compares service health indicators through A / B testing and stores the evaluation results in the dynamic knowledge base of the large model analysis layer.

[0048] The workflow for using this system is as follows:

[0049] Step 1: Component Deployment and Data Acquisition

[0050] Deploy the fluentd component in the microservice instance to collect service logs in real time; deploy the prometheus monitoring component to collect performance metrics; and deploy the tcpdump component to collect network communication data.

[0051] Step 2: Multimodal data alignment and feature fusion

[0052] After the log text is segmented by the IK word segmenter, keyword vectors are extracted using the TF-IDF algorithm; numerical performance indicators are normalized using the Z-Score algorithm to eliminate dimensional differences; network data are extracted using a feature dimensionality reduction algorithm to extract key feature vectors, and finally, multi-source heterogeneous data are integrated into a unified feature matrix representation.

[0053] Step 3: Generate candidate root cause hypotheses from the large model

[0054] Pre-trained model architecture: A large pre-trained model based on Transformer is adopted. The model contains a 12-layer encoder, each with 768 hidden units, using the GELU activation function. Unsupervised pre-training is performed through masked language modeling (MLM) and next-sentence prediction (NSP) tasks. In the microservice fault diagnosis scenario, a multi-classification layer is added for fault root cause classification, and a regression layer is added for fault severity assessment during further fine-tuning.

[0055] Model training process: Historical fault scenario data, including normal and abnormal state samples, are collected and divided into training, validation, and test sets in an 8:1:1 ratio. Mixed precision training is used to accelerate model convergence. The AdamW optimizer is used with an initial learning rate of 3e-5, a training period of 10 epochs, and 32 samples per batch. An early stopping mechanism is used to prevent overfitting.

[0056] Step 4: Causal credibility assessment based on knowledge graph

[0057] A causal reasoning model based on knowledge graphs is constructed, representing service call relationships and configuration dependencies in the microservice architecture as a directed acyclic graph (DAG). When an anomaly is detected, a graph neural network (GNN) algorithm is used to calculate the causal correlation between nodes, and an attention mechanism is used to highlight key fault propagation paths, generating a probability distribution of the root cause hypothesis.

[0058] Step 5: Generate a remediation strategy and simulate its verification.

[0059] A sandbox environment is created using Kubernetes namespace isolation technology, and the remediation strategy to be executed is first verified on the image replica service in the sandbox. The sandbox environment shares the same configuration parameters and dependent services as the production environment, but access to external production systems is restricted through network policies to prevent accidental operations from causing cascading failures.

[0060] Step 6: Execute the repair command issued through the Kubernetes Operator

[0061] Establish a bidirectional gRPC connection with the Kubernetes API Server to translate remediation strategies into native commands, such as rolling updates (kubectl rollout), resource scaling (kubectl scale), and configuration updates (kubectl apply). During execution, the watch mechanism monitors Pod state changes in real time, and if an anomaly occurs, rollback to the state before execution.

[0062] Step 7: Collect feedback data to update the model

[0063] Deploy an Airflow scheduling workflow in the cluster to trigger an online model learning task every hour. Sample the latest fault data from the production environment, update the parameters of the large model through incremental training, and use a shadow model to verify the performance of the new model. Only switch versions when the model accuracy improves by more than 0.5%.

[0064] The above description is merely a preferred embodiment of the present invention and is used only to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A microservice fault diagnosis and self-healing system based on a large model, characterized in that, The multimodal data acquisition layer is used to collect runtime data from each service instance in the microservice architecture and to preprocess the collected data. The large model analysis layer, connected to the data acquisition layer, uses a fault diagnosis and repair model to perform in-depth analysis on the preprocessed data, locate the root cause of microservice anomalies, and generate corresponding diagnostic results. The self-healing execution layer, connected to the large model analysis layer, is used to execute self-healing strategies and automatically repair microservice faults. The data acquisition layer establishes a real-time connection with microservice instances through a log collector, a performance monitoring agent, and a network traffic sniffer. The types of data collected include log information, performance metrics, network communication data, and service configuration information. Data cleaning algorithms are used to remove invalid and noisy data. Numerical data is standardized by normalization. Text data is preprocessed by word segmentation, stemming, and part-of-speech tagging, and key features are extracted from the data. The large model analysis layer adopts a pre-trained large model based on the Transformer architecture, which is trained on a large-scale microservice runtime dataset. It automatically extracts high-level semantic features of the data, generates vector representations of the microservice runtime status, and uses classifiers or regression models for anomaly detection and root cause localization. Based on the root cause diagnosis results and fault handling rule base, combined with the microservice architecture topology information, it generates self-healing strategies including restarting service instances, redeploying, and adjusting resource configuration. The large model analysis layer includes anomaly model identification, causal reasoning engine, and dynamic knowledge base. Anomaly model identification detects anomalies and finds the location of anomalies. The causal reasoning engine is responsible for analyzing the cause of the failure and providing repair suggestions. The dynamic knowledge base stores historical failure cases, repair solutions, and repair effect evaluation data, supporting the large model to continuously learn and optimize the diagnostic accuracy and repair effectiveness. The self-healing execution layer is integrated with the orchestration tools, configuration management tools, and container runtime environment of the microservice architecture. It can automatically perform self-healing operations, monitor the system status in real time during the execution process, verify the self-healing results, and record fault diagnosis and self-healing process information.

2. The system according to claim 1, characterized in that, The self-healing execution layer includes a security sandbox, a policy executor, and an effect verification module. The security sandbox supports the pre-execution of repair strategies in an isolated environment to prevent secondary failures during repair. The policy executor is responsible for performing repair operations. The effect verification module compares service health indicators through A / B testing and stores the evaluation results in the dynamic knowledge base of the large model analysis layer.

3. The system according to claim 1, characterized in that, The workflow is as follows: Step 1: Collect service logs, performance metrics, network communication data, and configuration data in real time; Step 2: Multimodal data alignment and feature fusion; Step 3: Generate candidate root cause hypotheses from the large model; Step 4: Causal credibility assessment based on knowledge graph; Step 5: Generate a repair strategy and simulate its verification; Step Six: Execute the repair command issued through the Kubernetes Operator; Step 7: Collect feedback data to update the model.

Citation Information

Patent Citations

  • Fault self-recovery method and device

    CN117370054A

  • Micro-service fault diagnosis method and device based on large language model and electronic equipment

    CN117891640A

Cited By

  • Large-model-driven IT system integrated resource intelligent configuration and cooperative scheduling method

    CN121919006A

  • A communication network self-healing method based on a fault root cause analysis large model

    CN122533926A