Fault processing method and device for micro-service architecture and computer equipment

By real-time monitoring and failure analysis of microservices in smart grids, self-healing solutions are determined and self-healing treatment or service downgrades are solved, the existing microservice architecture lacks flexibility in fault handling, and the flexibility of fault handling and system stability are improved.

CN120011124APending Publication Date: 2025-05-16SHENZHEN COMTOP INFORMATION TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510116240.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing microservice architecture lacks flexibility in troubleshooting, cannot differentiate the importance and fault characteristics of different services, and it is difficult to effectively deal with the complex and changing operating conditions in power grid systems.

Method used

By monitoring microservices in the smart grid in real time, determine the fault condition and its self-healing plan, and perform self-healing treatment based on the fault condition and self-healing plan or degrade service to microservices with low priority.

Benefits of technology

It improves the flexibility of microservice architecture fault handling, and can automatically determine fault self-healing or service downgrade strategies in the smart grid to ensure the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011124A_ABST
    Figure CN120011124A_ABST
Patent Text Reader

Abstract

The invention relates to a fault processing method and device for a micro-service architecture and computer equipment. The method comprises the following steps: performing real-time monitoring on a first micro-service in a smart power grid to obtain real-time monitoring data of the first micro-service; according to the real-time monitoring data, determining a fault condition of the first micro-service and a self-healing scheme corresponding to the fault condition; and according to the fault condition and the self-healing scheme, carrying out self-healing processing on the first micro-service, or carrying out service degradation on a second micro-service of which the priority is lower than that of the first micro-service. By adopting the method, the fault processing flexibility of the micro-service architecture can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart grid technology, and in particular to a fault handling method, apparatus and computer equipment for a microservice architecture. Background Art

[0002] The microservice architecture in smart grid is a new software design pattern that can decompose the complex power grid system into a series of small, loosely coupled, independently deployable service units. Each microservice focuses on implementing a specific function in the power grid, such as real-time monitoring, data analysis, equipment control, user interaction, etc. This architectural design enables smart grids to respond to dynamic changes in power systems and diverse user needs more flexibly and efficiently. With the development of information technology, smart grid systems are gradually adopting microservice architecture to improve the flexibility and scalability of the system. Microservice architecture decomposes traditional large-scale monolithic applications into multiple small, loosely coupled services, each of which implements specific business functions, making the system easier to manage and maintain.

[0003] However, the current microservice architecture has shortcomings in fault handling. For example, it cannot perform differentiated processing based on the importance and fault characteristics of different services, and cannot effectively cope with the complex and changeable operating conditions in the power grid system.

[0004] Therefore, the current microservice architecture fault handling technology is not flexible enough. Summary of the invention

[0005] Based on this, it is necessary to provide a fault handling method, apparatus, computer equipment, computer-readable storage medium and computer program product for a microservice architecture that can improve flexibility in response to the above technical problems.

[0006] In a first aspect, the present application provides a fault handling method for a microservice architecture, comprising:

[0007] Performing real-time monitoring on a first microservice in a smart grid to obtain real-time monitoring data of the first microservice;

[0008] Determine, according to the real-time monitoring data, a fault condition of the first microservice and a self-healing solution corresponding to the fault condition;

[0009] According to the fault condition and the self-healing solution, the first microservice is self-healed, or a second microservice having a lower priority than the first microservice is downgraded.

[0010] In one embodiment, the self-healing process is performed on the first microservice according to the fault condition and the self-healing solution, or the service is downgraded for the second microservice whose priority is lower than the first microservice, including:

[0011] If the self-healing solution complies with a preset solution, self-healing the first microservice is performed; the preset solution includes restarting the service, switching to a standby node, or resetting resource restrictions;

[0012] When the self-healing solution does not conform to the preset solution or the fault condition conforms to the preset fault, the second microservice is downgraded; the preset fault includes excessive system load or critical service fault.

[0013] In one of the embodiments, when the self-healing solution complies with a preset solution, performing self-healing processing on the first microservice includes:

[0014] When the self-healing solution complies with the preset solution, the number of self-healing successes, the number of self-healing failures, and the average recovery time of the first microservice are obtained;

[0015] Determine the self-healing success rate according to the number of self-healing successes, determine the self-healing failure rate according to the number of self-healing failures, and determine the recovery efficiency according to the average recovery time;

[0016] The self-healing success rate, the self-healing failure rate, and the recovery efficiency are summed to obtain a reputation parameter of the first microservice;

[0017] Perform self-healing processing on the first microservice according to the reputation parameter.

[0018] In one embodiment, performing self-healing processing on the first microservice according to the reputation parameter includes:

[0019] When the reputation parameter exceeds a preset parameter threshold, increasing the processing coefficient of the self-healing process;

[0020] According to the increased processing coefficient, self-healing processing is performed on the first microservice.

[0021] In one embodiment, determining the fault condition of the first microservice and the self-healing solution corresponding to the fault condition according to the real-time monitoring data includes:

[0022] Extracting target data from the real-time monitoring data according to a preset sliding window;

[0023] Determining data differences between the target data and historical monitoring data;

[0024] When the data difference exceeds a preset difference threshold, determining that the first microservice fails;

[0025] The failure condition of the first microservice is determined according to the log of the first microservice.

[0026] In one embodiment, determining the fault condition of the first microservice and the self-healing solution corresponding to the fault condition according to the real-time monitoring data includes:

[0027] The real-time monitoring data is input into a trained recognition model to obtain the fault condition and the self-healing solution.

[0028] In a second aspect, the present application also provides a fault handling device for a microservice architecture, including:

[0029] A real-time monitoring module, used to perform real-time monitoring on a first microservice in a smart grid and obtain real-time monitoring data of the first microservice;

[0030] A solution determination module, used to determine the fault condition of the first microservice and the self-healing solution corresponding to the fault condition according to the real-time monitoring data;

[0031] The fault processing module is used to perform self-healing processing on the first microservice according to the self-healing solution, or to downgrade the service of a second microservice whose priority is lower than that of the first microservice.

[0032] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0033] Performing real-time monitoring on a first microservice in a smart grid to obtain real-time monitoring data of the first microservice;

[0034] Determine, according to the real-time monitoring data, a fault condition of the first microservice and a self-healing solution corresponding to the fault condition;

[0035] According to the fault condition and the self-healing solution, the first microservice is self-healed, or a second microservice having a lower priority than the first microservice is downgraded.

[0036] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0037] Performing real-time monitoring on a first microservice in a smart grid to obtain real-time monitoring data of the first microservice;

[0038] Determine, according to the real-time monitoring data, a fault condition of the first microservice and a self-healing solution corresponding to the fault condition;

[0039] According to the fault condition and the self-healing solution, the first microservice is self-healed, or a second microservice having a lower priority than the first microservice is downgraded.

[0040] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0041] Performing real-time monitoring on a first microservice in a smart grid to obtain real-time monitoring data of the first microservice;

[0042] Determine, according to the real-time monitoring data, a fault condition of the first microservice and a self-healing solution corresponding to the fault condition;

[0043] According to the fault condition and the self-healing solution, the first microservice is self-healed, or a second microservice having a lower priority than the first microservice is downgraded.

[0044] The fault handling method, apparatus, computer equipment, computer-readable storage medium and computer program product of the above-mentioned microservice architecture obtain real-time monitoring data of the first microservice by real-time monitoring of the first microservice in the smart grid, determine the fault condition of the first microservice and the self-healing plan corresponding to the fault condition based on the real-time monitoring data, and perform self-healing processing on the first microservice according to the fault condition and the self-healing plan, or downgrade the service of the second microservice with a lower priority than the first microservice; when a microservice in the smart grid fails, it can automatically determine whether to perform fault self-healing on the current microservice or to downgrade the service of other microservices with lower priorities, thereby improving the flexibility of fault handling of the microservice architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0046] Figure 1 A schematic diagram of a process flow of a method for handling a fault in a microservice architecture in one embodiment;

[0047] Figure 2A flowchart of a method for self-healing and service degradation of a microservice architecture in a smart grid in one embodiment;

[0048] Figure 3 A schematic diagram of a process flow of a method for handling a fault in a microservice architecture in another embodiment;

[0049] Figure 4 It is a structural block diagram of a fault handling device of a microservice architecture in one embodiment;

[0050] Figure 5 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0052] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0053] In an exemplary embodiment, Figure 1 As shown, a fault handling method for a microservice architecture is provided. This embodiment uses the method applied to a terminal as an example for illustration. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0054] Step S102: monitor the first microservice in the smart grid in real time to obtain real-time monitoring data of the first microservice.

[0055] The first microservice may be a key service for realizing core functions in the smart grid. The real-time monitoring data may be the operating status, performance indicators, logs, etc. of the first microservice monitored in real time.

[0056] In a specific implementation, the terminal can monitor the operating status, performance indicators, logs, etc. of the first microservice in the smart grid in real time to obtain real-time monitoring data.

[0057] For example, Prometheus (monitoring system) can be used as a monitoring tool, and a customized collector can be configured for each microservice in the smart grid. The collector can be used to regularly extract KPIs (Key Performance Indication) from the microservices, such as request volume, error rate, latency, etc., as well as operating status data and logs, as real-time monitoring data.

[0058] Step S104: determine the fault condition of the first microservice and a self-healing solution corresponding to the fault condition according to the real-time monitoring data.

[0059] The fault condition may be a specific service instance and cause of the first microservice fault. The self-healing solution may be a strategy for the first microservice to automatically repair the fault.

[0060] In a specific implementation, the terminal can determine whether the first microservice fails based on the real-time monitoring data. If no failure occurs, no processing is required and monitoring can continue. Otherwise, if it is determined that the first microservice fails based on the real-time monitoring data, the fault situation can be further determined, and a self-healing solution can be automatically recommended for the fault situation.

[0061] For example, the real-time monitoring data can be time series data. The time series data is analyzed using a sliding window. The real-time monitoring data in the sliding window is compared with the corresponding historical monitoring data. The difference between the two is used to determine whether the first microservice has failed. Specifically, the Z-Score (standard score) algorithm can be used to quantify the degree of abnormality between the real-time monitoring data and the historical monitoring data in the sliding window. When the obtained Z-Score exceeds the preset threshold, it is determined that the first microservice has failed. After that, the real-time monitoring data can be analyzed to obtain the fault situation and self-healing plan.

[0062] For another example, an artificial intelligence model can be pre-trained, and real-time monitoring data can be input into the artificial intelligence model. Potential faults can be predicted by the artificial intelligence model to obtain the fault condition of the first microservice. The artificial intelligence model can also recommend the best self-healing strategy for potential faults as the self-healing solution corresponding to the fault condition.

[0063] Step S106: According to the fault condition and the self-healing solution, the first microservice is self-healed, or the second microservice having a lower priority than the first microservice is downgraded.

[0064] Among them, the self-healing process can be to automatically repair the first microservice. Service degradation can be to reduce the priority of the microservice. The second microservice can be a non-critical service that implements non-core functions in the smart grid. It is understandable that the first microservice can be a critical service that is critical to the operation of the system, such as real-time monitoring and equipment failure, and the second microservice can be a non-critical service that is relatively less important, such as data analysis and user interaction.

[0065] In a specific implementation, the terminal may prioritize self-healing the first microservice when the self-healing solution meets the specified solution, and downgrade the service of the second microservice with a lower priority than the first microservice when the self-healing solution does not meet the specified solution or the fault condition meets the specified fault.

[0066] For example, when a fault is detected and a self-healing strategy is determined, if the self-healing strategy is a strategy that can quickly resolve the fault, such as restarting the service, switching to a backup node, or resetting resource limits, the self-healing strategy is first tried to be automatically executed. Otherwise, if the self-healing strategy cannot quickly resolve the fault, or the fault is caused by excessive system load or a critical service failure, the service degradation strategy can be triggered. Specifically, the current microservice can be regarded as a critical service, and non-critical services with a lower priority than the current microservice can be downgraded to reduce the resource consumption of non-critical services or shut down non-critical services to release resources, ensuring that critical services have sufficient resources to maintain operation, thereby ensuring that the core functions of the system are not affected.

[0067] The fault handling method of the above-mentioned microservice architecture obtains real-time monitoring data of the first microservice by real-time monitoring of the first microservice in the smart grid, determines the fault condition of the first microservice and the self-healing plan corresponding to the fault condition according to the real-time monitoring data, and performs self-healing processing on the first microservice according to the fault condition and the self-healing plan, or downgrades the service of the second microservice with a lower priority than the first microservice; when a microservice in the smart grid fails, it can automatically determine whether to perform fault self-healing on the current microservice or to downgrade the service of other microservices with lower priorities, thereby improving the flexibility of fault handling of the microservice architecture.

[0068] In an exemplary embodiment, the above step S106 may specifically include: when the self-healing solution complies with the preset solution, performing self-healing processing on the first microservice; the preset solution includes restarting the service, switching to a backup node, or resetting resource restrictions; when the self-healing solution does not comply with the preset solution or the fault condition complies with the preset fault, downgrading the service of the second microservice; the preset fault includes excessive system load or critical service failure.

[0069] Among them, the preset scheme can be a preset self-healing strategy. Restarting the service can be restarting the first microservice. Switching to the standby node can switch the first microservice to be operated by the standby node. Resetting resource limits can be resetting resource limits, for example, increasing resource limits. Preset faults can be preset fault service instances, causes, etc. System overload can be that the overall load of the smart grid exceeds a preset threshold. Critical service failures can be failures of microservices that are critical to system operation, for example, failures of real-time monitoring, device control, etc.

[0070] In a specific implementation, the terminal can perform self-healing processing on the first microservice when the determined self-healing solution is to restart the service, switch to a backup node, or reset resource restrictions. If the determined self-healing solution is not to restart the service, switch to a backup node, or reset resource restrictions, or the determined fault condition is excessive load or a critical service failure, the non-critical service, i.e., the second microservice, is downgraded.

[0071] In this embodiment, by performing self-healing processing on the first microservice when the self-healing solution conforms to the preset solution, and downgrading the service of the second microservice when the self-healing solution does not conform to the preset solution or the fault condition conforms to the preset fault, when a fault is detected, the self-healing strategy can be tried first, and when the self-healing strategy cannot quickly resolve the fault, or when the fault is caused by excessive system load or critical service failure, the service degradation strategy can be further triggered, thereby ensuring the reliable execution of critical services in the smart grid.

[0072] In an exemplary embodiment, the above-mentioned step of performing self-healing processing on the first microservice when the self-healing scheme conforms to the preset scheme may specifically include: when the self-healing scheme conforms to the preset scheme, obtaining the number of self-healing successes, the number of self-healing failures and the average recovery time of the first microservice; determining the self-healing success rate according to the number of self-healing successes, determining the self-healing failure rate according to the number of self-healing failures, and determining the recovery efficiency according to the average recovery time; summing the self-healing success rate, the self-healing failure rate and the recovery efficiency to obtain the reputation parameter of the first microservice; and performing self-healing processing on the first microservice according to the reputation parameter.

[0073] The number of successful self-healing times may be the number of times the self-healing operation is successfully executed. The number of failed self-healing times may be the number of times the self-healing operation is failed to be executed. The average recovery time may be the average time for the microservice to recover after the self-healing operation is successful. The self-healing success rate may be the probability of the self-healing operation being successfully executed. The self-healing failure rate may be the probability of the self-healing operation being failed to be executed. The recovery efficiency may be the efficiency of the microservice recovery after the self-healing operation is successful. The reputation parameter may be a parameter reflecting the self-healing performance.

[0074] In a specific implementation, when the determined self-healing solution is to restart the service, switch to a backup node, or reset resource restrictions, the terminal can obtain the number of self-healing successes S, the number of self-healing failures F, and the average recovery time A of the first microservice, and then calculate the self-healing success rate S / T, the self-healing failure rate F / T, and the recovery efficiency 1-A / B, where T=S+F is the total number of self-healing operations, and B is the benchmark recovery time, which is the ideal recovery time set by the system. Let the weight coefficient of the self-healing success rate be α, with a value range of [0, 1], 1-α represents the weight coefficient of the recovery efficiency, let the weight coefficient of the self-healing failure rate be β, with a value range of [0, 1], and take the weighted sum of the self-healing success rate, self-healing failure rate, and recovery efficiency to obtain the reputation parameter of the first microservice:

[0075] R = α (S / T) + (1-α) (1-(A / B)) - β (F / T),

[0076] According to the value of the reputation parameter, the proportion of the self-healing solution is adjusted. If R is higher than the first threshold Th1, the degree of automation of the self-healing operation can be increased. If R is lower than the second threshold Th2, the self-healing operation can be reduced and manual intervention can be increased.

[0077] It should be noted that the statistical time for the number of self-healing successes, the number of self-healing failures and the average recovery time may be a preset period, for example, one day or one week.

[0078] In this embodiment, when the self-healing plan conforms to the preset plan, the number of self-healing successes, the number of self-healing failures and the average recovery time of the first microservice are obtained, the self-healing success rate is determined according to the number of self-healing successes, the self-healing failure rate is determined according to the number of self-healing failures, and the recovery efficiency is determined according to the average recovery time. The self-healing success rate, the self-healing failure rate and the recovery efficiency are summed to obtain the reputation parameter of the first microservice. According to the reputation parameter, the first microservice is self-healed, and a balance can be achieved between fault self-healing and manual intervention to ensure reliable repair of the microservice architecture.

[0079] In an exemplary embodiment, the step of performing self-healing processing on the first microservice according to the reputation parameter may specifically include: when the reputation parameter exceeds a preset parameter threshold, increasing a processing coefficient of the self-healing processing; and performing self-healing processing on the first microservice according to the increased processing coefficient.

[0080] The preset parameter threshold may be a preset threshold of a reputation parameter. The processing coefficient may be a proportion of self-healing operations.

[0081] In a specific implementation, when the calculated reputation parameter exceeds the preset parameter threshold Th1, the terminal can increase the processing coefficient of the self-healing operation to obtain an increased processing coefficient. At this time, the processing coefficient of manual intervention is correspondingly reduced. Afterwards, the degree of automation of the recovery of the first microservice can be improved according to the increased processing coefficient to realize the self-healing operation of the first microservice.

[0082] In this embodiment, when the reputation parameter exceeds the preset parameter threshold, the processing coefficient of the self-healing process is increased, and the first microservice is self-healed according to the increased processing coefficient. When the reliability of the self-healing operation is high, the proportion of the self-healing operation can be increased, thereby improving the degree of automation of microservice recovery.

[0083] In an exemplary embodiment, the above step S104 may specifically include: extracting target data from real-time monitoring data according to a preset sliding window; determining the data difference between the target data and the historical monitoring data; when the data difference exceeds a preset difference threshold, determining that the first microservice has failed; and determining the failure condition of the first microservice according to the log of the first microservice.

[0084] The preset sliding window may be a sliding window of a preset length. The target data may be real-time monitoring data within the preset sliding window. The historical monitoring data may be historical monitoring data within the preset sliding window. The preset difference threshold may be a preset data difference threshold.

[0085] In a specific implementation, the terminal can extract the real-time monitoring data within a preset sliding window to obtain the target data, compare the target data with the historical monitoring data within the preset sliding window, and obtain the data difference between the target data and the historical monitoring data. If the data difference does not exceed the preset difference threshold, it is determined that the first microservice has not failed. Otherwise, if the data difference exceeds the preset difference threshold, it is determined that the first microservice has failed. At this time, the specific service instance or cause of the failure of the first microservice can be located by analyzing the log of the first microservice.

[0086] In this embodiment, by extracting target data from real-time monitoring data according to a preset sliding window, determining the data difference between the target data and the historical monitoring data, and when the data difference exceeds a preset difference threshold, determining that the first microservice has failed, and determining the failure condition of the first microservice according to the log of the first microservice, it is possible to automatically identify whether a microservice has failed and the specific failure condition when the failure occurs, thereby improving the efficiency of microservice fault handling.

[0087] In an exemplary embodiment, the above step S104 may specifically include: inputting the real-time monitoring data into a trained recognition model to obtain a fault condition and a self-healing solution.

[0088] Among them, the recognition model can be but is not limited to various artificial intelligence models.

[0089] In the specific implementation, the terminal can pre-train the recognition model, input the collected real-time monitoring data into the trained recognition model, predict the fault situation through the recognition model, and recommend the optimal self-healing strategy, so as to obtain the fault situation of the microservice and the corresponding self-healing plan.

[0090] In this embodiment, by inputting real-time monitoring data into a trained recognition model to obtain fault conditions and self-healing solutions, the fault conditions and self-healing solutions can be automatically obtained through artificial intelligence, thereby improving the efficiency of microservice fault handling.

[0091] In order to facilitate those skilled in the art to have a deeper understanding of the embodiments of the present application, a specific example will be described below.

[0092] The present application discloses a fault self-healing and service degradation method for a microservice architecture in a smart grid. In this method, when system resources are tight or some service failures occur, the service degradation strategy can ensure that the core business is not affected, and at the same time, by strategically sacrificing non-core services, the availability of the overall service is guaranteed. This flexible service management method not only reduces the economic losses caused by service interruptions, but also improves user satisfaction with the service. During the peak period of the power grid, by intelligently downgrading non-emergency services, the smooth operation of key grid operations such as dispatching and monitoring can be ensured, thereby ensuring the stability of the power grid while maintaining the normal electricity demand of users. In general, this fault self-healing and service degradation method provides efficient operation and maintenance support for the smart grid microservice architecture, improves the system's automation management level, and provides a strong guarantee for the stable operation and high-quality service of the power grid.

[0093] Figure 2 A flowchart of the fault self-healing and service degradation method of the microservice architecture in a smart grid is provided. Figure 2 ,The fault self-healing and service degradation method of the microservice architecture in the smart grid may include the following steps:

[0094] Step S201: Analyze the business requirements of microservices in the power grid, clarify the functions, performance indicators and dependencies of each microservice, design the microservice architecture, ensure high availability, high concurrency and scalability, and lay the foundation for fault self-healing and service degradation.

[0095] Step S202: deploy a real-time monitoring system to monitor the running status, performance indicators, logs, etc. of the microservices in real time. Set monitoring thresholds and issue warnings in a timely manner when indicators are abnormal.

[0096] Step S203: Utilize the data collected by the monitoring system and analyze the algorithm to automatically detect the failure of the microservice, locate the cause of the failure, and distinguish whether it is a problem with the service itself or a problem with the dependent service.

[0097] Step S204: Design corresponding self-healing strategies for different types of faults, such as restarting services, switching to standby nodes, etc. Introduce artificial intelligence algorithms to achieve fault prediction and intelligent recommendation of self-healing strategies. Train prediction models based on historical fault data and real-time monitoring data, predict potential faults in advance, and recommend the best self-healing strategy.

[0098] Step S205, set degradation priorities for different microservices according to business importance. Design degradation strategies, such as limiting access frequency and returning to default values. Returning to default values ​​means that when a microservice fails to work properly, the system can return a preset default value or simplified result to ensure the overall availability of the system, rather than completely stopping the service.

[0099] Step S206, when a fault is detected, the self-healing strategy is automatically executed. When the system load is too high or a critical service fails, the service degradation strategy is triggered. Specifically, when a fault is detected, the system will first try to automatically execute the self-healing strategy, which may include restarting the service, switching to a backup node, resetting resource limits, etc. If the self-healing strategy cannot quickly solve the problem, or the fault is caused by excessive system load or a critical service failure, the system will further trigger the service degradation strategy. The purpose of the service degradation strategy is to protect the core functions of the system and release resources by reducing the resource consumption of non-core functions or shutting down non-core functions.

[0100] When a service degradation policy is triggered, non-critical services will not be completely ignored, but the priority of non-critical services will be lower than that of critical services. The system may limit the resource usage of non-critical services, reduce the execution frequency of non-critical services, or temporarily stop non-critical services in some cases to ensure the normal operation of critical services. For example, during peak hours of the power grid, real-time monitoring and equipment control are critical services, while user interactions (such as non-urgent queries or report generation) may be non-critical services. At this time, user interaction services can be downgraded.

[0101] Step S207, simulate and test the self-healing and degradation strategies to verify their effectiveness and reliability. During the actual operation, continuously collect data and optimize strategies.

[0102] Step S208, compile a detailed fault self-healing and service degradation operation manual, including the architecture diagram of the self-healing system, a detailed description of the self-healing credit score algorithm, an operation guide for the service degradation strategy, a configuration tutorial for the monitoring system, and the steps of the fault handling process; at the same time, train the operation and maintenance personnel to ensure that they can master the self-healing and degradation operations.

[0103] Step S209: Continuously optimize the self-healing and degradation strategies according to the actual operation conditions, and regularly review and adjust the monitoring thresholds to improve the accuracy of fault detection and location.

[0104] In one embodiment, the above step S201 may specifically include: designing a flexible microservice architecture, which adopts the Spring Cloud (microservice architecture solution) framework and combines Netflix OSS (microservice architecture solution) components, such as Eureka (service discovery framework) for service registration and discovery, Hystrix (fault tolerance and delay tolerance library) to provide circuit breaker mode, and Zuul (API gateway server) as API (Application Programming Interface) gateway. Ensure that each microservice is independent and replaceable, and can be deployed through Docker (application container engine) containers, and automatically expanded and managed using Kubernetes (container orchestration engine).

[0105] In one embodiment, the above step S202 may specifically include: selecting Prometheus as a monitoring tool because it supports multi-dimensional data models and powerful query languages. Customized indicator collectors are configured for each microservice, which regularly extract key performance indicators from the service, such as request volume, error rate, and latency. At the same time, a dashboard is created using Grafana (monitoring tool) to intuitively display the above indicators. In addition, a multi-layer alarm mechanism can be set up, including alarm rules based on Prometheus and alarm notifications of Alertmanager (alarm management tool), to ensure that relevant personnel can be notified in time by email, SMS or instant messaging tools when indicators are abnormal.

[0106] In one embodiment, the above step S203 may specifically include: using sliding window technology to analyze time series data, and identifying anomalies by calculating the statistical difference between the data points in the window and the historical data. Specifically, the Z-Score algorithm is used to quantify the degree of abnormality of the data point. When the Z-Score of the data point exceeds a certain threshold, it is considered that a fault has been detected. In order to locate the fault, log analysis and distributed tracing technology are combined. Use ELK Stack (ElasticsearchLogstash Kibana Stack, a stack composed of collection, processing, and display) to collect and analyze logs, and Zipkin (distributed tracing system) or Jaeger (distributed tracing system) to track the call chain between services. Through the above tools, the specific service instance and cause of the failure can be quickly determined.

[0107] In one embodiment, the above step S204 may specifically include:

[0108] Initialization parameters: Set the baseline recovery time B, which is a preset ideal recovery time threshold. Set the weight coefficients α and β, which can be adjusted according to the specific needs of the system.

[0109] Data collection: After each self-healing operation, record S (number of successes), F (number of failures), and A (average recovery time).

[0110] Credit points calculation:

[0111] The reputation score R is calculated using the following formula:

[0112] R = α (S / T) + (1-α) (1-(A / B)) - β (F / T),

[0113] Among them, R is the reputation score of the microservice, which reflects the strength of its self-healing ability. S is the number of self-healing successes, which indicates the number of times the self-healing operation is successfully executed. F is the number of self-healing failures, which indicates the number of times the self-healing operation is failed. T is the total number of self-healing operations, that is, S+F. A is the average recovery time, which is the average time for service recovery after the self-healing operation is successful. B is the benchmark recovery time, which is the ideal recovery time set by the system and is used to compare the actual recovery efficiency. α is the weight coefficient of the success rate, with a value range of [0, 1], which is used to adjust the impact of the number of successes on the reputation score. β is the weight coefficient of the failure rate, with a value range of [0, 1], which is used to adjust the impact of the number of failures on the reputation score. Furthermore,

[0114] α (S / T) indicates the contribution of the self-healing success rate to the credit score. The more successful times, the greater the contribution.

[0115] (1-α) (1-(A / B)) represents the contribution of recovery efficiency to credit score. The closer the recovery time is to the benchmark time, the greater the contribution.

[0116] -β (F / T) indicates the negative impact of the self-healing failure rate on the credit score. The more failures there are, the greater the negative impact.

[0117] Adjust the self-healing strategy based on the value of R. If R is higher than a certain threshold, the degree of automation of the self-healing operation can be increased; if R is lower than a certain threshold, automation can be reduced and manual intervention can be increased.

[0118] Periodic updates: Reputation points are recalculated regularly every day or week to reflect the latest self-healing effects.

[0119] In one embodiment, the above step S204 may further specifically include: introducing a time decay factor γ to consider the impact of historical data on the current credit score.

[0120] R(t) = R(t-1) γ + ΔR(t). R(t) is the reputation score at the current time point. R(t-1) is the reputation score at the previous time point. γ is the time decay factor, which ranges from [0, 1] and is used to reduce the impact of historical data. ΔR(t) is the change in the reputation score at the current time point, and is calculated as follows:

[0121] ΔR(t) = α (S(t) / T(t)) + (1-α) (1-(A(t) / B)) - β (F(t) / T(t));

[0122] Where S(t) is the number of successful self-healing within time period t.

[0123] F(t) is the number of self-healing failures in time period t.

[0124] T(t) is the total number of self-healing operations within time period t, that is, T(t)=S(t)+F(t).

[0125] A(t) is the average recovery time in time period t.

[0126] The time period t is a specific time range, i.e., a period of time from a certain starting point to the current time point, which is used to calculate the change in the credit score to evaluate the effectiveness of the self-healing strategy. The time range can be fixed (for example, daily, weekly, or monthly) or dynamic.

[0127] In one embodiment, the above step S205 may specifically include: dynamically adjusting the strength of the self-healing operation according to the reputation score of the microservice through a feedback loop. The calculation formula of the reputation score is: Reputation score = (S / T) (1-(A / B)) -(F / T), where S is the number of successful self-healing attempts, T is the number of self-healing attempts, A is the average recovery time, B is the baseline recovery time, and F is the number of self-healing failures. When the reputation score of a microservice is higher than a certain threshold, the system automatically executes a series of predefined self-healing scripts, such as restarting the service, redeploying containers, or switching to redundant instances. If the reputation score is lower than the threshold, the system will reduce automated operations and trigger a manual review process instead.

[0128] In one embodiment, the above step S206 may specifically include: first, implementing a circuit breaker mode based on Hystrix, which can automatically disconnect requests when the backend service response time is too long or the failure rate is too high to prevent an avalanche effect. In addition, the client load balancing and retry mechanism are implemented using Ribbon and Feign (two important components in the Spring Cloud architecture) to ensure that healthy instances can be automatically switched when some service instances are unavailable. For non-core functions, a set of function degradation schemes are designed, which automatically disable certain functions when the system load is too high, such as reducing the frequency of data synchronization, shutting down non-essential data processing tasks, etc. These degradation strategies are managed through configuration files and can be dynamically adjusted at runtime.

[0129] In one embodiment, the above step S207 may specifically include: using a chaos engineering method to evaluate the actual effect of the self-healing and service degradation strategies. The Chaos Monkey (fault injection system) tool is used to randomly terminate service instances in the test environment to simulate real failure scenarios. At the same time, failure modes such as network delay, CPU pressure, and memory overflow are introduced through the Gremlin platform (chaos engineering platform). The response behavior of the system is recorded, including the triggering of self-healing operations, the execution of service degradation, and the overall recovery time of the system. Through these tests, it is verified that the self-healing strategy can work effectively under different failure scenarios, and the service degradation strategy can ensure the availability of core services.

[0130] In one embodiment, the above step S209 may specifically include: regularly collecting users' evaluations of system stability and service quality by designing a detailed questionnaire, wherein the questionnaire covers multiple aspects such as service interruptions encountered by users during use, recovery speed, effectiveness of notification mechanisms, and user acceptance of service degradation measures. The collected data is statistically analyzed to obtain user satisfaction scores, as well as user needs and expectations under different failure scenarios. Based on these feedbacks, targeted adjustments are made to the self-healing and service degradation strategies.

[0131] In this application, the beneficial effects of the fault self-healing and service degradation method of the microservice architecture in the smart grid are reflected in the stability and reliability of the system. Through the self-healing algorithm, the system can respond quickly when a fault is detected and automatically execute the repair process. This automated fault handling mechanism greatly shortens the fault recovery time and reduces delays and errors that may be caused by human intervention. The introduction of self-healing algorithms, such as the use of machine learning to predict failure modes and recommend repair measures, enables the system to make intelligent decisions when faced with unknown failures, thereby improving the robustness of the system. In addition, the seamless connection of the self-healing process ensures the continuity of critical business services. For an industry such as the power grid that has extremely high requirements for stability, this means that large-scale power outages can be effectively avoided, ensuring the safety and efficiency of power grid operation.

[0132] In this application, when system resources are tight or some service failures occur, the service degradation strategy can ensure that the core business is not affected, and at the same time, by strategically sacrificing non-core services, the availability of the overall service is guaranteed. This flexible service management method not only reduces the economic losses caused by service interruptions, but also improves user satisfaction with the service. During the peak period of the power grid, by intelligently downgrading non-emergency services, the smooth operation of key power grid operations such as dispatching and monitoring can be ensured, thereby ensuring the stability of the power grid while maintaining the normal electricity demand of users. Overall, this fault self-healing and service degradation method provides efficient operation and maintenance support for the smart grid microservice architecture, improves the system's automated management level, and provides a strong guarantee for the stable operation and high-quality services of the power grid.

[0133] In one embodiment, Figure 3 As shown, a fault handling method for a microservice architecture is provided, which is described by taking the method applied to a terminal as an example, and includes the following steps:

[0134] Step S301, real-time monitoring of a first microservice in a smart grid to obtain real-time monitoring data of the first microservice;

[0135] Step S302, input the real-time monitoring data into the trained recognition model to obtain the fault situation and the self-healing solution corresponding to the fault situation;

[0136] Step S303: if the self-healing solution complies with the preset solution, self-healing processing is performed on the first microservice; the preset solution includes restarting the service, switching to a backup node, or resetting resource restrictions;

[0137] Step S304: when the self-healing solution does not conform to the preset solution or the fault condition conforms to the preset fault, downgrade the service of the second microservice whose priority is lower than the first microservice; the preset fault includes excessive system load or critical service fault.

[0138] In a specific implementation, the terminal can monitor the operating status, performance indicators, logs, etc. of the first microservice in the smart grid in real time, obtain real-time monitoring data, input the real-time monitoring data into a pre-trained recognition model, predict faults through the recognition model, and recommend the optimal self-healing strategy to obtain the fault condition of the first microservice and the corresponding self-healing plan. If the self-healing plan is to restart the service, switch to a backup node, or reset resource restrictions, the self-healing plan of the first microservice is automatically executed. If the self-healing plan is not to restart the service, switch to a backup node, or reset resource restrictions, or the fault is caused by excessive system load or a critical service failure, the service of the second microservice with non-core functions is downgraded.

[0139] The fault handling method of the above-mentioned microservice architecture obtains real-time monitoring data of the first microservice by real-time monitoring of the first microservice in the smart grid, inputs the real-time monitoring data into a trained recognition model, obtains a fault condition and a self-healing plan corresponding to the fault condition, and performs self-healing processing on the first microservice when the self-healing plan meets the preset plan; when the self-healing plan does not meet the preset plan or the fault condition meets the preset fault, downgrades the service of the second microservice with a lower priority than the first microservice; when a microservice in the smart grid fails, it can automatically determine whether to perform fault self-healing on the current microservice or to downgrade the service of other microservices with lower priorities, thereby improving the flexibility of fault handling of the microservice architecture.

[0140] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0141] Based on the same inventive concept, the embodiment of the present application also provides a microservice architecture fault handling device for implementing the above-mentioned microservice architecture fault handling method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above-mentioned method, so the specific limitations in the embodiments of one or more microservice architecture fault handling devices provided below can refer to the limitations of the microservice architecture fault handling method above, and will not be repeated here.

[0142] In an exemplary embodiment, Figure 4 As shown, a fault handling device of a microservice architecture is provided, including: a real-time monitoring module 402, a solution determination module 404 and a fault handling module 406, wherein:

[0143] A real-time monitoring module 402 is used to monitor the first microservice in the smart grid in real time and obtain real-time monitoring data of the first microservice;

[0144] A solution determination module 404 is used to determine the fault condition of the first microservice and the self-healing solution corresponding to the fault condition according to the real-time monitoring data;

[0145] The fault processing module 406 is used to perform self-healing processing on the first microservice according to the self-healing solution, or to downgrade the service of a second microservice whose priority is lower than that of the first microservice.

[0146] In an exemplary embodiment, the above-mentioned fault processing module 406 is also used to perform self-healing processing on the first microservice when the self-healing plan conforms to a preset plan; the preset plan includes restarting the service, switching to a backup node, or resetting resource restrictions; when the self-healing plan does not conform to the preset plan or the fault condition conforms to a preset fault, downgrade the service of the second microservice; the preset fault includes excessive system load or critical service failure.

[0147] In an exemplary embodiment, the above-mentioned fault processing module 406 is also used to obtain the number of self-healing successes, the number of self-healing failures and the average recovery time of the first microservice when the self-healing plan conforms to the preset plan; determine the self-healing success rate according to the number of self-healing successes, determine the self-healing failure rate according to the number of self-healing failures, and determine the recovery efficiency according to the average recovery time; sum the self-healing success rate, the self-healing failure rate and the recovery efficiency to obtain the reputation parameter of the first microservice; and perform self-healing processing on the first microservice according to the reputation parameter.

[0148] In an exemplary embodiment, the fault processing module 406 is further configured to increase a processing coefficient of the self-healing process when the reputation parameter exceeds a preset parameter threshold; and perform self-healing processing on the first microservice according to the increased processing coefficient.

[0149] In an exemplary embodiment, the above-mentioned solution determination module 404 is also used to extract target data from the real-time monitoring data according to a preset sliding window; determine the data difference between the target data and the historical monitoring data; when the data difference exceeds a preset difference threshold, determine that the first microservice has a fault; and determine the fault condition of the first microservice according to the log of the first microservice.

[0150] In an exemplary embodiment, the solution determination module 404 is further configured to input the real-time monitoring data into a trained recognition model to obtain the fault condition and the self-healing solution.

[0151] Each module in the fault handling device of the above microservice architecture can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each of the above modules.

[0152] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (Near Field Communication, NFC) or other technologies. When the computer program is executed by the processor, a fault handling method of a microservice architecture is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0153] Those skilled in the art will understand that Figure 5The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0154] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0155] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0156] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0157] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0158] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0159] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0160] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A fault handling method for a microservice architecture, characterized in that: The method comprises: Performing real-time monitoring on a first microservice in a smart grid to obtain real-time monitoring data of the first microservice; Determine, according to the real-time monitoring data, a fault condition of the first microservice and a self-healing solution corresponding to the fault condition; According to the fault condition and the self-healing solution, the first microservice is self-healed, or a second microservice having a lower priority than the first microservice is downgraded.

2. The method according to claim 1, characterized in that The performing self-healing processing on the first microservice according to the fault condition and the self-healing solution, or downgrading the service of a second microservice having a lower priority than the first microservice, includes: If the self-healing solution complies with a preset solution, self-healing the first microservice is performed; the preset solution includes restarting the service, switching to a standby node, or resetting resource restrictions; When the self-healing solution does not conform to the preset solution or the fault condition conforms to the preset fault, the second microservice is downgraded; the preset fault includes excessive system load or critical service fault.

3. The method according to claim 2, characterized in that When the self-healing solution complies with the preset solution, performing self-healing processing on the first microservice includes: When the self-healing solution complies with the preset solution, the number of self-healing successes, the number of self-healing failures, and the average recovery time of the first microservice are obtained; Determine the self-healing success rate according to the number of self-healing successes, determine the self-healing failure rate according to the number of self-healing failures, and determine the recovery efficiency according to the average recovery time; The self-healing success rate, the self-healing failure rate, and the recovery efficiency are summed to obtain a reputation parameter of the first microservice; Perform self-healing processing on the first microservice according to the reputation parameter.

4. The method according to claim 3, characterized in that The performing self-healing processing on the first microservice according to the reputation parameter includes: When the reputation parameter exceeds a preset parameter threshold, increasing the processing coefficient of the self-healing process; According to the increased processing coefficient, self-healing processing is performed on the first microservice.

5. The method according to claim 1, characterized in that Determining the fault condition of the first microservice and the self-healing solution corresponding to the fault condition according to the real-time monitoring data includes: Extracting target data from the real-time monitoring data according to a preset sliding window; Determining data differences between the target data and historical monitoring data; When the data difference exceeds a preset difference threshold, determining that the first microservice fails; The failure condition of the first microservice is determined according to the log of the first microservice.

6. The method according to claim 1, characterized in that Determining the fault condition of the first microservice and the self-healing solution corresponding to the fault condition according to the real-time monitoring data includes: The real-time monitoring data is input into a trained recognition model to obtain the fault condition and the self-healing solution.

7. A fault handling device for a microservice architecture, characterized in that: The device comprises: A real-time monitoring module, used to perform real-time monitoring on a first microservice in a smart grid and obtain real-time monitoring data of the first microservice; A solution determination module, used to determine the fault condition of the first microservice and the self-healing solution corresponding to the fault condition according to the real-time monitoring data; The fault processing module is used to perform self-healing processing on the first microservice according to the self-healing solution, or to downgrade the service of a second microservice whose priority is lower than that of the first microservice.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Back-end system fault self-recovery method and device, storage medium and computer equipment

    CN120768800A