Method and system for microservice operation and maintenance
By introducing automated operation and maintenance systems of SLO detectors, trend analyzers, causal inferences and switch processors in microservice operation and maintenance, the problem of DevOps operation and maintenance relying on manual access is solved, and the operation and maintenance efficiency and system stability are improved.
Patent Information
- Application Number
- CN202210128858.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-11
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-02-11
AI Technical Summary
The existing DevOps operation and maintenance methods rely on manual access, resulting in low operation and maintenance efficiency, high cost, and difficulty in dealing with online problems and discovering implicit problems in a timely manner, affecting technology promotion and iteration.
A method and system for microservice operation and maintenance is proposed, combining SLO detectors, trend analyzers, causal inferences and switch processors to realize hierarchical management and fault elimination through automated operation and maintenance processes.
It improves operation and maintenance efficiency, reduces operation and maintenance costs, realizes automated operation and maintenance, reduces manual intervention, and improves system stability and fault handling capabilities.
Smart Images

Figure CN114462644B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network operation and maintenance, and in particular to a method and system for microservice operation and maintenance. Background Art
[0002] Many organizations divide development and system management into different departments. The driving force of the development department is usually "frequent delivery of new features", while the operation department is more concerned with the reliability of IT services and the efficiency of IT cost investment. The mismatch between the two goals creates a gap between the development and operation departments, thereby slowing down the speed at which IT delivers application value. As software engineering practices become more mature in major Internet companies,
[0003] In order to improve the above problems, technical means in the field of operation and maintenance are also gradually upgraded. For example, DevOps, which is quite popular at the current stage, is used to integrate the three aspects of development, technical operation and quality assurance within the enterprise to promote communication, collaboration and integration between development, technical operation and quality assurance departments. However, DevOps still relies more on manual access as the last mile to solve online problems. With limited labor costs, this may lead to a large number of alarms accumulating during the operation and maintenance process, making it difficult to focus on core problems, online problems are not handled in time, resulting in a chain reaction of failures, and hidden problems cannot be discovered in time in the early stage, resulting in more serious application failures. A series of problems, which are not conducive to the promotion and iteration of technology, and the cost is also huge.
[0004] Therefore, the industry urgently needs a more effective and reasonable solution that can improve operation and maintenance efficiency and reduce operation and maintenance labor costs. Summary of the invention
[0005] The embodiments of the present application propose a method and system for microservice operation and maintenance, which uses digitization, hierarchy, and evolution as the starting point for automated operation and maintenance, improves the efficiency of operation and maintenance, reduces the cost of operation and maintenance, and liberates part of the productivity of operation and maintenance personnel.
[0006] In a first aspect, an embodiment of the present application provides a microservice operation and maintenance system, the system comprising: an SLO detector, for collecting and outputting target time series data of the current level, the target time series data of the current level comprising time series data formed by indicator measurement values of the detection object corresponding to the current level and target time series data collected by the SLO detector of the next level; a trend analyzer, for analyzing the target time series data of the current level to identify problems and sending the problems to a causal inferencer; the causal inferencer, for inferring the cause of the problem, and determining a switch plan corresponding to processing the problem, and sending information indicating the switch plan to a switch processor, the switch plan being a strategy set to ensure stable operation of an application; the switch processor, for receiving the information indicating the switch plan and executing the switch plan.
[0007] In the second aspect, an embodiment of the present application proposes a method for microservice operation and maintenance, including: an SLO detector collects and outputs target time series data of the current level, the target time series data of the current level includes time series data formed by the indicator measurement value of the detection object corresponding to the current level and the target time series data collected by the SLO detector of the next level; a trend analyzer analyzes the target time series data of the current level to identify the problem, and sends the problem to a causal inference device; the causal inference device infers the cause of the problem, determines the switch plan corresponding to the problem, and sends information indicating the switch plan to the switch processor, the switch plan is a strategy set to ensure the stable operation of the application; the switch processor receives the information indicating the switch plan and executes the switch plan.
[0008] In an embodiment of the present application, a method and system for microservice operation and maintenance are proposed. The system combines SLO with switch plans to perform hierarchical management of various indicators in the operation and maintenance process, improves information density and value layer by layer, and triggers corresponding switches to eliminate risks or faults after analyzing the causes of the problems. A closed loop is formed between SLO and switch plans, and risk and fault response measures are continuously precipitated, thereby improving the model's automated operation and maintenance capabilities. Pattern self-nesting can be achieved at each level, which is convenient for model expansion and dynamic evolution, thereby improving operation and maintenance efficiency and reducing operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0010] Figure 1 This is a system architecture diagram of a microservice operation and maintenance model according to an embodiment of the present application;
[0011] Figure 2 It is a schematic diagram of the association relationship between indicators of SLO levels in one embodiment of the present application;
[0012] Figure 3 is a schematic flow chart of forming a switch plan according to an embodiment of the present application;
[0013] Figure 4 is a schematic block diagram of a method for executing a switch plan according to an embodiment of the present application;
[0014] Figure 5 It is a schematic diagram of the working principle of a trend analyzer according to an embodiment of the present application;
[0015] Figure 6It is a flowchart of a method for microservice operation and maintenance according to an embodiment of the present application.
[0016] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0018] To facilitate understanding, several terms involved in the embodiments of the present application are first introduced.
[0019] DevOps: DevOps is the abbreviation of Develop and Operations. It is the integration of development, technical operations and quality assurance within an enterprise. It is used to promote communication, collaboration and integration between development, technical operations and quality assurance departments. It is a culture, movement or practice that emphasizes communication and cooperation between "software developers (Dev)" and "IT operation and maintenance technicians (Ops)". By automating the processes of "software delivery" and "architecture changes", it makes building, testing and releasing software faster, more frequent and more reliable.
[0020] AIOps: Artificial Intelligence Information Technology (IT) Operations, refers to the method of applying artificial intelligence (AI) to improve IT operations. AIOps is a high-level implementation of enterprise-level DevOps on the operation and maintenance (technical operations) side. Specifically, AIOps uses big data, analysis, and machine learning capabilities to do the following: collect and aggregate the ever-increasing amount of operational data generated by multiple IT infrastructure components, applications, and performance inspection tools; intelligently filter "signals" from "noise" to identify important events and patterns related to system performance and availability issues; diagnose the root cause and report it to the IT department so that they can respond and remedy quickly, or in some cases automatically resolve these problems without human intervention.
[0021] By replacing multiple individual manual IT operations tools with a single intelligent automated IT operations platform, AIOps enables IT operations teams to respond faster and even proactively handle slowdowns and outages, significantly reducing workload.
[0022] Microservice architecture: It is an expansion of the distributed cloud architecture and a further extension of service-oriented architecture. By further refining services to form microservices and running them on a separate container platform, the elasticity and agility of the cloud can be achieved, such as elastic scaling, consistent online and offline environments, and improved automated operation and maintenance capabilities.
[0023] Operation and maintenance automation: convert some operation and maintenance work from manual execution to autonomous execution by models or systems.
[0024] Service level agreement (SLA): This can refer to a contract between an external service provider and a customer, or a contract between an IT department and the internal department it serves. The agreement document includes the services that the service provider or IT department will provide, as well as the performance standards that are expected to be achieved.
[0025] Service level objective (SLO): specifies a desired state of the functionality provided by the service, which can be a performance indicator associated with an SLA. SLOs are usually measured in terms of describing a benchmark or goal set by the parties regarding the service provided by the provider to the customer in a given period of time. For example, when used as a metric for a call center, an SLO may mean that 80% of incoming calls are answered by a service representative within one minute.
[0026] Switch plan: a method and means set up to ensure stable operation of the application and to stop the loss of failure in a timely manner. Optionally, in the embodiment of the present application, the switch plan can also be referred to as a switch.
[0027] Application: It can refer to a programming language provided to meet user needs in different fields and different problems, as well as a collection of application programs written in various programming languages.
[0028] The embodiments of the present application propose a microservice operation and maintenance model and corresponding training methods and usage methods to address the problems that the microservice DevOps operation and maintenance method is more dependent on organizational management during the implementation stage and is difficult to implement in a process-based manner, and the value stream transfer efficiency is low, resulting in missed opportunities for fault management. This model combines SLO with the switch plan to hierarchically manage various indicators in the operation and maintenance process, gradually improve the information density and value, and after analyzing the cause of the problem, trigger the corresponding switch to eliminate the risk or fault, forming a closed loop between the SLO and the switch plan, continuously precipitating risk and fault response measures, and improving the model's automated operation and maintenance capabilities. At the same time, an alarm interface is opened to facilitate relevant personnel to understand the system situation. This model can realize self-nesting of patterns at each level, which is convenient for model expansion and dynamic evolution. This model is suitable for microservice operation and maintenance scenarios, and can greatly improve operation and maintenance efficiency and reduce operation and maintenance costs.
[0029] It should be noted that in the embodiments of the present application, the system and the model are inter-connected concepts, that is, the microservice operation and maintenance model can also be referred to as the microservice operation and maintenance system.
[0030] The model proposed in the embodiment of the present application focuses on the three aspects of digitalization, hierarchy and evolution of operation and maintenance, in order to improve the efficiency of operation and maintenance, reduce the cost of operation and maintenance, and liberate part of the productivity of operation and maintenance personnel. Among them, digitalization attempts to use the effective indicator measurement values appearing in the system as the basis for decision-making, open up the information links between different modules within the system, avoid the emergence of information islands, and improve the accuracy of decision-making. Hierarchy emphasizes that there are hierarchical differences in the importance of indicators and the information density of alarms. Indicator data at different levels support decisions at different levels to improve decision-making efficiency. Indicators with problems or potential problems are processed through reasonable switch plans, and problems are eliminated and closed at the current level as much as possible. Evolution is based on the iteration of the indicator processing analysis model, and a series of problems that occur during the operation and iteration of the system are formed into a knowledge graph to provide guarantees for the effectiveness of system operation and maintenance, and even guide the technical evolution of the system.
[0031] Figure 1 : is a system architecture diagram of a microservice operation and maintenance model according to an embodiment of the present application. Figure 1 As shown in FIG. 1 , the model can be divided into multiple layers, each layer has a corresponding SLO. The highest layer SLO is the total SLO. As an example, the multiple layers from top to bottom may include the business layer, application layer, operating system layer, hardware layer, etc.
[0032] like Figure 1 As shown, each level includes an SLO detector, a trend analyzer, an anomaly alarm, a causal inference unit, and a switch processor.
[0033] Among them, each level has a corresponding direct detection object, and the direct detection object refers to the detection object corresponding to each layer of SLO itself. As an example, if the current level is at the operating system layer, the direct detection correspondence may include the amount of remaining memory space of the CPU. Optionally, the above-mentioned direct detection object may also be referred to as a detection object or object. The above-mentioned object may have one or more indicator measurement values to facilitate determining the status of the analysis object based on the indicator measurement value. Among them, the indicator measurement value refers to the data used to quantify the indicator. In some examples, the concepts of indicator measurement value and indicator can be interchangeable.
[0034] The SLO detector at each level is used to collect target time series data and use the collected target time series data as the input of the trend analyzer. Specifically, the target time series data collected by the SLO detector at each level includes two parts: the first part is the time series data formed by the indicator measurement value of the detection object corresponding to the current level, and the second part is the target time series data collected by the SLO detector at the next level.
[0035] The trend analyzer is used to analyze the received target time series data. If the target time series data deviates from the preset pre-made range, it can identify the abnormal scenario and match the abnormal pattern. If the problem identified by the abnormal pattern is a known problem, the known problem is output to the causal analyzer. A known problem may refer to a problem stored in a problem library, or in other words, a known problem is a problem that has been trained. If the trend analyzer fails to identify the problem, the problem can be determined as an unknown problem, and the unknown problem can be handled by manual intervention or algorithmic analysis. For example, you can first use algorithmic analysis to handle the unknown problem. If it still cannot be solved, you can apply for manual intervention to solve it.
[0036] The cause-effect analyzer is used to infer the cause of the problem based on the problem output by the trend analyzer, and then trigger the switch processor to execute the corresponding switch plan.
[0037] Optionally, if the trend analyzer cannot handle the problem according to the existing question library or learning algorithm, an alarm message may be sent through the abnormal alarm, and the alarm message is used to indicate that the abnormal problem should be handled by manual intervention.
[0038] Next, we will introduce the construction process and training method of the above model in detail with the attached figures.
[0039] The construction of the microservice operation and maintenance model can be divided into three parts: SLO decomposition, switch setting, and fault reasoning. The main functions of each part are described as follows.
[0040] SLO decomposition part: Decompose the total SLO into SLOs at each level, and determine a reasonable SLO for each level to try to solve the drawbacks of empirical protection thresholds. The above multiple levels can refer to different hierarchical architectures in a computer system. As an example, the multiple levels from top to bottom may include business layer, application layer, operating system layer, hardware layer, etc.
[0041] Switch setting section: Set up switch plans to implement effective response measures for objects that exceed the protection threshold.
[0042] Fault reasoning part: Introduce analysis components for potential problems in the model, predict potential problems and risks, and at the same time precipitate the protection effect of the switch plan to form an effective fault response set, partially replacing the work of operation and maintenance personnel.
[0043] After completing the construction of the three parts, the model can follow the iteration and evolution of the application by dynamically replacing the basic modules. The following describes the construction process of each part.
[0044] Part 1: SLO Decomposition
[0045] The main function of the SLO decomposition part is to decompose the SLO into multiple levels and determine the SLO of each level. The SLO of each level can be determined based on the SLO of the previous level. The SLO of the current level cannot be determined by the current information, but needs to be determined based on the information input of the previous level. The SLO of the highest level is the overall application goal, which exists outside the system. At the same time, SLO is also used as an input parameter of the SLO detector module of this level.
[0046] The principle of SLO decomposition is to identify the related indicators that affect the current indicators, and to trigger the relevant switch plan in time when there are potential problems or existing problems in the related indicators, to stop the loss or inform the relevant personnel to intervene. Among them, this model can determine the SLO of each level according to the correlation and causal relationship between indicators. For example, the correlation between indicators can be used as the analysis entry point and gradually upgraded to the causal relationship. Without loss of generality, the relationship between objects can be divided into three types: positive correlation, negative correlation and irrelevant. This section identifies the changes in the metric values of the indicators at the current level and the changes in the metric values of the indicators at the next level, and performs statistical analysis to obtain indicators with positive and negative correlation as the input parameters of the SLO detector.
[0047] For example, assuming that the SLO has been decomposed to the current level, the SLO of each indicator of the next level needs to be determined. The specific process of determining the SLO of the next level will be described below. Determining the SLO includes three stages: identifying related indicators, contributing model, and determining the SLO.
[0048] A. Identification of related indicators
[0049] Assume that the indicators of the current level are M and N, and the SLO has been determined. It is necessary to determine the indicators associated with the next level and determine the correlation. Assume that the indicators identified for the next level are X, Y, Z, and Q.
[0050] One possible way to identify the correlation relationship is to simulate external scenarios, observe the fluctuations of indicators M, N and X, Y, Z, and use data statistical tools to determine the relationship between indicators M, N and X, Y, Z, Q. This relationship can be for a specific scenario or for statistical results in all scenarios.
[0051] One possible way to identify association relationships is to analyze the cause-effect relationship between indicators based on application experience or theoretical basis to determine the correlation.
[0052] Figure 2 Schematic diagram of the correlation relationship between indicators of SLO levels in one embodiment of the present application. Figure 2 As shown, without loss of generality, X has a positive correlation with M, Y has a negative correlation with M, Y has a negative correlation with N, Z has a positive correlation with N, and Q has no correlation with M or N.
[0053] It should be noted that the current level and the next level may be located within the same microservice or between different microservices, depending on the object and level of indicator decomposition, and the embodiments of the present application do not limit this.
[0054] B. Contribution Model
[0055] After establishing the lower-level indicators associated with the current indicator, it is necessary to further quantify the degree to which the current indicator is affected by the lower-level indicators. This can be expressed by the following formula:
[0056]
[0057] Among them, N represents the level of the current indicator, which can also be called the current layer, and the N+1 layer represents the next level of the current layer; represents the SLO of the ith indicator at the Nth layer. Similarly, represents the SLO of the jth indicator in the N+1th layer; m represents the number of all indicator types in the Nth layer; a i express right The contribution of i The value is in the range of real numbers.
[0058] Optionally, according to formula (1), the SLO of all indicators of the N+1th layer can be expressed as follows:
[0059] SLO N+1 =A*SLO N ; (2)
[0060] in, SLO N+1 represents the SLO of all indicators in the N+1th layer, K represents the number of all indicator types in the N+1th layer, and m represents the number of all indicator types in the Nth layer.
[0061] This contribution model is used to explain the impact of the lower-level indicators on the current indicators, providing a decision basis for switch settings and providing a reference for the next-level SLO.
[0062] Optionally, a possible way to determine the coefficient is to use a single factor variable method, control other factors to maintain a specific level, observe the influence of a factor in the next layer on the current layer indicator, and use the normalized correlation coefficient as the contribution coefficient.
[0063] C. Determine SLO
[0064] According to the contribution model, after determining the degree of influence of the lower-level indicators on the upper-level indicators, the lower-level indicator SLO can be expressed by the following formula:
[0065] SLO N =A -1 *SLO N+1 ; (3)
[0066] Among them, SLO N Indicates the SLO of all indicators at layer N, SLO N+1 Indicates the SLO of all indicators at layer N+1.
[0067] In some examples, for situations where the above decomposition is not satisfied, a critical value heuristic may be used to detect the SLO of the next layer when all SLOs of the current layer are satisfied.
[0068] In some examples, after the SLO is derived, it can be adjusted accordingly based on actual conditions.
[0069] It should be understood that the role of SLO decomposition is to provide a decision basis for whether to execute the subsequent switch plan, which is the foundation of the effectiveness of automated operation and maintenance. Next, we will continue to introduce the switch setting part.
[0070] Part 2: Switch Setting Part
[0071] The main function of the switchgear part is to set up switch plans, so as to implement effective response measures for objects that exceed the protection threshold. The initial setting of the switch plan requires the intervention of operation and maintenance personnel, developers, and experts to set up appropriate switch plans for different types of faults, so as to precipitate online manual operation and maintenance measures into the switch plan execution means in the model.
[0072] Figure 3FIG. 1 is a schematic flow chart of forming a switch plan according to an embodiment of the present application. Figure 3 As shown, the logic of the switch plan formation can be expressed as the following process:
[0073] S301. Sort out the possible problems at the current level;
[0074] S302, sort out solutions to problems at the current level;
[0075] S303, converting the above-mentioned solution into a switch plan.
[0076] Figure 4 is a schematic block diagram of a method for executing a switch plan according to an embodiment of the present application. Figure 4 The switch in can refer to Figure 1 The switch processor in . Figure 4 As shown in the figure, when the metric value of the indicator object is detected to be lower than the SLO, the corresponding switch plan needs to be triggered to restore the SLO. The switch plan can act on the current layer or the next layer. It should be noted that the rationality of acting on the next layer switch is that the current layer collects all the data of the next layer, and the decision stage refers to more information, and the decision accuracy is relatively high. Therefore, if the current layer SLO can be improved by executing the next layer switch, the next layer switch is executed first, thereby reducing the impact of the problem.
[0077] In the embodiment of the present application, all factors that cause SLO reduction due to failures or potential problems can be abstracted as resources, and the effect of switches on the stability of the entire system can be abstracted into three dimensions: resource redundancy guarantee, resource scheduling capability improvement, and resource utilization improvement.
[0078] In terms of specific implementation, there is a progressive relationship between the efficiency and cost of problem handling in the above three dimensions. Resource redundancy can be guaranteed by timely releasing resources through corresponding switches to ensure the execution of core logic; resource scheduling involves scheduling algorithms under different scenarios and making plan decisions, which reduces timeliness and increases costs accordingly; improving resource utilization involves improving the collaboration efficiency between functional modules within the system and improving the execution efficiency within each module.
[0079] For example, if a CPU surge causes the SLO to decrease, it indicates that there is insufficient CPU resource redundancy; if insufficient load instances cause the SLO to decrease, it indicates a problem with resource scheduling; if an interface response timeout causes the SLO to decrease, it indicates low resource utilization efficiency.
[0080] In the embodiment of the present application, constraining the switch behavior under the above three dimensions can ensure the effectiveness of the switch execution, and at the same time facilitate relevant personnel to analyze system failures and trace the root causes of the problems. Therefore, the process of switch setting can be attributed to the classification problem of switches under the above three dimensions. Optionally, the following will describe the setting method of the switch plan that may exist under the above three dimensions, wherein switches under different dimensions can be set at each SLO layer, and there is no corresponding relationship between the switch dimension and the SLO level.
[0081] A. Resource Redundancy
[0082] For problems such as insufficient resources, in order to restore the SLO of core applications, resources can be allocated to core applications by downgrading the resource usage of non-core applications to ensure the resource usage of core applications. Commonly used switch plans can be to increase the current limiting intensity of non-core applications, perform circuit breaking and downgrading of non-core applications, etc.
[0083] B. Resource Scheduling
[0084] Resource scheduling can ensure that resources can be effectively used between different applications and different system modules. The corresponding switching plans can be to trigger dynamic scaling, replace the scheduling algorithm to match the application scenario, or reduce the loss when switching resource usage, such as reducing the number of threads.
[0085] C. Resource utilization
[0086] Resource utilization rate involves the collaboration efficiency between the functional modules within the system and the execution efficiency within each module. This method can be used for solutions that have different application execution logics set up in different application scenarios. After switching to the corresponding scenario, switch to the corresponding application execution code. This method is suitable for the continuous evolution of applications, and gradually improves resource utilization during the function iteration process. Compared with the first two methods, this method has a relatively long cycle and requires more intervention from relevant personnel.
[0087] Part III: Fault Reasoning
[0088] The fault reasoning part is the link between the SLO and the switch plan. It makes decisions based on the degree to which different indicators fail to reach the SLO to trigger the corresponding switch plan. After the model of the fault reasoning part is built, the automated operation and maintenance of microservices can be realized.
[0089] Among them, the fault reasoning part needs to count massive historical operation and maintenance data to simulate expert reasoning behavior. The core is to discover problems or potential risks in the system and infer the executableness of switches that can be used to deal with the problems.
[0090] refer to Figure 1As shown in the figure, the fault reasoning part mainly includes three components: trend analyzer, causal inference and abnormal alarm. Among them, the trend analyzer is used to analyze the time series of indicator measurement values to find potential problems or known problems. The causal inference is used to infer the cause of the problem based on the problem identified by the trend analyzer, and then trigger the corresponding switch plan. The abnormal alarm is used to alarm through the abnormal alarm when the trend analyzer identifies a problem that cannot be handled by the existing problem library or learning algorithm, and output the problem to the outside for relevant personnel to intervene and handle. The working principles of the above three components are introduced below.
[0091] A. Trend Analyzer
[0092] The trend analyzer can be used to receive and analyze the time series data of the indicator measurement values, identify abnormal scenarios, and obtain abnormal patterns that match the abnormal scenarios. In terms of abnormal pattern matching, this model believes that the system has inherent characteristics. Under certain conditions, abnormal behaviors will recur and manifest themselves in regular changes in indicator measurement values. Therefore, the process of trend analysis can be transformed into the identification and matching of abnormal patterns. The accuracy of trend analysis lies in the accuracy of abnormal pattern recognition, or in other words, in the number of abnormal patterns that have been identified. The more abnormal patterns that have been identified, the higher the accuracy of trend analysis and the higher the efficiency.
[0093] Figure 5 Schematic diagram of the working principle of the trend analyzer of one embodiment of the present application. Figure 5 As shown, the process of trend analysis can be broken down into:
[0094] 1. If the trend analysis identifies a known problem by matching the abnormal pattern, the corresponding switch plan is executed;
[0095] 2. If the trend analysis cannot match an abnormal pattern, it is necessary to further determine whether it is an unknown problem. If the system fails immediately after the unmatched indicator timing pattern occurs, this indicator timing pattern can be defined as an abnormal pattern, and the unknown problem can be handled by manual intervention or algorithm analysis. Otherwise, it is considered that this pattern will not affect the system and the analysis ends.
[0096] 3. After identifying unknown problems, set them as known problems, set up emergency plans for such anomalies, and enter the abnormal patterns to provide abnormal pattern matching basis for subsequent trend analysis;
[0097] It can be understood that the core of trend analysis lies in the identification of abnormal patterns and matching them with problems.
[0098] In some examples, one implementation of a trend analyzer may be to use artificial intelligence, machine learning algorithms to learn problem patterns.
[0099] In some examples, one implementation of the trend analyzer may be that after relevant personnel identify a problem, they enter the corresponding abnormal pattern and problem.
[0100] B. Causal Inference
[0101] After the trend analyzer has identified the problem, it is necessary to analyze the specific cause so that the right remedy can be found, the corresponding switch can be executed, and the fault can be repaired.
[0102] In some examples, one possible implementation of a causal inference device is to use a Bayesian posterior probability approach to perform root cause inference.
[0103] In some examples, one possible implementation of a causal inference engine is to form a knowledge graph based on the experience of relevant personnel, and then match the root cause in the knowledge graph when a problem occurs.
[0104] C Abnormal alarm
[0105] After the trend analyzer identifies a problem, if it can be handled automatically, there is no need to notify the maintenance personnel, or the maintenance personnel can be selectively notified of the identified problem and the solution; if it cannot be handled automatically, the relevant maintenance personnel need to be notified through the abnormal alarm to intervene.
[0106] The structure, principle and construction process of the microservice operation and maintenance model of the embodiment of the present application are introduced above. Next, the method of microservice operation and maintenance of the embodiment of the present application will be introduced in conjunction with the accompanying drawings. The method can be operated based on the microservice operation and maintenance model.
[0107] Figure 6 It is a flowchart of a method for microservice operation and maintenance according to an embodiment of the present application. Figure 6 The method can be based on Figure 1 The microservice operation and maintenance model is implemented. Figure 6 As shown, the method includes the following steps.
[0108] S601, the SLO detector collects and outputs the target time series data of the current level, wherein the target time series data of the current level includes the time series data formed by the indicator measurement value of the detection object corresponding to the current level and the target time series data collected by the SLO detector of the next level.
[0109] S602: The trend analyzer analyzes the target time series data at the current level to identify problems.
[0110] Optionally, the trend analyzer analyzes the target time series data of the current level to identify problems, and sends the problems to the causal inference device, including: the trend analyzer analyzes the target time series data to identify abnormal scenarios; when the trend analyzer identifies abnormal scenarios, it determines whether there is an abnormal pattern that matches the abnormal scenario; when there is an abnormal pattern that matches the abnormal scenario, the trend analyzer determines a known problem corresponding to the abnormal pattern, and sends the known problem to the causal inference device.
[0111] Part S602 also includes: the trend analyzer determines whether the system fails due to the time series data of the indicator measurement value corresponding to the abnormal scenario when there is no abnormal pattern matching the abnormal scenario; if the system fails, the trend analyzer determines the abnormal pattern and matches it with an unknown problem, wherein the unknown problem is handled by manual intervention or algorithm analysis; if the system does not fail, the trend analyzer ends this analysis.
[0112] S603: The causal inference unit receives the problem identified by the trend analyzer, infers the cause of the problem, and determines a switch plan corresponding to processing the problem.
[0113] S604: The switch processor receives information indicating the switch plan and executes the switch plan, where the switch plan is a strategy set to ensure stable operation of the application.
[0114] In the embodiment of the present application, this model combines SLO and switch plans to hierarchically manage various indicators in the operation and maintenance process, improve information density and value layer by layer, and after analyzing the cause of the problem, trigger the corresponding switch to eliminate risks or faults, forming a closed loop between SLO and switch plans, continuously precipitating risk and fault response measures, improving the model's automated operation and maintenance capabilities, and opening an alarm interface to facilitate relevant personnel to understand the system situation. This model can realize self-nesting of patterns at all levels, which is convenient for model expansion and dynamic evolution. This model is suitable for microservice operation and maintenance scenarios, which can greatly improve operation and maintenance efficiency and reduce operation and maintenance costs.
[0115] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0116] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0118] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0119] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0120] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0121] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0122] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A system for microservice operation and maintenance, It is characterized in that The system comprises multiple layers, each layer comprising: A service level objective (SLO) detector is used to collect and output the target time series data of the current level, wherein the target time series data of the current level includes the time series data formed by the indicator measurement value of the detection object corresponding to the current level and the target time series data collected by the SLO detector of the next level; wherein each level corresponds to an SLO, and the SLO of each level is determined according to the SLO of the previous level; Trend Analyzer, which is used to analyze the target time series data of the current level to identify problems; A causal inference device, configured to receive the problem identified by the trend analyzer, infer the cause of the problem, and determine a switch plan corresponding to handling the problem; The switch processor is used to receive information indicating the switch plan and execute the switch plan, wherein the switch plan is a strategy set to ensure stable operation of the application; wherein the switch plan is used for the current level or the next level.
2. The system according to claim 1, It is characterized in that In terms of the trend analyzer analyzing the target time series data of the current level to identify problems and sending the problems to the causal inference device, the trend analyzer is specifically used to: Analyze the target time series data to identify abnormal scenarios; In the case of identifying an abnormal scenario, determining whether there is an abnormal pattern matching the abnormal scenario; In the case that there is an abnormal pattern matching the abnormal scenario, a known problem corresponding to the abnormal pattern is determined, and the known problem is sent to the causal inference device.
3. The system according to claim 2, It is characterized in that The trend analyzer is also used to: In the absence of an abnormal pattern matching the abnormal scenario, determining whether the system fails due to time series data of the indicator measurement value corresponding to the abnormal scenario; If the system fails, an abnormal pattern is determined and an unknown problem is matched to it, wherein the unknown problem is handled by manual intervention or algorithm analysis; If the system does not fail, the analysis ends.
4. The system according to claim 1 or 2, It is characterized in that The system further comprises: The abnormal alarm is used to send an alarm message when the trend analyzer cannot handle the problem through the existing question library or learning algorithm. The alarm message is used to indicate manual intervention to handle the problem.
5. A method for microservice operation and maintenance, It is characterized in that include: The service level objective SLO detector collects and outputs the target time series data of the current level, wherein the target time series data of the current level includes the time series data formed by the indicator measurement value of the detection object corresponding to the current level and the target time series data collected by the SLO detector of the next level; wherein each level corresponds to the SLO, and the SLO of each level is determined according to the SLO of the previous level; The trend analyzer analyzes the target time series data at the current level to identify problems; The causal inference device receives the problem identified by the trend analyzer, infers the cause of the problem, and determines a switch plan corresponding to processing the problem; The switch processor receives information indicating the switch plan and executes the switch plan, wherein the switch plan is a strategy set to ensure stable operation of the application; wherein the switch plan is used for the current level or the next level.
6. The method according to claim 5, It is characterized in that The trend analyzer analyzes the target time series data of the current level to identify problems and sends the problems to the causal inference device, including: The trend analyzer performs analysis based on the target time series data to identify abnormal scenarios; The trend analyzer, when identifying an abnormal scenario, determines whether there is an abnormal pattern matching the abnormal scenario; When there is an abnormal pattern matching the abnormal scenario, the trend analyzer determines a known problem corresponding to the abnormal pattern and sends the known problem to the causal inference device.
7. The method according to claim 6, It is characterized in that The method further comprises: The trend analyzer determines whether the system fails due to the time series data of the indicator measurement value corresponding to the abnormal scenario when there is no abnormal pattern matching the abnormal scenario; If a system failure occurs, the trend analyzer determines the abnormal pattern and matches it with an unknown problem, wherein the unknown problem is handled by manual intervention or algorithmic analysis; If no failure occurs in the system, the trend analyzer ends the analysis.
8. The method according to claim 5 or 6, It is characterized in that The method further comprises: When the trend analyzer cannot process a problem through an existing problem library or a learning algorithm, the abnormal alarm sends an alarm message, and the alarm message is used to indicate manual intervention to process the problem.
Citation Information
Patent Citations
A method and system for determining fault of heterogeneous system based on machine learning
CN111209131A
Cloud native system-oriented micro-service root cause positioning method
CN113014421A