Fault management method, device and equipment of service orchestration system and storage medium
By introducing machine learning and decision tree algorithms into the anomaly detection model in the service orchestration system, the problem of inaccurate fault monitoring is solved, efficient fault detection and automated recovery are achieved, and the stability and reliability of the system are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-17
- Publication Date
- 2026-03-24
AI Technical Summary
Inaccurate fault monitoring in existing technologies, along with false alarms and missed alarms, leads to insufficient stability and reliability of service orchestration in financial business systems.
We employ machine learning and anomaly detection algorithms, construct an anomaly detection model using decision tree algorithm, and use historical and current system data from the service orchestration system for feature selection and training to identify potential faults and perform accurate detection.
It improves the accuracy of fault detection, reduces false alarms and false negatives, ensures the stability and reliability of the service orchestration system, supports automated fault recovery and self-healing, and shortens fault repair time.
Smart Images

Figure CN117313012B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of financial technology and internet technology, and in particular to a fault management method, apparatus, device and storage medium for a service orchestration system. Background Technology
[0002] Self-healing and automatic repair of service orchestration faults in financial business systems is an important solution technology in the field of information technology. By using automated mechanisms and strategies, it is possible to identify faults and take appropriate measures to recover automatically.
[0003] Currently, there are some fault self-healing products and technologies in the industry, such as Kubernetes container management tools, which mainly provide scalable container deployment and management capabilities, ensuring high availability of applications through monitoring and self-healing functions. However, existing fault self-healing products generally suffer from inaccurate fault monitoring, false alarms, and missed alarms, failing to guarantee the stability and reliability of the system. Summary of the Invention
[0004] The main objective of this application is to provide a fault management method, apparatus, device, and storage medium for a service orchestration system, which can solve the technical problems of inaccurate fault monitoring, false alarms, and missed alarms in the prior art.
[0005] To achieve the above objectives, the first aspect of this application provides a fault management method for a service orchestration system, the method comprising:
[0006] Acquire historical system data from the service orchestration system for different time periods. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data.
[0007] Feature selection is performed on historical system data, and a dataset is constructed based on the selected features;
[0008] An anomaly detection model to be trained is constructed using the dataset and trained to obtain a trained anomaly detection model. The anomaly detection model to be trained is constructed based on the decision tree algorithm.
[0009] Target features are filtered from the current system data of the acquired service orchestration system, which includes current log data, current metric data, and current operation monitoring data.
[0010] The target features are input into the trained anomaly detection model to obtain the anomaly detection results of the service orchestration system.
[0011] To achieve the above objectives, a second aspect of this application provides a fault management apparatus for a service orchestration system, the apparatus comprising:
[0012] The first data acquisition module is used to acquire historical system data of the service orchestration system at different time periods. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data.
[0013] The dataset construction module is used to select features from historical system data and construct datasets based on the selected features.
[0014] The training module is used to build and train an anomaly detection model using a dataset to obtain a trained anomaly detection model. The anomaly detection model to be trained is built based on the decision tree algorithm.
[0015] The second data acquisition module is used to filter target features from the current system data of the acquired service orchestration system, wherein the current system data includes current log data, current indicator data and current operation monitoring data;
[0016] The anomaly detection module is used to input target features into a trained anomaly detection model to obtain anomaly detection results for the service orchestration system.
[0017] To achieve the above objectives, a third aspect of this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the following steps:
[0018] Acquire historical system data from the service orchestration system for different time periods. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data.
[0019] Feature selection is performed on historical system data, and a dataset is constructed based on the selected features;
[0020] An anomaly detection model to be trained is constructed using the dataset and trained to obtain a trained anomaly detection model. The anomaly detection model to be trained is constructed based on the decision tree algorithm.
[0021] Target features are filtered from the current system data of the acquired service orchestration system, which includes current log data, current metric data, and current operation monitoring data.
[0022] The target features are input into the trained anomaly detection model to obtain the anomaly detection results of the service orchestration system.
[0023] To achieve the above objectives, a fourth aspect of this application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:
[0024] Acquire historical system data from the service orchestration system for different time periods. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data.
[0025] Feature selection is performed on historical system data, and a dataset is constructed based on the selected features;
[0026] An anomaly detection model to be trained is constructed using the dataset and trained to obtain a trained anomaly detection model. The anomaly detection model to be trained is constructed based on the decision tree algorithm.
[0027] Target features are filtered from the current system data of the acquired service orchestration system, which includes current log data, current metric data, and current operation monitoring data.
[0028] The target features are input into the trained anomaly detection model to obtain the anomaly detection results of the service orchestration system.
[0029] The embodiments of this application have the following beneficial effects:
[0030] This application supports intelligent fault detection, improving existing fault detection mechanisms by introducing machine learning and anomaly detection algorithms. Through analysis of large amounts of logs, indicator data, and behavioral patterns, it can automatically identify anomalies and more accurately detect and identify potential faults. This reduces false alarms and false negatives, ensuring the stability and reliability of financial business system service orchestration. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] in:
[0033] Figure 1 This is a diagram illustrating the application environment of the fault management method for the service orchestration system in this application embodiment;
[0034] Figure 2 This is a flowchart of the fault management method of the service orchestration system in the embodiments of this application;
[0035] Figure 3 This is a structural block diagram of the fault management device of the service orchestration system in the embodiments of this application;
[0036] Figure 4 This is a structural block diagram of the computer device in the embodiments of this application. Detailed Implementation
[0037] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0038] Figure 1 This is a diagram illustrating the application environment of a fault management method for a service orchestration system in one embodiment. (Refer to...) Figure 1The fault management method of this service orchestration system is applied to the fault management system of the service orchestration system. The fault management system of this service orchestration system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, tablet computer, or laptop computer. The server 120 can be implemented using a standalone server or a server cluster consisting of multiple servers. Terminal 110 is used to send various configuration data of the fault management system of the service orchestration system to server 120. Server 120 is used to obtain historical system data of the service orchestration system at different time periods. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data. Features are selected from the historical system data, and a dataset is constructed based on the selected features. An anomaly detection model to be trained is constructed using the dataset and trained to obtain a trained anomaly detection model. The anomaly detection model to be trained is constructed based on a decision tree algorithm. Target features are filtered from the current system data of the service orchestration system, which includes current log data, current indicator data, and current operation monitoring data. The target features are input into the trained anomaly detection model to obtain the anomaly detection results of the service orchestration system.
[0039] like Figure 2 As shown, in one embodiment, a fault management method for a service orchestration system is provided, which specifically includes the following steps:
[0040] S100: Obtain historical system data from the service orchestration system for different time periods. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data.
[0041] Specifically, service orchestration systems are an important technology in the low-code field. Compared with traditional IT platforms, low-code platforms are closer to the actual development process, providing developers with a visual development environment, reducing or eliminating the need for native code writing in application development, and thus providing a solution for convenient application building.
[0042] Common failures in service orchestration systems include: high concurrency, large volume of requests, network congestion, downtime, and unavailability of interface request services.
[0043] The service orchestration system may experience anomalies or normal operation at different times in the past, and the anomalies that occur at different times may also be different. Each time period corresponds to a historical system data set, which includes some of the raw data collected, as well as processed data that has been processed and statistically analyzed from the raw data.
[0044] Historical network log data includes at least one of historical network traffic data, historical connection status data, and historical access log data; historical system log data includes at least one of historical system operation data, historical system error information data, and historical system warning data; historical operation log data includes at least one of system operation status data, system operation error data, system operation anomaly data, and system operation warning data.
[0045] Historical system processing capacity metrics include various indicators such as historical request error rate, historical request success rate, historical exception handling count, historical access volume, and historical transaction success rate, which are used to indicate the system's stability and error handling capabilities.
[0046] Historical system response capability metrics include response time metrics, such as historical request timeout rate, number of historical request timeouts, and average historical request response time.
[0047] Historical operation monitoring data includes task execution details, such as error messages and warnings for task failures (e.g., file read / write errors) or task timeouts, the number of abnormal events that occurred within a certain time period, task error rate, network traffic data, and resource utilization such as CPU utilization and memory utilization.
[0048] Historical system data is obtained by cleaning and transforming raw data. The purpose of cleaning is to preprocess the raw data, repair and correct errors and missing data, resolve inconsistencies in data format, and ensure the accuracy and integrity of the data.
[0049] S200: Select features from historical system data and construct a dataset based on the selected features.
[0050] Specifically, machine learning algorithms such as decision trees can be used to evaluate the importance of each feature in historical system data and select important features. Alternatively, genetic algorithms can be used to select appropriate features from historical system data. In this case, each parameter in the historical system data represents a feature.
[0051] In addition, appropriate features can be selected according to specific business needs, and this application does not impose any restrictions on this.
[0052] The selected features are data from historical system data that are highly correlated with system failures. By selecting features, the influence of irrelevant factors can be eliminated, thereby improving the accuracy of the model in fault identification and detection.
[0053] S300: Construct an anomaly detection model to be trained using the dataset and train it to obtain a trained anomaly detection model. The anomaly detection model to be trained is constructed based on the decision tree algorithm.
[0054] Specifically, the decision tree algorithm is a method for approximating discrete function values. It is a typical classification method that first processes the data, using inductive algorithms to generate readable rules and a decision tree, and then uses these rules to analyze new data. Essentially, a decision tree is a process of classifying data through a series of rules.
[0055] Decision tree algorithms construct decision trees to discover classification rules inherent in data. The core of the decision tree algorithm lies in constructing decision trees with high accuracy and small size. Decision tree construction can be divided into two steps. The first step is decision tree generation: the process of generating a decision tree from a training sample set. Generally, the training sample dataset is a historical dataset with a certain degree of aggregation, used for data analysis and processing, based on actual needs. The second step is decision tree pruning: decision tree pruning is the process of verifying, correcting, and removing the decision tree generated in the previous stage. It mainly involves using data from a new sample dataset (called the test dataset) to validate the initial rules generated during the decision tree generation process, and pruning branches that affect the accuracy of the prediction.
[0056] This embodiment uses a dataset to build and train an anomaly detection model, and optimizes the model parameters to obtain a trained anomaly detection model.
[0057] S400: Filter target features from the current system data of the acquired service orchestration system, where the current system data includes current log data, current metric data, and current operation monitoring data.
[0058] Specifically, the current system data refers to system data collected within a certain time period. For example, an anomaly check is performed on the service orchestration system every preset time interval. The current system data used in each anomaly check is the system data statistically collected within the preset time period in the past.
[0059] S500: Input the target features into the trained anomaly detection model to obtain the anomaly detection results of the service orchestration system.
[0060] Specifically, the anomaly detection results are used to indicate whether there is an anomaly in the current service orchestration system, such as whether the current service orchestration system has an anomaly or not.
[0061] Alternatively, the anomaly detection results can be used to indicate whether there are anomalies in the current service orchestration system and their classification. For example, fault types or anomaly classifications include: application errors, network problems, hardware problems, etc. This application does not limit the scope of these classifications. Application errors include program logic errors, algorithm errors, configuration errors, etc.
[0062] This embodiment supports intelligent fault detection, improving existing fault detection mechanisms by introducing machine learning and anomaly detection algorithms. By analyzing large amounts of logs, indicator data, and behavioral patterns, it can automatically identify anomalies and more accurately detect and identify potential faults. This reduces false alarms and false negatives, ensuring the stability and reliability of financial business system service orchestration.
[0063] In one embodiment, the method further includes:
[0064] If the anomaly detection result indicates that an anomaly has occurred in the service orchestration system, then key indicator data will be obtained from the current indicator data and the current operation monitoring data.
[0065] Based on each key indicator data and the corresponding fault recovery triggering rules, determine whether to trigger an alarm and fault recovery;
[0066] If an alarm and fault recovery are triggered, the corresponding fault recovery operation will be performed.
[0067] Specifically, key performance indicators include at least one of the following: request error rate or request success rate, number of exceptions handled, request timeout rate or number of timeouts, average request response time, resource utilization, and transaction success rate.
[0068] Each key performance indicator (KPI) corresponds to a fault recovery trigger rule. If, based on at least one KPI and its corresponding fault recovery trigger rule, a corresponding alarm and fault recovery are determined to have been triggered, then a fault recovery operation is performed on the service orchestration system. Alternatively, if, based on at least a preset number of KPIs and their corresponding fault recovery trigger rules, a corresponding alarm and fault recovery are determined to have been triggered, then a fault recovery operation is performed on the service orchestration system.
[0069] The fault recovery triggering rules are configured with corresponding thresholds, with different key metrics corresponding to different thresholds. More specifically, if at least one of the following conditions is met, or at least a preset number of conditions are met: request error rate exceeds the first threshold (or request success rate is less than the second threshold), number of exception handling exceeds the third threshold, request timeout rate exceeds the fourth threshold (or number of request timeouts exceeds the eighth threshold), average request response time exceeds the fifth threshold, resource utilization does not exceed the sixth threshold, and transaction success rate does not exceed the seventh threshold, then an alarm and fault recovery for the service orchestration system are triggered.
[0070] For example, theoretically, the average response time of a system request should not exceed 3 seconds. If the number of times a request response time exceeds 3 seconds is greater than 100 times within 1 minute, an alarm and fault recovery will be triggered for the service orchestration system.
[0071] Of course, the above is merely an illustrative example. The specific threshold should be configured according to the actual scenario. This application does not limit the selection of the threshold.
[0072] In this embodiment, when an anomaly is determined to occur in the service orchestration system, it further determines whether to trigger an alarm and fault recovery for the service orchestration system. While further confirming the anomaly, if the anomaly is determined to be serious, an alarm and fault recovery are triggered, which effectively prevents misjudgment and enables timely fault recovery and troubleshooting of the service orchestration system, thus achieving system self-healing.
[0073] In one embodiment, performing the corresponding fault recovery operation includes:
[0074] Based on the current log data, current metric data, and current operational monitoring data, determine the cause of the current failure;
[0075] Match the fault recovery strategy corresponding to the fault cause, and execute the corresponding fault recovery operation according to the matched fault recovery strategy.
[0076] Specifically, the cause of the current fault can be identified by comparing the data during the fault period with the data during normal operation to find the differences.
[0077] Additionally, the fault type or anomaly classification corresponding to each set of anomaly samples can be provided during the training of the anomaly detection model. The anomaly detection model can learn the system data distribution patterns corresponding to different fault types during training. Therefore, during the detection phase, if the anomaly detection result output by the trained anomaly detection model indicates an anomaly in the service orchestration system, it will also specify the type of anomaly. These fault types include: application errors such as program logic errors, algorithm errors, and configuration errors; as well as network problems and hardware errors.
[0078] The cause of a fault can be basically determined based on the type of fault.
[0079] Alternatively, the final cause of the failure can be determined by combining the above data comparison methods and model detection methods.
[0080] Different causes of failures have corresponding failure recovery strategies, and different failure recovery strategies execute corresponding failure recovery operations.
[0081] This embodiment supports a predictable self-healing strategy. For complex service orchestration scenarios, it first uses historical data and establishes a fault detection model to predict potential faults, including the possible types and extent of faults, and then formulates corresponding self-healing strategies. By taking appropriate preventative measures before faults occur, the stability and reliability of the system are improved.
[0082] This embodiment supports automated fault analysis, introducing automated fault analysis methods. Through root cause analysis and other techniques, combined with monitoring data and log analysis, it quickly locates the cause of the fault, finds the cause to match the corresponding fault recovery strategy, and performs fault recovery operations on the service orchestration system according to the matched fault recovery strategy. This allows for timely, rapid, targeted, and accurate fault recovery of the service orchestration system, shortening fault repair time, preventing blind recovery from backfiring, and effectively ensuring the normal operation of the service orchestration system.
[0083] In one embodiment, a fault recovery strategy corresponding to the cause of the fault is matched, and the corresponding fault recovery operation is executed according to the matched fault recovery strategy, including:
[0084] If the failure is caused by a service crash or a stop responding, restart at least one of the relevant services, processes, and containers in the service orchestration system.
[0085] If the cause of the problem is a network connectivity issue, the system will automatically detect the network and attempt to reconnect and / or adjust network settings.
[0086] If the failure is due to a configuration error or failure, then perform one of the following: automatically roll back the configuration, reload the correct configuration, or restore to the backup configuration.
[0087] Specifically, the appropriate recovery strategy is selected based on the type and nature of the failure. For example, if a service crashes or becomes unresponsive, the service orchestration system services and / or related processes and / or containers are automatically restarted. If there is a network outage or connectivity problem, the network outage is automatically detected and reconnection is attempted, or network settings or configurations are adjusted. If there is a configuration error or failure, the configuration is automatically rolled back, the correct configuration is reloaded, or the system is restored to a backup configuration.
[0088] In this context, "rollback configuration" refers to rolling back the service orchestration system from its current configuration to a target version configuration. The target version configuration is an existing configuration, such as the default version configuration or the previous version configuration. Of course, the target version configuration can be set according to actual needs, and this application does not impose any restrictions on it.
[0089] Reloading the correct configuration means reloading the current configuration or reloading another version of the certified correct configuration.
[0090] Restoring to a backup configuration means switching the service orchestration system from its current configuration to a backup configuration. A backup configuration is a certified correct configuration among existing configurations. While the configuration level of a backup configuration may be lower than the current configuration, it guarantees most of the functionality of the service orchestration system and is less prone to errors.
[0091] This embodiment activates corresponding fault repair strategies based on different fault causes, which can effectively and specifically restore the performance of the service orchestration system.
[0092] This application extends the design by focusing on key aspects such as intelligent fault detection, automated fault troubleshooting, and adaptive fault recovery. For complex service orchestration scenarios, it uses historical data and fault models to predict potential faults and takes corresponding measures to improve the accuracy, reliability, and self-healing capabilities of fault detection.
[0093] In one embodiment, a fault recovery strategy corresponding to the cause of the fault is matched, and the corresponding fault recovery operation is executed according to the matched fault recovery strategy, including:
[0094] If the cause of the failure is a primary resource failure or anomaly, the system will switch to a backup resource, which includes at least one of the following: server resources, network devices, applications, and databases.
[0095] Specifically, in addition to software failures, service orchestration systems may also experience physical failures. When a major resource failure or anomaly is detected, scripts and automation tools can be used for operation and management to automatically trigger fault recovery strategies and switch the service orchestration system from the currently used resource to a backup resource.
[0096] Backup resources refer to alternative resources such as primary server resources, network equipment resources, application resources, and database resources.
[0097] This embodiment implements fault tolerance and redundancy mechanisms to ensure that a quick switch to backup resources can be made when resource failures occur, thus guaranteeing the normal operation of the service orchestration system.
[0098] In one embodiment, performing corresponding fault recovery operations according to the matched fault recovery strategy includes:
[0099] Determine the scope of the current fault's impact based on the current log data and current operational monitoring data;
[0100] Determine the priority level of the current fault based on its impact range;
[0101] Based on the current fault priority, the corresponding fault recovery operation is encapsulated as a task and published to the target location in the fault handling queue for execution. The target location is determined based on the current fault priority and the priority of existing tasks in the fault handling queue.
[0102] Specifically, based on current operational monitoring data, system performance and service availability are analyzed to determine the scope of the first failure's impact; simultaneously, log data is used to identify affected components or services, further determining the scope of the second failure's impact. The total impact scope of the current failure is then calculated based on the first and second failure impact scopes.
[0103] The larger the scope of a fault's impact and the higher its severity, the higher its priority level, and the more urgent the fault needs to be addressed. Based on this, this embodiment determines the priority level of the current fault according to its scope of impact.
[0104] Service orchestration systems may experience various failures at different times. New failures may occur before historical failures have been resolved. The severity of different failures varies, so their processing priorities are also different.
[0105] The priority level of a current fault can be determined based on a pre-stored mapping table between fault impact range and priority level. Alternatively, it can be determined by comparing the impact range of all unresolved faults in the current service orchestration system; the larger the impact range, the higher the priority level. Unresolved faults include historical unresolved faults that are yet to be resolved, as well as newly generated current faults.
[0106] Of course, the priority level of a fault can also be assessed based on both the scope of its impact and the importance of that impact. More specifically, a fault score is obtained by quantifying and weighting the scope of the fault's impact and its importance. The higher the fault score, the higher the priority level.
[0107] To ensure that the service orchestration system can maintain normal operation of most functions even in the event of a small-scale failure, some threads can be used to resolve the failures one by one. This allows some thread resources to be reserved for handling the normal business and services of the service orchestration system. Based on this, this embodiment uses a fault handling queue to cache pending fault recovery tasks.
[0108] Existing fault recovery tasks are queued in the fault handling queue according to their priority. These tasks wait for thread processing in sequence. Newly generated fault recovery tasks can be inserted into any target position in the fault handling queue based on their priority. In other words, the order in the fault handling queue is dynamically adjusted. When a new high-priority fault appears in the queue, the queue is dynamically adjusted, and the scheduling order is automatically adjusted.
[0109] In addition, it can support adaptive fault recovery strategies, which can dynamically adjust fault recovery strategies according to the current status of the service orchestration system and environmental conditions. Different recovery strategies can be adopted during high load periods to ensure system performance, while during non-critical periods, the focus can be on resource utilization to perform recovery.
[0110] More specifically, the number of threads for handling fault recovery tasks is allocated based on the availability of resources in the service orchestration system and the current load, in order to determine whether to use single-threaded processing or multi-threaded parallel processing.
[0111] In addition, to ensure the correctness and reliability of the solution, the function of scheduling fault recovery operation tasks using a fault handling queue can be pre-tested and verified automatically before being put into the production environment.
[0112] This embodiment introduces priority management and intelligent scheduling. By analyzing the impact range and severity of faults, it makes decisions on priority management and fault recovery. Based on priority management, it manages the fault processing queue and performs intelligent scheduling to handle different faults sequentially, while ensuring that critical tasks and services receive priority processing, avoiding chain failures and cascading failures. This helps the service orchestration system better allocate resources and handle the most critical and urgent issues when faults occur.
[0113] In one embodiment, the method further includes:
[0114] Generate and display a real-time monitoring report of the service orchestration system based on current system data;
[0115] And / or,
[0116] The system provides a visual representation of the recovery progress of fault recovery operations.
[0117] Specifically, the current system data is updated every preset time interval. Therefore, the real-time monitoring report is also updated every preset time interval. The preset time interval can be, for example, 1 minute, 5 minutes, 10 minutes, etc. This application does not limit this.
[0118] The progress of fault recovery operations can be represented by the degree of recovery of the affected area.
[0119] The system is designed with visual monitoring and reporting functions to visualize monitoring data, including fault status, recovery progress, and performance indicators. This helps operations and maintenance administrators quickly identify faults and understand the operational status of the service orchestration system.
[0120] This application also supports custom self-healing strategies and rules, allowing users to configure and adjust self-healing solutions according to specific circumstances. It also provides a visual interface and tools to track and monitor the self-healing process of the service orchestration system. Operations administrators can view logs and execution steps in real time during the self-healing process, thereby better understanding and adjusting the self-healing mechanism. It offers good flexibility and scalability, adapting to different business needs.
[0121] This application can also support cross-system collaboration and service recovery. When a service fails, it can automatically notify the service orchestration related systems and services to make corresponding adjustments in order to maintain the normal operation of the entire system.
[0122] refer to Figure 3 This application also provides a fault management device for a service orchestration system, the device comprising:
[0123] The first data acquisition module 100 is used to acquire historical system data of the service orchestration system at different time periods. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data.
[0124] The dataset construction module 200 is used to select features from historical system data and construct a dataset based on the selected features.
[0125] The training module 300 is used to construct and train an anomaly detection model using a dataset to obtain a trained anomaly detection model. The anomaly detection model to be trained is constructed based on a decision tree algorithm.
[0126] The second data acquisition module 400 is used to filter target features from the current system data of the acquired service orchestration system, wherein the current system data includes current log data, current indicator data and current operation monitoring data;
[0127] The anomaly detection module 500 is used to input target features into the trained anomaly detection model to obtain the anomaly detection results of the service orchestration system.
[0128] In one embodiment, the device further includes:
[0129] The third data acquisition module is used to acquire key indicator data from the current indicator data and the current operation monitoring data if the anomaly detection result indicates that an anomaly has occurred in the service orchestration system.
[0130] The trigger judgment module is used to determine whether to trigger an alarm and perform fault recovery based on each key indicator data and the corresponding fault recovery trigger rules;
[0131] The fault recovery module is used to perform corresponding fault recovery operations if it is determined that an alarm has been triggered and fault recovery has occurred.
[0132] In one embodiment, the fault recovery module includes:
[0133] The fault analysis module is used to determine the cause of the current fault based on the current log data, current indicator data, and current operation monitoring data.
[0134] The strategy matching and fault recovery module is used to match fault recovery strategies corresponding to the fault causes and to execute corresponding fault recovery operations based on the matched fault recovery strategies.
[0135] In one embodiment, the policy matching and fault recovery module includes:
[0136] The first fault recovery module is used to restart at least one of the relevant services, processes, and containers of the service orchestration system if the fault is caused by a service crash or a stop responding.
[0137] The second fault recovery module is used to automatically detect the network and attempt to reconnect and / or adjust network settings if the cause of the fault is a network connection problem.
[0138] The third fault recovery module is used to perform any one of the following if the fault is caused by a configuration error or failure: automatic configuration rollback, reloading the correct configuration, or restoration to the backup configuration.
[0139] In one embodiment, the policy matching and fault recovery module includes:
[0140] The fourth fault recovery module is used to switch to backup resources if the fault is caused by a primary resource failure or abnormality. The backup resources include at least one of server resources, network devices, applications, and databases.
[0141] In one embodiment, the policy matching and fault recovery module includes:
[0142] The impact scope determination module is used to determine the impact scope of the current fault based on the current log data and the current operation monitoring data.
[0143] The priority determination module is used to determine the priority level of the current fault based on the scope of its impact.
[0144] The task publishing module is used to encapsulate the corresponding fault recovery operation as a task and publish it to the target location in the fault processing queue for execution based on the priority level of the current fault. The target location is determined based on the priority level of the current fault and the priority level of the existing tasks in the fault processing queue.
[0145] In one embodiment, the device further includes:
[0146] The first display module is used to generate and display real-time monitoring reports of the service orchestration system based on current system data;
[0147] And / or,
[0148] The second display module is used to visually demonstrate the progress of the fault recovery operation.
[0149] This application enables rapid fault detection and response, and automatically executes corresponding recovery operations. It can shorten fault recovery time, ensure high service availability and continuity, and improve troubleshooting efficiency by up to 80%. Maintenance costs and workload are reduced by 50%, decreasing reliance on manual intervention and the workload of troubleshooting and repair, allowing the team to focus more on other important tasks.
[0150] This application supports predictable self-healing strategies, taking appropriate preventative measures before failures occur to improve system stability and reliability. It also supports automated monitoring and alarms, providing real-time monitoring and alarm functions. Operations administrators can quickly receive notifications of abnormal key indicators, promptly identify and address faults, and prevent the escalation of their impact.
[0151] This application can automatically detect and recover from faults, reduce the impact of faults on the system, quickly diagnose faults and take appropriate measures to minimize service interruption time and reduce business impact, thereby improving the reliability and stability of the system.
[0152] This application supports business continuity and disaster recovery, automatically switching to backup resources in the event of a failure and ensuring the continued operation of critical business operations.
[0153] Figure 4 An internal structural diagram of a computer device in one embodiment is shown. This computer device can specifically be a terminal or a server. Figure 4As shown, the computer device includes a processor, memory, and network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program causes the processor to perform the steps in the above-described method embodiments. The internal memory may also store a computer program, which, when executed by the processor, causes the processor to perform the steps in the above-described method embodiments. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0154] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps:
[0155] Acquire historical system data from the service orchestration system for different time periods. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data.
[0156] Feature selection is performed on historical system data, and a dataset is constructed based on the selected features;
[0157] An anomaly detection model to be trained is constructed using the dataset and trained to obtain a trained anomaly detection model. The anomaly detection model to be trained is constructed based on the decision tree algorithm.
[0158] Target features are filtered from the current system data of the acquired service orchestration system, which includes current log data, current metric data, and current operation monitoring data.
[0159] The target features are input into the trained anomaly detection model to obtain the anomaly detection results of the service orchestration system.
[0160] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the following steps:
[0161] Acquire historical system data from the service orchestration system for different time periods. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data.
[0162] Feature selection is performed on historical system data, and a dataset is constructed based on the selected features;
[0163] An anomaly detection model to be trained is constructed using the dataset and trained to obtain a trained anomaly detection model. The anomaly detection model to be trained is constructed based on the decision tree algorithm.
[0164] Target features are filtered from the current system data of the acquired service orchestration system, which includes current log data, current metric data, and current operation monitoring data.
[0165] The target features are input into the trained anomaly detection model to obtain the anomaly detection results of the service orchestration system.
[0166] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0167] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0168] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A fault management method for a service orchestration system, characterized in that, The method includes: The system acquires historical system data from different time periods of the service orchestration system. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data. Feature selection is performed on the historical system data, and a dataset is constructed based on the selected features; The dataset is used to construct and train an anomaly detection model to obtain a trained anomaly detection model, wherein the anomaly detection model to be trained is constructed based on a decision tree algorithm; Target features are filtered from the current system data of the service orchestration system, wherein the current system data includes current log data, current indicator data, and current operation monitoring data; The target features are input into the trained anomaly detection model to obtain the anomaly detection results of the service orchestration system; If the anomaly detection result indicates that the service orchestration system has an anomaly, then key indicator data is obtained from the current indicator data and the current operation monitoring data. The key indicator data includes at least one of the following: request error rate or request success rate, number of anomaly handling, request timeout rate or number of request timeouts, average request response time, resource utilization, and transaction success rate. Based on each key indicator data and the corresponding fault recovery triggering rules, determine whether to trigger an alarm and fault recovery; If an alarm and fault recovery are triggered, the corresponding fault recovery operation will be performed.
2. The method according to claim 1, characterized in that, The execution of the corresponding fault recovery operation includes: Based on the current log data, current metric data, and current operational monitoring data, determine the cause of the current fault; Match the fault recovery strategy corresponding to the fault cause, and execute the corresponding fault recovery operation according to the matched fault recovery strategy.
3. The method according to claim 2, characterized in that, The process of matching a fault recovery strategy with the fault cause and executing corresponding fault recovery operations based on the matched fault recovery strategy includes: If the cause of the failure is a service crash or a stop responding, then restart at least one of the relevant services, processes, and containers of the service orchestration system; If the cause of the failure is a network connectivity problem, the network will be automatically detected and an attempt will be made to reconnect and / or adjust network settings. If the cause of the failure is a configuration error or failure, then perform one of the following: automatically roll back the configuration, reload the correct configuration, or restore to the backup configuration.
4. The method according to claim 2, characterized in that, The process of matching a fault recovery strategy with the fault cause and executing corresponding fault recovery operations based on the matched fault recovery strategy includes: If the cause of the failure is a primary resource failure or anomaly, then the system switches to a backup resource, which includes at least one of server resources, network devices, applications, and databases.
5. The method according to claim 2, characterized in that, The step of performing corresponding fault recovery operations according to the matched fault recovery strategy includes: The scope of the current fault's impact is determined based on the current log data and current operational monitoring data. The priority level of the current fault is determined based on the scope of its impact. Based on the priority level of the current fault, the corresponding fault recovery operation is encapsulated as a task and published to the target location in the fault processing queue for execution. The target location is determined based on the priority level of the current fault and the priority level of the existing tasks in the fault processing queue.
6. The method according to claim 2, characterized in that, The method further includes: Generate and display a real-time monitoring report of the service orchestration system based on the current system data; And / or, The recovery progress of the fault recovery operation is displayed visually.
7. A fault management device for a service orchestration system, characterized in that, The device includes: The first data acquisition module is used to acquire historical system data of the service orchestration system at different time periods. The historical system data includes historical log data, historical indicator data, and historical operation monitoring data. The historical log data includes at least one of historical network log data, historical system log data, and historical operation log data. The historical indicator data includes at least one of historical system processing capacity indicator data and historical system response capacity indicator data. The dataset construction module is used to select features from the historical system data and construct a dataset based on the selected features. The training module is used to construct and train an anomaly detection model to be trained using the dataset, thereby obtaining a trained anomaly detection model, wherein the anomaly detection model to be trained is constructed based on a decision tree algorithm. The second data acquisition module is used to filter target features from the acquired current system data of the service orchestration system, wherein the current system data includes current log data, current indicator data and current operation monitoring data; An anomaly detection module is used to input the target features into a trained anomaly detection model to obtain the anomaly detection results of the service orchestration system; The third data acquisition module is used to acquire key indicator data from the current indicator data and the current operation monitoring data if the anomaly detection result indicates that an anomaly has occurred in the service orchestration system. The key indicator data includes at least one of the following: request error rate or request success rate, number of anomaly handling, request timeout rate or number of request timeouts, average request response time, resource utilization, and transaction success rate. The trigger judgment module is used to determine whether to trigger an alarm and perform fault recovery based on each key indicator data and the corresponding fault recovery trigger rules; The fault recovery module is used to perform corresponding fault recovery operations if it is determined that an alarm has been triggered and fault recovery has occurred.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the processor performs the steps of the method as described in any one of claims 1 to 6.
9. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
A method and system for determining fault of heterogeneous system based on machine learning
CN111209131A