Service fault processing method and device and related equipment

By collecting and analyzing business logs and middleware data, using abnormal detection and decision-making models to automatically identify faults and execute processing strategies, the problem of inefficient fault handling caused by manual intervention in the existing technology is solved, and rapid fault identification and automated recovery is achieved.

CN120498956APending Publication Date: 2025-08-15NEW H3C TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510472281.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The fault handling of existing systems relies on manual intervention and is difficult to quickly locate and recover, resulting in inefficient fault handling, affecting system availability and increasing operation and maintenance costs.

Method used

By collecting business logs and system middleware performance data, the anomaly detection module and decision-making model are used to automatically identify fault types and execute processing strategies, combining static rules and machine learning models to optimize decisions.

Benefits of technology

It realizes rapid fault identification and automated recovery, reduces fault processing time, improves system availability and reduces operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498956A_ABST
    Figure CN120498956A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network operation and maintenance, in particular to a service fault processing method and device and related equipment. The method comprises the following steps: collecting business log data of each key business service of a system and operation performance data of each system middleware; detecting the business log data and the operation performance data based on a preset anomaly detection module to judge whether abnormal data for representing a service fault exists or not; when the abnormal data is judged to exist, determining a target fault type of a target service fault corresponding to the abnormal data, and determining a target fault processing strategy corresponding to the target fault type based on a preset decision model; and executing the target fault processing strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network operation and maintenance technology, and in particular to a service fault handling method, apparatus, and related equipment. Background Art

[0002] With the widespread use of information technology, computer systems of all kinds play a key role in various fields. Whether it's the online services of internet companies or the information management systems of traditional enterprises, they all face the risk of system failures. If they can't be restored in a timely manner, it will seriously impact business operations.

[0003] Currently, system troubleshooting often relies on manual intervention. Typically, users detect anomalies and report them to operations and maintenance personnel, or they use simple monitoring tools to identify system anomalies. However, these monitoring tools can only detect basic metrics, such as server CPU usage and memory usage. They struggle to accurately detect and locate deeper faults caused by business logic errors or interactions between system middleware and business processes.

[0004] Existing fault location methods are inefficient. Operations and maintenance personnel must manually collect large amounts of business logs, system middleware logs, and other logs, then spend considerable time analyzing them to identify the root cause. For complex system architectures that include multiple types of system middleware (such as Kafka, MQ, and different types of databases), manual analysis is extremely difficult due to the wide variations in log formats and content, significantly increasing the time and cost of fault location.

[0005] Even after the fault is located, the recovery process still presents problems. Due to the lack of an automated recovery mechanism, operations and maintenance personnel need to manually perform a series of recovery operations, such as restarting related services and adjusting system configuration parameters. The entire recovery process is time-consuming and requires a high level of technical skills from the operations and maintenance personnel. Once an operation error occurs, it may lead to more serious consequences. For example, when restarting the database service, if the data is not properly backed up or the relevant components are not started in the correct order, data loss may occur or the system may not start normally. This situation of relying on manual intervention and long recovery time seriously affects system availability and business continuity, reduces the system's SLA (Service Level Agreement), and increases the company's operation and maintenance costs. In summary, the existing system fault handling methods have obvious shortcomings in terms of timeliness, degree of automation, and recovery efficiency, and urgently need to be improved. Summary of the Invention

[0006] The present application provides a service failure handling method, apparatus, and related equipment.

[0007] In a first aspect, the present application provides a method for handling service failures, the method comprising:

[0008] Collect business log data of each key business service of the system and the operating performance data of each system middleware;

[0009] Detecting the business log data and the operational performance data based on a preset anomaly detection module to determine whether there is abnormal data that characterizes a service failure;

[0010] When it is determined that abnormal data exists, determining a target fault type of the target service fault corresponding to the abnormal data, and determining a target fault handling strategy corresponding to the target fault type based on a preset decision model;

[0011] Execute the target fault handling strategy.

[0012] Optionally, each system middleware includes at least one of the following:

[0013] Distributed publish-subscribe messaging system, messaging middleware, and database middleware.

[0014] Optionally, the decision model includes a static rule engine and a dynamic machine learning model, wherein the static rule engine includes a mapping relationship between preset fault types and fault handling strategies, and the machine learning model is a model trained based on fault characteristics and fault handling strategies of historical service faults;

[0015] The step of determining a target fault handling strategy corresponding to the target fault type based on a preset decision model includes:

[0016] According to the mapping relationship between the preset fault types and fault handling strategies included in the static rule engine, a target fault handling strategy corresponding to the target fault type is determined.

[0017] Optionally, if the mapping relationship between fault types and fault handling strategies included in the static rule engine does not include a target fault handling strategy corresponding to the target fault type, the method further includes:

[0018] Extract the fault characteristics of the target service fault, and input the fault characteristics into the machine learning model to obtain a corresponding target fault handling strategy.

[0019] Optionally, after executing the target fault handling strategy, the method further includes:

[0020] Determine whether the target service failure has returned to normal;

[0021] If it is determined that the service has not returned to normal, the parameters of the machine learning model are adjusted, and the fault characteristics are input into the machine learning model to obtain the latest target fault handling strategy, and the steps of executing the latest target fault handling strategy are performed until the target service fault returns to normal.

[0022] Optionally, after determining that the target service failure has returned to normal, the method further includes:

[0023] Add the mapping relationship between the fault type of the target service fault and the latest target fault handling strategy to the static rule engine.

[0024] In a second aspect, the present application provides a service fault handling device, the device comprising:

[0025] The collection unit is used to collect business log data of each key business service of the system and the operating performance data of each system middleware;

[0026] a judgment unit, configured to detect the service log data and the operation performance data based on a preset anomaly detection module to determine whether there is abnormal data for characterizing a service failure;

[0027] a determination unit configured to, when determining that abnormal data exists, determine a target fault type of a target service fault corresponding to the abnormal data, and determine a target fault handling strategy corresponding to the target fault type based on a preset decision model;

[0028] An execution unit is configured to execute the target fault handling strategy.

[0029] Optionally, each system middleware includes at least one of the following:

[0030] Distributed publish-subscribe messaging system, messaging middleware, and database middleware.

[0031] Optionally, the decision model includes a static rule engine and a dynamic machine learning model, wherein the static rule engine includes a mapping relationship between preset fault types and fault handling strategies, and the machine learning model is a model trained based on fault characteristics and fault handling strategies of historical service faults;

[0032] When determining the target fault handling strategy corresponding to the target fault type based on the preset decision model, the determining unit is specifically configured to:

[0033] According to the mapping relationship between the preset fault types and fault handling strategies included in the static rule engine, a target fault handling strategy corresponding to the target fault type is determined.

[0034] Optionally, if the mapping relationship between fault types and fault handling strategies included in the static rule engine does not include a target fault handling strategy corresponding to the target fault type, the determining unit is further configured to:

[0035] Extract the fault characteristics of the target service fault, and input the fault characteristics into the machine learning model to obtain a corresponding target fault handling strategy.

[0036] Optionally, the device further comprises an adjustment unit:

[0037] The judging unit is further configured to judge whether the target service failure has been restored to normal;

[0038] If it is determined that the state has not returned to normal, the adjusting unit is specifically configured to adjust the parameters of the machine learning model;

[0039] The determining unit is further configured to input the fault characteristics into the machine learning model to obtain the latest target fault handling strategy; the executing unit is further configured to execute the latest target fault handling strategy;

[0040] Until the target service failure returns to normal.

[0041] Optionally, after determining that the target service failure has returned to normal, the apparatus further includes:

[0042] An adding unit is used to add a mapping relationship between the fault type of the target service fault and the latest target fault handling strategy to the static rule engine.

[0043] In a third aspect, an embodiment of the present application provides a service fault handling device, the service fault handling device comprising:

[0044] a memory for storing program instructions;

[0045] The processor is configured to call the program instructions stored in the memory and execute the steps of the method as described in any one of the first aspects above according to the obtained program instructions.

[0046] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable the computer to execute the steps of the method described in any one of the above-mentioned first aspects.

[0047] In summary, the service fault handling method provided in the embodiment of the present application collects the business log data of each key business service of the system and the operating performance data of each system middleware; based on a preset anomaly detection module, the business log data and the operating performance data are detected to determine whether there is abnormal data used to characterize the service fault; when it is determined that there is abnormal data, the target fault type of the target service fault corresponding to the abnormal data is determined, and based on a preset decision model, the target fault handling strategy corresponding to the target fault type is determined; and the target fault handling strategy is executed.

[0048] The service fault handling method provided in the embodiments of this application, through real-time analysis of service logs and continuous probing of system middleware, can quickly detect faults at their earliest stages, even in their infancy. Compared to traditional methods that rely on user feedback or regular inspections to detect faults, this method can significantly shorten fault detection time, enabling the system to initiate recovery processes in the shortest possible time, preventing faults from worsening and impacting services, and significantly improving the timeliness of fault handling. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments of the present application or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings of the embodiments of the present application.

[0050] Figure 1 A detailed flowchart of a service fault handling method provided in an embodiment of the present application;

[0051] Figure 2 A schematic diagram of the structure of a service fault handling device provided in an embodiment of the present application;

[0052] Figure 3 A schematic diagram of the hardware architecture of another service fault handling device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a," "the," and "the" used in this application and claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more associated listed items.

[0054] It should be understood that although the terms first, second, third, etc. may be used to describe various information in the embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" used may also be interpreted as "at the time of" or "when" or "in response to determining".

[0055] For example, see Figure 1 FIG. 1 is a detailed flow chart of a method for handling service failures provided in an embodiment of the present application, the method comprising the following steps:

[0056] Step 100: Collect business log data of each key business service of the system and operation performance data of each system middleware.

[0057] In practice, deploying the detection service requires servers with stable performance and ample computing and storage resources. For example, servers with multi-core CPUs, large memory capacities, and high-speed storage devices ensure efficient operation when processing large amounts of data and complex calculations. Depending on the system scale and workload, either a standalone or clustered deployment can be chosen. Small-scale systems utilize a standalone deployment; large-scale, highly concurrent systems utilize a cluster, using a load balancer to evenly distribute tasks across nodes, improving system reliability and processing power.

[0058] A detection service is a microservice that can perform task detection based on the detection configuration. Specifically, you can configure which business services are system-critical business services. In this way, you can deploy a log collector on the node (virtual machine / physical machine) where the system-critical business services are located to collect business log data in real time and quickly transmit the collected business log data to the detection service's log storage center via a preset network transmission protocol. Furthermore, the log collector can also perform preliminary verification and organization of the collected business log data to remove duplicate or invalid data and ensure data completeness and accuracy.

[0059] In an embodiment of the present application, the detection service can also detect the operating performance data of each system middleware in real time. It should be noted that in an embodiment of the present application, the system middleware may include at least one of the following: a distributed publish-subscribe message system (such as Kafka), a message middleware (such as queue MQ) and a database middleware.

[0060] Then, you can detect the operating performance data of each system middleware, for example:

[0061] Kafka monitoring includes message production detection: The detection service uses the Kafka management API to obtain real-time metrics such as producer message sending rate and success rate. If a producer's message sending rate suddenly drops or its success rate remains below the threshold, it is determined that there may be a network connection or code logic failure. Message consumption monitoring: The Kafka management API is used to monitor the consumer group's message consumption rate and backlog. When the message backlog exceeds the threshold, the analysis is conducted to determine whether it is due to insufficient consumer processing capacity or a Kafka cluster configuration issue, such as the number of partitions and the reasonableness of replica distribution.

[0062] MQ monitoring includes queue status monitoring: The detection service regularly queries the status of each MQ queue, such as queue length and message entry and exit rates. A continuously increasing queue length and slow exit rate may indicate message processing congestion. Message flow monitoring: A specific message tracking mechanism is set up in the MQ system. The detection service tracks the flow of messages from producer to consumer, quickly locating any remaining or lost messages.

[0063] Database monitoring includes connection monitoring: The detection service uses database connection pooling technology to regularly test database connection status to ensure availability. If the number of connection failures exceeds a threshold, it may indicate a database failure, such as server downtime or network outage. Read and write performance monitoring: Using the database's own or third-party monitoring plug-ins, database read and write operation response time, throughput, and other performance metrics are collected in real time. When read and write performance degrades, analysis is performed to determine if the cause is excessive load, index failure, or insufficient hardware resources.

[0064] Step 110: Detect the service log data and the operation performance data based on a preset anomaly detection module to determine whether there is abnormal data for characterizing a service failure.

[0065] In practice, the detection service integrates business log analysis results with system middleware monitoring data. For example, if frequent errors in a business module are detected in the business log and the corresponding database read and write performance is abnormal, a preliminary diagnosis is made that a fault exists in the interaction between the business module and the database. By analyzing system component dependencies and call chains, the scope of the fault can be determined, including a single business process, an entire business line, or multiple business lines.

[0066] Specifically, the detection service uses a log parsing module to convert unstructured log text into structured data based on regular expressions, semantic analysis, and other technologies. For example, for the log "[2024-01-01 10:00:00]INFO OrderService-Order

[12345] has been created successfully," it can extract key information such as the timestamp, log level, service name, and order number to determine whether the business process is normal, such as whether there are any anomalies such as order creation failure or processing timeout.

[0067] Furthermore, a machine learning anomaly detection model based on time series analysis is used to monitor the parsed business log data in real time. This model learns the patterns of normal business operation log data and identifies anomalies when data deviates significantly from these patterns. If the interval between order creations changes unexpectedly, the model issues an alert, and the detection service further analyzes the scope of the anomaly's impact on the business.

[0068] For the detected operating performance data of each system middleware, corresponding thresholds can be set in advance based on empirical values for each performance data. In this way, it is possible to determine whether each performance data is abnormal by referring to the preset thresholds of each performance data.

[0069] Step 120: When it is determined that abnormal data exists, a target fault type of the target service fault corresponding to the abnormal data is determined, and based on a preset decision model, a target fault handling strategy corresponding to the target fault type is determined.

[0070] In an embodiment of the present application, when abnormal data is determined, the fault type of the corresponding service failure can be determined based on the abnormal data. For example, the detection service obtains indicators such as the producer message sending rate and success rate in real time through the Kafka management API. If a producer's message sending rate suddenly drops or its success rate remains below a threshold, it is determined that there may be a fault in the network connection, code logic, etc. At this time, a fault handling strategy for resolving the fault can be automatically determined based on the decision model included in the detection service.

[0071] Step 130: Execute the target fault handling strategy.

[0072] In an embodiment of the present application, the decision model includes a static rule engine and a dynamic machine learning model, wherein the static rule engine includes a mapping relationship between preset fault types and fault handling strategies, and the machine learning model is a model trained based on the fault characteristics and fault handling strategies of historical service faults;

[0073] Then, when determining the target fault handling strategy corresponding to the target fault type based on the preset decision model, a preferred implementation method is:

[0074] According to the mapping relationship between the preset fault types and fault handling strategies included in the static rule engine, a target fault handling strategy corresponding to the target fault type is determined.

[0075] In practice, the rule engine pre-defines a series of rules for handling different fault types (pre-set based on operational experience and used for subsequent fault handling and recovery). These rules are based on in-depth analysis of various common system faults and extensive operational experience. They provide the decision-making model with the ability to quickly respond and make preliminary judgments.

[0076] For example, consider rules for Kafka failures: When Kafka experiences a message backlog, the rules engine makes a judgment based on predefined rules. For example, if the message backlog exceeds the threshold for normal system operation and persists for a certain period of time, the rules engine triggers an instruction to increase the number of partitions. This is because increasing the number of partitions improves the parallel processing capabilities of the Kafka cluster, thereby accelerating message consumption and alleviating the backlog. Furthermore, the rules engine further determines the specific number of partitions to increase based on factors such as the growth trend of the backlog and the business's requirements for message timeliness. For example, for business scenarios with extremely high timeliness requirements, if the backlog reaches a threshold, the engine will aggressively increase the number of partitions to quickly restore message processing speed. However, for businesses with less timeliness requirements, the engine will be more cautious and gradual in increasing partitions to avoid wasting system resources due to excessive partition additions.

[0077] Another example is a rule for database failures: when the number of database connections is detected to be too high, the rule engine determines that the database may be overloaded. At this point, it adjusts the connection pool parameters according to pre-set rules. For example, it might increase the maximum number of connections in the connection pool to accommodate more business requests, or adjust the connection timeout to prevent business requests from being blocked due to long waits for connections. Furthermore, if the database's read and write performance experiences anomalies, such as long read response times, the rule engine might determine whether this is due to index failure, triggering rules to rebuild indexes or optimize query statements.

[0078] Furthermore, in an embodiment of the present application, if the mapping relationship between fault types and fault handling strategies included in the static rule engine does not include a target fault handling strategy corresponding to the target fault type, the method may further include the following steps:

[0079] Extract the fault characteristics of the target service fault, and input the fault characteristics into the machine learning model to obtain a corresponding target fault handling strategy.

[0080] Machine learning algorithms play a key role in optimization and adaptation within decision-making models. Through deep learning of historical failure data and corresponding processing results, they continuously improve the accuracy and rationality of decision-making models, enabling them to better cope with complex and diverse failure scenarios.

[0081] Specifically, the system first collects a large amount of historical fault data. This data covers detailed information on various system middleware (such as Kafka, MQ, databases, etc.) under different fault conditions, including the system status at the time of the fault, the values of relevant monitoring indicators, the treatment measures taken, and the final treatment results. Then, this raw data is preprocessed, including data cleaning, feature extraction and selection. For example, key features closely related to the fault are extracted from complex log data, such as the number of partitions, the number of consumers, and the message generation rate when Kafka messages are backlogged; for database faults, features such as the number of connections, read and write throughput, and query statement execution time are extracted. This preprocessed data will serve as training samples for the machine learning algorithm.

[0082] Then, appropriate machine learning algorithms, such as decision trees, random forests, and support vector machines, are used to train the preprocessed data. During training, the algorithm learns the inherent relationship between fault characteristics and handling outcomes, building a model that can predict the optimal recovery approach. For example, by learning from a large amount of Kafka message backlog failure data, the model can determine whether the most effective recovery approach is to increase the number of partitions, adjust consumer configurations, or perform other actions, given varying backlog levels, system resource availability, and business needs. Furthermore, to improve the model's generalization and accuracy, cross-validation and regularization techniques are used to optimize the model and avoid overfitting or underfitting.

[0083] In summary, machine learning algorithms play a key role in optimization and adaptation within decision-making models. Through deep learning of historical fault data and corresponding processing results, they continuously improve the accuracy and rationality of decision-making models, enabling them to better cope with complex and diverse fault scenarios. The model can be trained using historical data to produce a trained machine learning model. If, after identifying a fault, no corresponding fault handling strategy is found within the static rule engine's mapping between fault type and fault handling strategy, the fault's fault characteristics can be extracted and fed into the machine learning model to generate the corresponding fault handling strategy.

[0084] In an embodiment of the present application, after executing the fault handling strategy, it is further determined whether the target service fault has returned to normal; if it is determined that it has not returned to normal, the parameters of the machine learning model are adjusted, and the fault characteristics are input into the machine learning model to obtain the latest target fault handling strategy, and the steps of executing the latest target fault handling strategy are performed until the target service fault returns to normal.

[0085] In other words, the decision model is not static but rather has the ability to learn in real time. During actual operation, every time the system handles a fault, new fault data and handling results are fed back to the machine learning algorithm. The algorithm uses this new data to update and optimize the model in real time, enabling it to continuously adapt to changes in the system's operating environment and emerging fault types. For example, if the introduction of a new business module in the system causes a new fault pattern in Kafka message processing, the decision model can learn from this new fault data and promptly adjust its decision-making strategy to better respond to similar faults.

[0086] In the embodiment of the present application, the decision model may also be used as follows:

[0087] The rule engine and machine learning algorithm work together, complementing each other's strengths. When the detection service detects a system failure, the rule engine first makes a preliminary diagnosis based on predefined rules and provides an empirically-based recovery recommendation. Simultaneously, the machine learning algorithm, based on the current failure characteristics and historical learning experience, generates a recovery recommendation informed by data analysis and model predictions. The decision model then comprehensively considers these two recommendations, using a weighting and integration strategy to determine the final recovery approach. For example, for common failures with relatively fixed patterns, the rule engine's recommendations may be given higher weight; for complex and rare failures, the machine learning algorithm's recommendations, based on big data analysis, may be given greater weight. This combined approach allows the decision model to leverage the rules engine's rapid response and determinism to make swift decisions in common failure scenarios, while also leveraging the machine learning algorithm's intelligence and adaptability to provide more precise and effective recovery solutions in complex failure scenarios, thereby comprehensively improving the system's ability to cope with various types of failures.

[0088] In an embodiment of the present application, once the decision model determines the recovery means, the detection service will automatically perform these operations through automated scripts or calling the management API of the corresponding system. For example, for the operation of increasing the number of partitions in Kafka, the detection service will call Kafka's AdminClient API to complete it. Specifically, it will construct the corresponding request parameters based on the number of partitions increased by the decision model and send them to the controller of the Kafka cluster. The controller is responsible for creating new partitions in the cluster and coordinating relevant nodes to perform operations such as data allocation and synchronization. For the operation of adjusting the database connection pool parameters, the detection service will modify the configuration file of the database connection pool based on the decision results. Taking a common database connection pool (such as HikariCP) as an example, the detection service will locate its configuration file, modify key parameters such as the maximum number of connections and the minimum number of idle connections, and notify the database connection pool to reload the configuration by sending specific signals or instructions, thereby adjusting the connection pool parameters. After executing the recovery operation, the detection service will continue to monitor the system status and obtain the system's operating indicators in real time through the same business log analysis and system middleware monitoring methods as before. If it is found that the system has not yet returned to normal, the detection service will restart the fault judgment and decision-making process, adjust the recovery strategy, and ensure that the fault is completely resolved and the system returns to normal operation.

[0089] Based on the same inventive concept as the above-mentioned embodiment, for example, refer to Figure 2 FIG. 1 is a schematic diagram of a structure of a service fault handling device provided in an embodiment of the present application, the device comprising:

[0090] The collection unit 20 is used to collect business log data of each key business service of the system and the operating performance data of each system middleware;

[0091] A judgment unit 21 is configured to detect the service log data and the operation performance data based on a preset anomaly detection module to determine whether there is abnormal data that is used to characterize a service failure;

[0092] The determining unit 22 is configured to determine, when determining that abnormal data exists, a target fault type of the target service fault corresponding to the abnormal data, and determine a target fault handling strategy corresponding to the target fault type based on a preset decision model;

[0093] The execution unit 23 is configured to execute the target fault handling strategy.

[0094] Optionally, each system middleware includes at least one of the following:

[0095] Distributed publish-subscribe messaging system, messaging middleware, and database middleware.

[0096] Optionally, the decision model includes a static rule engine and a dynamic machine learning model, wherein the static rule engine includes a mapping relationship between preset fault types and fault handling strategies, and the machine learning model is a model trained based on fault characteristics and fault handling strategies of historical service faults;

[0097] When determining the target fault handling strategy corresponding to the target fault type based on the preset decision model, the determining unit 22 is specifically configured to:

[0098] According to the mapping relationship between the preset fault types and fault handling strategies included in the static rule engine, a target fault handling strategy corresponding to the target fault type is determined.

[0099] Optionally, if the mapping relationship between fault types and fault handling strategies included in the static rule engine does not include a target fault handling strategy corresponding to the target fault type, the determining unit 22 is further configured to:

[0100] Extract the fault characteristics of the target service fault, and input the fault characteristics into the machine learning model to obtain a corresponding target fault handling strategy.

[0101] Optionally, the device further comprises an adjustment unit:

[0102] The judging unit 21 is further configured to judge whether the target service failure has been restored to normal;

[0103] If it is determined that the state has not returned to normal, the adjusting unit is specifically configured to adjust the parameters of the machine learning model;

[0104] The determining unit 22 is further configured to input the fault characteristics into the machine learning model to obtain the latest target fault handling strategy; the executing unit 23 is further configured to execute the latest target fault handling strategy;

[0105] Until the target service failure returns to normal.

[0106] Optionally, after determining that the target service failure has returned to normal, the apparatus further includes:

[0107] An adding unit is used to add a mapping relationship between the fault type of the target service fault and the latest target fault handling strategy to the static rule engine.

[0108] The above units may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, when a unit is implemented by scheduling program code through a processing element, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these units may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0109] Furthermore, the service fault handling device provided in the embodiment of the present application, from the hardware level, the hardware architecture diagram of the service fault handling device can be found in Figure 3 As shown, the service fault processing device may include: a memory 30 and a processor 31,

[0110] The memory 30 is used to store program instructions. The processor 31 calls the program instructions stored in the memory 30 and executes the above method embodiment according to the obtained program instructions. The specific implementation method and technical effect are similar and will not be repeated here.

[0111] Optionally, the present application also provides a service fault processing device, comprising at least one processing element (or chip) for executing the above method embodiment.

[0112] Optionally, the present application also provides a program product, such as a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable the computer to execute the above method embodiments.

[0113] Here, the machine-readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as CD, DVD, etc.), or similar storage media, or a combination thereof.

[0114] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0115] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0116] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0117] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0118] Furthermore, these computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0120] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for handling service failures, characterized in that: The method comprises: Collect business log data of each key business service of the system and the operating performance data of each system middleware; Detecting the business log data and the operational performance data based on a preset anomaly detection module to determine whether there is abnormal data that characterizes a service failure; When it is determined that abnormal data exists, determining a target fault type of the target service fault corresponding to the abnormal data, and determining a target fault handling strategy corresponding to the target fault type based on a preset decision model; Execute the target fault handling strategy.

2. The method according to claim 1, wherein Each system middleware includes at least one of the following: Distributed publish-subscribe messaging system, messaging middleware, and database middleware.

3. The method according to claim 1 or 2, wherein: The decision model includes a static rule engine and a dynamic machine learning model. The static rule engine includes a mapping relationship between preset fault types and fault handling strategies, and the machine learning model is a model trained based on the fault characteristics and fault handling strategies of historical service faults. The step of determining a target fault handling strategy corresponding to the target fault type based on a preset decision model includes: According to the mapping relationship between the preset fault types and fault handling strategies included in the static rule engine, a target fault handling strategy corresponding to the target fault type is determined.

4. The method according to claim 3, wherein If the mapping relationship between fault types and fault handling strategies included in the static rule engine does not include a target fault handling strategy corresponding to the target fault type, the method further includes: Extract the fault characteristics of the target service fault, and input the fault characteristics into the machine learning model to obtain a corresponding target fault handling strategy.

5. The method according to claim 4, wherein After executing the target fault handling strategy, the method further includes: Determine whether the target service failure has returned to normal; If it is determined that the service has not returned to normal, the parameters of the machine learning model are adjusted, and the fault characteristics are input into the machine learning model to obtain the latest target fault handling strategy, and the steps of executing the latest target fault handling strategy are performed until the target service fault returns to normal.

6. The method according to claim 5, wherein After determining that the target service failure has returned to normal, the method further includes: Add the mapping relationship between the fault type of the target service fault and the latest target fault handling strategy to the static rule engine.

7. A service fault handling device, characterized in that: The device comprises: The collection unit is used to collect business log data of each key business service of the system and the operating performance data of each system middleware; a judgment unit, configured to detect the service log data and the operation performance data based on a preset anomaly detection module to determine whether there is abnormal data for characterizing a service failure; a determination unit configured to, when determining that abnormal data exists, determine a target fault type of a target service fault corresponding to the abnormal data, and determine a target fault handling strategy corresponding to the target fault type based on a preset decision model; An execution unit is configured to execute the target fault handling strategy.

8. The device according to claim 7, wherein The decision model includes a static rule engine and a dynamic machine learning model. The static rule engine includes a mapping relationship between preset fault types and fault handling strategies, and the machine learning model is a model trained based on the fault characteristics and fault handling strategies of historical service faults. When determining the target fault handling strategy corresponding to the target fault type based on the preset decision model, the determining unit is specifically configured to: According to the mapping relationship between the preset fault types and fault handling strategies included in the static rule engine, a target fault handling strategy corresponding to the target fault type is determined.

9. The device according to claim 8, wherein If the target fault handling strategy corresponding to the target fault type does not exist in the mapping relationship between the fault type and the fault handling strategy included in the static rule engine, the determining unit is further configured to: Extract the fault characteristics of the target service fault, and input the fault characteristics into the machine learning model to obtain a corresponding target fault handling strategy.

10. A service fault handling device, characterized in that: The service failure processing device includes: a memory for storing program instructions; The processor is configured to call the program instructions stored in the memory, and execute the steps of the method according to any one of claims 1 to 6 according to the obtained program instructions.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable the computer to execute the steps of the method according to any one of claims 1 to 6.