Fault processing method and system, electronic equipment and storage medium

By setting up tracking points in the microservice architecture to obtain status data and perform fault correlation analysis, fault instances can be automatically identified and repaired, solving the problem of long fault recovery time in the microservice architecture and improving the stability and availability of business systems.

CN121008946APending Publication Date: 2025-11-25BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511050228.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In a microservices decomposition architecture, each microservice is handled by a different team, resulting in a lack of unified fault handling measures. This leads to long fault recovery times, hinders rapid and automated fault recovery, and reduces the overall availability and stability of the business system.

Method used

By setting up tracking points for each service instance in the call chain, obtaining status data, parsing the call relationships between service instances, performing fault correlation analysis, identifying fault instances, and determining and executing fault repair strategies based on fault types, automated fault repair is achieved.

Benefits of technology

Without the need for manual troubleshooting and coordination, it can quickly locate faulty service instances and implement targeted repairs, shortening fault recovery time, improving the stability and availability of business systems, and ensuring business continuity and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121008946A_ABST
    Figure CN121008946A_ABST
Patent Text Reader

Abstract

The invention provides a fault processing method and system, electronic equipment and a storage medium, and the method comprises the steps: obtaining the state data of each service instance through a burying point of each service instance preset in a call link, and the call link comprises a plurality of service instances; analyzing a calling relation between the service instances according to the state data, performing fault correlation analysis on the calling link based on the state data and the calling relation, and identifying the service instance with a fault as a fault instance; and determining the fault type of the fault instance, and determining and executing a fault recovery strategy for the fault instance based on the fault type. Therefore, one-by-one manual troubleshooting and coordination are not needed, fault service instances can be rapidly positioned and targeted repair can be implemented by means of an automatic process, the fault recovery time is greatly shortened, the service interruption duration is reduced, the service system can operate more stably, the overall availability of the service system during complex function processing is effectively improved, and the service efficiency is improved. And the continuity and the high efficiency of the service are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a fault handling method, system, electronic device and storage medium. Background Technology

[0002] In the current software development field, in order to improve development efficiency, large business systems generally adopt a microservice decomposition architecture, which breaks down a large business system into a group of small, independent microservices. Each microservice is built around a specific business function, can run independently, and interact through a lightweight communication mechanism to jointly complete the functions of the entire business system.

[0003] In this architecture, when a business interface needs to complete a complex function, it often needs to call interfaces provided by multiple different microservices to work together. Moreover, the calls between these microservices do not have to be parallel, but rather occur sequentially according to a certain business logic order, forming a serial call chain. If any link in the chain fails, the entire business function can be interrupted.

[0004] However, since each microservice is usually handled by a different team, there is a lack of unified fault handling measures. When problems occur, they often rely on manual coordination, which not only prolongs the fault recovery time and makes it impossible to achieve fast and automated fault recovery, but also reduces the overall availability and stability of the business system. Summary of the Invention

[0005] To address the aforementioned technical problems, this application discloses a fault handling method, system, electronic device, and storage medium, to at least resolve the issue that microservice decomposition architectures in related technologies cannot achieve rapid and automated fault recovery, thus reducing the overall availability and stability of business systems. The technical solution disclosed herein is as follows:

[0006] In a first aspect, this application discloses a fault handling method, the method comprising:

[0007] By pre-setting instrumentation points for each service instance in the call chain, the status data of each service instance is obtained, and the call chain includes multiple service instances;

[0008] The call relationship between the service instances is parsed based on the status data, and based on the status data and the call relationship, a fault correlation analysis is performed on the call chain to identify the service instance that has failed as a fault instance.

[0009] Determine the fault type of the fault instance, and based on the fault type, determine and execute a fault repair strategy for the fault instance.

[0010] Secondly, embodiments of the present invention provide a fault handling system, including:

[0011] The monitoring center is used to obtain the status data of each service instance by pre-setting the instrumentation points of each service instance in the call chain, which includes multiple service instances;

[0012] The scheduling center is used to synchronize the status data to the policy center;

[0013] The strategy center is used to parse the call relationship between the service instances based on the status data, and perform fault correlation analysis on the call chain based on the status data and the call relationship to identify the service instance that has failed as a fault instance; determine the fault type of the fault instance, and determine the fault repair strategy for the fault instance based on the fault type.

[0014] The dispatch center is also used to synchronize the fault repair strategy to the processing center;

[0015] The processing center is used to execute fault repair strategies for the faulty instances.

[0016] Thirdly, embodiments of the present invention provide an electronic device, including:

[0017] processor;

[0018] Memory used to store the processor's executable instructions;

[0019] The processor is configured to execute the instructions to implement the fault handling method described in any of the preceding claims.

[0020] Fourthly, embodiments of the present invention provide a computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by a processor of a fault-handling electronic device, the fault-handling electronic device is able to perform any of the fault-handling methods described above.

[0021] Fifthly, embodiments of the present invention provide a computer program product, including a computer program that, when executed by a processor, implements the fault handling method described in any of the preceding claims.

[0022] Compared with the prior art, this application has the following advantages:

[0023] In this application, by setting up instrumentation points for each service instance in the call chain, the status data of each service instance in the call chain is obtained. Then, the call relationship between service instances is parsed and fault correlation analysis is performed to identify faulty instances. Subsequently, fault repair strategies are determined and executed based on the fault type. This process does not require manual investigation and coordination one by one. It can quickly locate faulty service instances and implement targeted repairs through automated processes, which greatly shortens the fault recovery time, reduces business interruption time, enables the business system to run more stably, effectively improves the overall availability of the business system when handling complex functions, and ensures business continuity and efficiency. Attached Figure Description

[0024] Figure 1 This is a flowchart of the steps of a fault handling method according to this application;

[0025] Figure 2 This is a schematic diagram of a fault handling system according to this application;

[0026] Figure 3 This is a structural block diagram of a fault handling system according to this application;

[0027] Figure 4 This is a schematic diagram of an electronic device according to this application;

[0028] Figure 5 This is a block diagram of a fault handling device according to this application. Detailed Implementation

[0029] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0031] In related technologies, large-scale business systems commonly adopt a microservices decomposition architecture, breaking down a large business system into a set of small, independent microservices. Each microservice is built around a specific business function, can run independently, and interacts through lightweight communication mechanisms to collectively complete the functionality of the entire business system. When a business interface needs to complete a complex function, it often needs to call interfaces provided by multiple different microservices to collaboratively complete the task. The call relationships between these microservices do not have to be parallel, but rather occur sequentially according to a certain business logic order, forming a serial call chain.

[0032] For example, an "order placement" interface may need to call the "product" microservice to query inventory, call the "user" microservice to verify user information, call the "inventory" microservice to deduct inventory, and call the "order" microservice to create an order. Among these, inventory must be successfully deducted before an order can be created.

[0033] However, when any link in the call chain fails, the entire business function can be interrupted. Moreover, since each microservice is usually handled by a different team, there is a lack of unified fault handling measures. When problems occur, manual coordination is often required, which not only prolongs the fault recovery time and fails to achieve rapid and automated fault recovery, but also reduces the overall availability and stability of the business system. Based on this, the fault handling method in this application is proposed to solve the above problems.

[0034] Reference Figure 1 The diagram illustrates a flowchart of a fault handling method according to this application, which may specifically include the following steps:

[0035] In step S11, the status data of each service instance is obtained by pre-setting the instrumentation points for each service instance in the call chain. The call chain includes multiple service instances.

[0036] In this step, the status data of each service instance in the call chain can be obtained. In a microservice splitting architecture, the call chain consists of multiple service nodes, and each service node contains at least one service instance. The status data can include, but is not limited to, information such as CPU utilization, memory utilization, interface success rate, response latency, error logs, and communication status with other service instances for each service instance.

[0037] One approach is to deploy a monitoring center throughout the microservice architecture, which can embed tracking points at the code level of each service instance to collect the status data of each service instance in real time.

[0038] Specifically, the first step is to design a data tracking scheme in advance based on the business characteristics and interaction logic of each service instance in the call chain, and to clarify the links that need to be monitored and the types of status data that need to be collected in each service instance.

[0039] Then, according to the designed tracking scheme, tracking programs are embedded in the code of each service instance. When the service instance runs, the tracking program will automatically trigger the data collection operation and encapsulate the collected data in a unified data format to obtain status data, so as to ensure the standardization and parsability of the status data.

[0040] Furthermore, the encapsulated status data can be transmitted to the monitoring center through a pre-defined communication mechanism, enabling centralized collection and storage of status data for each service instance, and providing data support for subsequent call relationship parsing and fault analysis.

[0041] In step S12, the call relationship between service instances is parsed based on the status data, and based on the status data and call relationship, a fault correlation analysis is performed on the call chain to identify the service instance that has failed as a fault instance.

[0042] In this step, the call paths and dependencies between various service instances can be identified based on the status data. Then, by combining these call relationships and status data, a correlation analysis of the fault can be performed from the perspective of the overall call chain. This allows for the precise location of the service instance that has failed in the call chain, thus identifying it as a faulty instance and clarifying the specific location of the fault, laying the foundation for subsequent fault repair.

[0043] Specifically, we can first analyze the topology of the entire call chain to clarify the call relationships between various service instances, such as instance A calling instance B, and instance B calling instance C, etc. Then, we can conduct fault correlation analysis by combining the status data of each service instance to determine the propagation path of the fault in the call chain, thereby identifying the service instance that caused this series of faults as the fault instance.

[0044] Fault correlation analysis can employ methods such as traffic analysis, load analysis, and middleware analysis. Traffic analysis refers to identifying faults caused by abnormal changes in request traffic, such as connection request overload or malicious request attacks, by analyzing abnormal changes in request traffic. Load analysis refers to identifying faults caused by resource exhaustion or unreasonable configuration by analyzing the resource usage of service nodes, such as excessive CPU utilization, excessive memory utilization, and full connection pools. Middleware analysis refers to identifying the chain reaction caused by middleware failures by analyzing the status of middleware that the service depends on (such as databases, caches, and message queues), such as cache failures and message queue blocking.

[0045] In step S13, the fault type of the fault instance is determined, and based on the fault type, a fault repair strategy for the fault instance is determined and executed.

[0046] It is understandable that different fault types usually require different fault repair strategies. Therefore, in this step, after identifying the fault instance, it is necessary to accurately determine its fault type.

[0047] For example, when the CPU utilization of the service node where the faulty instance is located is consistently above 95%, memory utilization spikes, and interface response latency increases significantly, it indicates that the fault type is insufficient resource carrying capacity of the service node where the faulty instance is located.

[0048] When the success rate of the service instance's interface drops sharply, and error logs frequently show messages such as "connection timeout" and "unable to access downstream services," and the network packet loss rate increases abnormally, it indicates that the fault type is a network communication fault.

[0049] When some business interfaces of a service instance return error data, and the logs contain keywords such as "business logic verification failed," but other basic operational metrics are normal, it indicates that the fault type is a business logic error.

[0050] Then, based on the type of fault, the corresponding fault repair strategy can be determined and implemented.

[0051] For example, if the fault type is that the service node where the faulty instance is located has insufficient resources, the fault repair strategy can be to expand the service instance capacity of the service node where the faulty instance is located, thereby distributing the business pressure by increasing the number of service instances, and at the same time adjusting the resource configuration parameters to increase its ability to carry business.

[0052] If the fault type is a network communication fault, the fault repair strategy can be to check the network links of the faulty instance, perform network configuration adjustments and restart network components on the faulty network links, and if necessary, temporarily disconnect the communication between the faulty instance and downstream instances to prevent the fault from spreading.

[0053] If the fault type is a business logic error, the fault recovery strategy could be to temporarily block the faulty instance from calling the business interface and enable a fallback scheme, and then restore normal calls after the problem is resolved; etc.

[0054] In this way, by accurately identifying the type of fault, we can more effectively determine and implement repair strategies to restore the normal operation of the faulty node as soon as possible, and ensure the stability and reliability of the entire call chain and business system.

[0055] In one implementation, step S12 involves parsing the call relationships between service instances based on the state data, including:

[0056] Based on the status data, determine the upstream and / or downstream service instances of the service instance to obtain the calling relationship between the service instances;

[0057] Furthermore, based on status data and call relationships, fault correlation analysis is performed on the call chain to identify service instances that have failed, which are considered fault instances, including:

[0058] Select status data that meets the pre-defined abnormal conditions as abnormal data;

[0059] Based on the call relationship, root cause analysis is performed on the abnormal data to identify the service instance that caused the failure, which is then designated as the failure instance.

[0060] In this implementation, firstly, the call relationship between service instances can be parsed based on the status data. Specifically, the source information of the request received by each service instance and the target information of the request sent can be extracted from the status data. If the target of the request sent by instance A is instance B, then A is the upstream instance of B and B is the downstream instance of A. In this way, the upstream and downstream associations between all service instances can be sorted out, thereby clarifying the call relationship of service instances in the entire call chain.

[0061] Based on this, anomaly analysis can be performed on the status data of each service instance in the call chain. By setting anomaly conditions, status data that meets the conditions can be filtered out. These anomaly conditions can be set based on business requirements and historical operational data. For example, anomaly conditions may include response time exceeding a preset threshold, the number of failed calls reaching a specified proportion, or the appearance of a specific error code; the specific conditions are not limited.

[0062] It's understandable that service instances have calling relationships, and their operation is interconnected. For example, an excessively long response time for a service instance might be due to incorrect request data sent by an upstream instance, or a failed call to a service instance might cause a sharp drop in the call volume of downstream instances, and so on. However, identifying abnormal data can only reflect the abnormal behavior of each service instance itself, but it cannot determine whether these abnormalities are due to problems within the service instance itself or are affected by other service instances.

[0063] Therefore, root cause analysis of abnormal data can be performed by combining the established service instance call relationships. For example, if the status data of a service instance is abnormal, the abnormal data of its upstream and downstream instances will be traced. Then, based on the abnormal data of multiple traced service instances, the root cause of the failure can be further analyzed, and the service instance containing the root cause can be identified as the faulty instance. Through this layer-by-layer correlation analysis, indirect effects caused by failures of downstream or upstream instances can be eliminated, ultimately accurately identifying the root cause in the call chain that caused the abnormality and triggered a chain reaction.

[0064] For example, status data could be API success rate. When the API success rate of instance A drops, we can check if the response latency of its downstream instance B is abnormal. If instance B is abnormal, we can further investigate the status data of its downstream instance C. If instance C's API success rate is normal, then instance B is determined to be a faulty instance. Alternatively, status data could be connection pool status. When instance B's connection pool frequently runs out of resources, and its API success rate for calling instance C drops sharply, while instance C's resource utilization abnormally increases, correlation analysis can determine that instance C is a faulty instance.

[0065] Through this progressive correlation analysis, indirect effects caused by downstream or upstream instance failures are eliminated, and the instance node that first experienced an anomaly in the call chain and triggered a chain reaction is finally accurately identified, i.e., the fault node.

[0066] This allows for precise identification of the root cause of the failure, ensuring accurate location of the fault instance. It provides a clear and accurate target for implementing targeted fault repair strategies, helping to more efficiently troubleshoot and resolve faults in the microservice architecture and ensuring the stable operation of the business system.

[0067] In one implementation, determining the fault type of a fault instance includes:

[0068] Get the connection pool status of the faulty instance. The connection pool status is the resource pool status used to establish connections between the faulty instance and other service instances.

[0069] If the connection pool status meets the overload condition, the fault type of the faulty instance is determined to be connection quantity overload;

[0070] Therefore, based on the fault type, the fault repair strategy for the faulty node is determined and implemented, including:

[0071] If the fault type is connection overload, increase the connection pool size of the faulty instance from the current first value to the second value.

[0072] As can be understood, a connection pool is a resource pool established by a service instance to efficiently manage network connections with other service instances. By pre-creating and maintaining a certain number of connections in the connection pool, service instances can directly reuse them when they need to communicate with other service nodes, avoiding the performance loss caused by frequently creating and closing connections. Each service instance usually maintains its own corresponding connection pool.

[0073] In this implementation, the failure type of a failed instance can be determined by obtaining the connection pool status of the failed instance. The connection pool status includes, but is not limited to, data such as the current number of active connections, the number of idle connections, the number of connection timeouts, the maximum number of connections, and the length of the connection waiting queue, all of which are crucial for determining the failure type.

[0074] When the connection pool status meets the overload condition, it means that the number of connections managed by the current faulty instance exceeds its reasonable capacity. In other words, the number of available connections in the connection pool is insufficient to meet the business call requirements of the service instance. Therefore, the fault type of the faulty instance can be determined as connection overload.

[0075] Overload conditions are usually set based on business scenarios and system performance, such as the number of active connections reaching or exceeding the maximum allowed number of connections in the connection pool, the length of the request queue waiting for connections continuously increasing and exceeding the preset length, the number of timeouts for obtaining connections exceeding the preset number, and so on.

[0076] Therefore, the fault repair strategy could be to increase the connection pool size of the faulty instance from the current first value to a second value. The setting of the second value needs to be considered in conjunction with the service load of the faulty instance and the system resource capacity. For example, if the original maximum number of connections in the connection pool of the faulty instance (the first value) was 100, after analyzing its peak service calls and the remaining resources such as server memory and CPU, the maximum number of connections could be doubled to 200 (the second value).

[0077] When implementing this fault recovery strategy, the connection pool parameters of the faulty instance can be dynamically updated, and the new connection pool size will take effect without restarting the service instance. In this way, the number of connections that the faulty instance can manage is increased, enabling it to cope with the current connection request pressure, thereby alleviating the fault caused by connection overload, restoring the normal operation of the faulty node, avoiding the waste of resources caused by blind expansion, and ensuring the stability and efficiency of the entire business system.

[0078] In one implementation, the call chain includes multiple service nodes, each service node deploys at least one service instance, and after increasing the connection pool size of the faulty instance from the current first value to a second value, it also includes:

[0079] Returning to the steps of obtaining the status data of each service instance through pre-set instrumentation points in each service instance in the call chain, and determining whether the faulty instance has recovered to normal operation based on the latest obtained status data;

[0080] If normal operation is restored, the connection pool size of the faulty instance will be restored from the second value to the first value;

[0081] If normal operation is not restored, the service node where the faulty instance is located will be designated as the faulty node, and the other service instances in the faulty node will be designated as candidate instances. The timeout of the candidate instances will be reduced from the current third value to the fourth value. The timeout is the maximum waiting time for the candidate instances to communicate with other service instances outside the faulty node.

[0082] In this implementation, after increasing the connection pool size of the faulty instance from the current first value to the second value, the verification and adjustment phase can begin. At this point, the process returns to the step of obtaining the status data of each service instance in the call chain. Using the latest obtained status data, a determination is made as to whether the faulty instance has resumed normal operation.

[0083] If the assessment results indicate that the faulty instance has returned to normal operation, this means that the previous operation of increasing the connection pool size successfully resolved the fault and was not a long-term necessity. Therefore, the connection pool size of the faulty instance can be restored from the second value to the first value. This avoids the connection pool being maintained at an excessively large state for an extended period, reducing resource waste. Furthermore, too many idle connections consume system memory and port resources, which can negatively impact the overall performance of the service instance. Therefore, restoring the initial configuration after the fault is resolved aligns with the principle of rational resource utilization.

[0084] If the latest status data indicates that the faulty instance has not returned to normal operation, it means that simply increasing the connection pool size has not effectively solved the problem. In this case, the timeout period of the candidate instance can be reduced from the current third value to the fourth value. The candidate instance is another service instance in the faulty node where the faulty instance is located.

[0085] Timeout refers to the maximum waiting time set by a service instance when calling other service instances or processing requests. If no response is received or processing is completed within this time, it is considered a timeout, thus avoiding resource waste or process blockage caused by continuous waiting.

[0086] Therefore, by reducing the timeout of these candidate instances from the current third value to the fourth value, the waiting time of the candidate instances can be shortened, the resource consumption caused by the candidate instances waiting can be reduced, and their processing efficiency can be improved. This will allow the candidate instances to release resources more quickly, prioritize the processing of new valid requests, reduce the overall load pressure on the faulty nodes, and avoid affecting the faulty instances that have been processed by expanding the connection pool, which will help the faulty instances to resume normal operation as soon as possible.

[0087] In one implementation, after reducing the timeout period of the candidate instance from the current third value to the fourth value, it also includes:

[0088] Returning to the steps of obtaining the status data of each service instance through pre-set instrumentation points in each service instance in the call chain, and determining whether the faulty instance has recovered to normal operation based on the latest obtained status data;

[0089] If normal operation is restored, the connection pool size of the faulty instance will be restored from the second value to the first value, and the timeout time of the candidate instance will be restored from the fourth value to the third value.

[0090] If normal operation is not restored, the circuit breaker failure instance will continue to call downstream service nodes in its call chain.

[0091] In this implementation, after reducing the timeout of the candidate instance from the current third value to the fourth value, the status of the faulty instance can continue to be monitored. At this point, the process returns to the step of obtaining the status data of each service instance in the call chain, and based on the latest obtained data, it is determined whether the faulty instance has returned to normal operation.

[0092] If the assessment result indicates that the faulty instance has recovered and is running normally, this means that the previous operation of narrowing down the timeout time of candidate instances was effective. To restore the system resource configuration to its optimal state before the fault occurred, two recovery operations can be performed: first, restore the connection pool size of the faulty instance from the previously expanded second value back to the original first value; second, restore the timeout time of candidate instances from the narrowed fourth value back to the initial third value. This eliminates the additional consumption of system resources caused by temporary adjustments, restores the initial running configuration of the service nodes, and reserves reasonable handling space for subsequent business fluctuations.

[0093] If the latest status data indicates that the faulty instance has not yet returned to normal operation, it means that previous measures have failed to effectively resolve the issue. In this case, to prevent the fault from spreading further and affecting the stability of the entire call chain, a circuit breaker mechanism can be implemented. This means blocking calls from the faulty instance to downstream instances in the call chain. In other words, it can prevent the faulty instance from continuously requesting downstream instances, thus avoiding it from continuing to occupy connection resources or spreading abnormal pressure upstream and downstream.

[0094] For example, if a faulty instance calls the interface of downstream instance C and connection timeouts continue to occur, causing the entire call chain to become blocked, the service instance will stop calling instance C after the circuit breaker is applied, and instead return a pre-defined degraded response. This quickly releases the occupied connection resources and prevents the faulty instance from dragging down the entire call chain. Simultaneously, the circuit breaker operation will trigger a system alarm, prompting operations personnel to intervene and investigate deeper issues, such as code logic defects in the faulty instance or fundamental bottlenecks in downstream services, providing a basis for thoroughly resolving the fault.

[0095] In this way, the faulty instance is isolated from the downstream instance, preventing the fault from spreading and ensuring the normal operation of other parts of the call chain. It also buys time and space for further investigation and resolution of the fault.

[0096] In one implementation, based on the fault type, a fault repair strategy for the fault instance is determined and executed, including:

[0097] Determine the operational status of the faulty instance;

[0098] Based on the fault type and operational status, determine and implement fault repair strategies for fault instances.

[0099] In this implementation, once the fault type of a fault instance is identified, when further determining and executing the fault repair strategy for the fault instance, it is also necessary to comprehensively consider the running status of the fault instance, determine the specific fault repair strategy, and execute it. In other words, the same fault type may require completely different handling methods under different running states.

[0100] The operational status of faulty instances includes, but is not limited to, overall resource utilization (such as CPU, memory, and disk I / O utilization), service instance health status (such as the number of non-faulty instances and response latency), communication quality with upstream and downstream services (such as network latency and packet loss rate), and current business traffic characteristics (such as whether there is a sudden increase in request volume or whether it is during peak traffic periods).

[0101] For example, if the fault type is connection overload, and the operational status shows that the instance resources of the faulty node are sufficient (low CPU / memory usage) and the non-faulty instances are healthy, it indicates that the faulty node has room for expansion. The fault repair strategy can be to temporarily increase the connection pool size of the faulty instance, while appropriately expanding the capacity of other service instances to quickly handle more connections. If the operational status shows that the instance resources of the faulty node are saturated (CPU is consistently at 100%) and the non-faulty instances are also experiencing response delays, then blindly expanding the connection pool can exacerbate resource contention. The fault repair strategy can be adjusted to first implement circuit breaking and degradation on some non-core interfaces of the faulty node to reduce connection consumption, while triggering the expansion of the faulty node to fundamentally alleviate resource pressure.

[0102] Alternatively, if the fault type is network communication fault, and the operation status shows that only the faulty instance in the faulty node has abnormal communication with downstream, while other service instances are normal, it indicates that the problem may be in the network configuration of the faulty instance. The fault repair strategy can be to isolate the faulty instance and restart the network components of the faulty instance. If the operation status shows that the latency of all service instances in the faulty node communicating with downstream is soaring, it may be an overall network link problem. The fault repair strategy can be adjusted to temporarily shorten the timeout time of all service instances in the faulty node to reduce connection waiting.

[0103] This approach, which combines fault type with operational status, enables more targeted repair strategies, improves repair efficiency, and more effectively restores faulty instances to normal operation, ensuring the stability and efficiency of the entire system.

[0104] In one implementation, determining the fault type of a fault instance includes:

[0105] The success rate of the interface for obtaining faulty instances;

[0106] If the interface success rate is lower than the first threshold, the failure type of the faulty instance is determined to be low interface success rate;

[0107] Therefore, determining the operational status of the faulty instance includes:

[0108] Get the number of requests to the faulty instance, and determine whether the candidate instance can support a number of calls greater than or equal to the number of requests. The candidate instance is the other service instance in the service node where the faulty instance is located, excluding the faulty instance.

[0109] Correspondingly, based on the fault type and operational status, a fault repair strategy for the faulty service instance is determined and implemented, including:

[0110] If it can be supported, remove the faulty instance;

[0111] If the system cannot support the circuit breaker, the calls from the faulty instance to the corresponding downstream service instance will be suspended. After the corresponding downstream service instance is scaled up, the calls from the faulty instance to the corresponding downstream service instance will be restored.

[0112] In this implementation, when determining the fault type of a faulty instance, the interface success rate of that instance is first obtained. The interface success rate is the ratio of the number of requests successfully processed by the faulty instance to the total number of requests received. When the interface success rate is lower than a pre-set first threshold, it means that the faulty instance is frequently failing during request processing and cannot stably complete normal service interactions. At this time, the fault type of the faulty instance can be determined as low interface success rate.

[0113] Therefore, when determining the operational status of a failed instance, the carrying capacity of other service instances in the service node where the failed instance resides can be assessed. Specifically, the number of requests currently directed to the failed instance is first counted, reflecting the current business pressure that the failed instance needs to handle. At the same time, other service instances in the service node where the failed instance resides are considered as candidate instances. The current load, processing performance, and resource reserves of these candidate instances are analyzed to determine whether they can support a number of calls greater than or equal to the number of requests, that is, whether the candidate instances as a whole have the capacity to handle the original request volume of the failed instance.

[0114] Furthermore, when determining and executing fault repair strategies based on the above fault types and operating conditions, if the candidate instance can support the corresponding number of requests, it means that the normal operation of the business can be guaranteed without relying on the fault instance. In this case, the fault instance will be removed from the call chain, so that all requests are automatically routed to the candidate instance, avoiding the impact of the low success rate of the fault instance on the overall business process.

[0115] If the candidate instance cannot support the corresponding number of requests, a circuit breaker should be implemented first to suspend the faulty instance's calls to the corresponding downstream service instances, preventing the faulty instance from continuously transmitting erroneous requests downstream and causing a chain reaction of problems. Subsequently, the corresponding downstream service instances should be scaled up by increasing the number of service instances and improving resource configuration to enhance their processing capacity. After the downstream service instances have been scaled up and reached a stable state, calls from the faulty instance to them should be restored. In this way, the normal operation of the overall call chain can be gradually restored while ensuring the stability of downstream services.

[0116] Figure 2 This is a schematic diagram of a fault handling system according to this application. The system comprises four parts: a monitoring center, a dispatch center, a strategy center, and a handling center, wherein:

[0117] Monitoring Center: By pre-setting instrumentation points for each service instance (such as service A, service B, and service C) in the call chain, the status data of each service instance is obtained. The status data includes, but is not limited to, the interface success rate, CPU utilization, number of connection pools, downstream interface success rate, memory utilization, etc. of each service instance in the service node. The status data that meets the pre-set abnormal conditions is filtered out as abnormal data and alarmed to the scheduling center.

[0118] Dispatch Center: Sends abnormal data collected by the monitoring center to the policy center;

[0119] Strategy Center: Based on the status data, it parses the call relationship between service instances and performs fault correlation analysis on the call chain based on the status data and call relationship, including but not limited to traffic analysis, load analysis and middleware analysis, so as to achieve problem attribution, identify the service instance that has failed as the fault instance, then determine the fault type of the fault instance, and determine the fault repair strategy for the fault instance based on the fault type.

[0120] Dispatch Center: Based on the fault repair strategy, matches handling measures with the handling center;

[0121] Disposal Center: Locates the corresponding processing tools and executes fault repair strategies for faulty instances, including but not limited to adjusting rate limiting, registry center, and timeout for faulty nodes or instances, to achieve self-recovery of the call chain.

[0122] As can be seen from the above, the technical solution provided by the embodiments of this disclosure obtains the status data of each service instance in the call chain by setting up instrumentation points in each service instance, then parses the call relationship between service instances and performs fault correlation analysis to identify faulty instances, and then determines and executes fault repair strategies based on the fault type. This process does not require manual investigation and coordination one by one. It can quickly locate faulty service instances and implement targeted repairs by means of automated processes, which greatly shortens the fault recovery time, reduces the business interruption time, enables the business system to run more stably, effectively improves the overall availability of the business system when handling complex functions, and ensures the continuity and efficiency of business.

[0123] Reference Figure 3 The diagram shows a structural schematic of a fault handling system according to this application, which may specifically include:

[0124] Monitoring center 201 is used to obtain the status data of each service instance by pre-setting the instrumentation points of each service instance in the call chain, wherein the call chain includes multiple service instances;

[0125] The scheduling center 202 is used to synchronize the status data to the policy center;

[0126] The strategy center 203 is used to parse the call relationship between the service instances according to the status data, and perform fault correlation analysis on the call chain based on the status data and the call relationship to identify the service instance that has failed as a fault instance; determine the fault type of the fault instance, and determine the fault repair strategy for the fault instance based on the fault type.

[0127] The dispatch center 202 is also used to synchronize the fault repair strategy to the processing center;

[0128] The processing center 204 is used to execute fault repair strategies for the faulty instances.

[0129] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0130] Figure 4 This is a block diagram of an electronic device for fault handling according to an exemplary embodiment, including a processor and a memory, wherein the memory is used to store a computer program; and the processor is used to execute the program stored in the memory.

[0131] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0132] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0133] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions, which can be executed by a processor of an electronic device to perform the above-described method. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0134] In an exemplary embodiment, a computer program product is also provided that, when run on a computer, enables the computer to implement the above-described method for handling comment faults.

[0135] Figure 5 This is a block diagram illustrating an apparatus 800 for fault handling according to an exemplary embodiment.

[0136] For example, device 800 can be a mobile phone, computer, digital broadcasting electronic device, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0137] Reference Figure 5 The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0138] Processing component 802 typically controls the overall operation of device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0139] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of this data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0140] Power supply component 807 provides power to various components of device 800. Power supply component 807 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 800.

[0141] Multimedia component 808 includes a screen that provides an output interface between the device 800 and the account. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the account. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data to be processed. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0142] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0143] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0144] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, changes in position of device 800 or a component of device 800, the presence or absence of contact between an account and device 800, orientation or acceleration / deceleration of device 800, and temperature changes of device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0145] Communication component 816 is configured to facilitate wired or wireless communication between device 800 and other devices. Device 800 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0146] In an exemplary embodiment, the apparatus 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described in the first and second aspects.

[0147] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of the device 800 to perform the above-described method. Optionally, for example, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.

[0148] In an exemplary embodiment, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the fault handling methods described in the above embodiments.

[0149] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The embodiments of the invention are intended to cover any variations, uses, or adaptations of the embodiments of the invention that follow the general principles of the embodiments of the invention and include common knowledge or customary techniques in the art not disclosed in the embodiments of the invention. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the embodiments of the invention are indicated by the following claims.

[0150] It should be understood that the embodiments of the present invention are not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from their scope. The scope of the embodiments of the present invention is limited only by the appended claims.

Claims

1. A fault handling method, characterized in that, include: By pre-setting instrumentation points for each service instance in the call chain, the status data of each service instance is obtained, and the call chain includes multiple service instances; The call relationship between the service instances is parsed based on the status data, and based on the status data and the call relationship, a fault correlation analysis is performed on the call chain to identify the service instance that has failed as a fault instance. Determine the fault type of the fault instance, and based on the fault type, determine and execute a fault repair strategy for the fault instance.

2. The method according to claim 1, characterized in that, The step of parsing the call relationship between the service instances based on the status data includes: Based on the status data, determine the upstream service instance and / or downstream service instance of the service instance, and obtain the calling relationship between the service instances; Based on the status data and the call relationship, the call chain is subjected to fault correlation analysis to identify the service instance that has failed, which is designated as a fault instance. Select status data that meets the pre-defined abnormal conditions as abnormal data; Based on the call relationship, root cause analysis is performed on the abnormal data to identify the service instance that caused the failure, which is then identified as the failure instance.

3. The method according to claim 1, characterized in that, Determining the fault type of the fault instance includes: Obtain the connection pool status of the faulty instance, where the connection pool status is the resource pool status used to establish connections between the faulty instance and other service instances. If the connection pool status meets the overload condition, the fault type of the fault instance is determined to be connection quantity overload; The step of determining and executing a fault repair strategy for the faulty node based on the fault type includes: In the case where the fault type is connection overload, the connection pool size of the fault instance will be increased from the current first value to the second value.

4. The method according to claim 3, characterized in that, The call chain includes multiple service nodes, and at least one service instance is deployed on each service node. After increasing the connection pool size of the faulty instance from the current first value to the second value, the process further includes: Returning to the step of obtaining the status data of each service instance by pre-setting the instrumentation points of each service instance in the call chain, based on the latest obtained status data, it is determined whether the faulty instance has recovered to normal operation; If normal operation is restored, the connection pool size of the faulty instance will be restored from the second value to the first value; If normal operation is not restored, the service node where the faulty instance is located will be designated as the faulty node, and other service instances in the faulty node other than the faulty instance will be designated as candidate instances. The timeout time of the candidate instances will be reduced from the current third value to the fourth value. The timeout time is the maximum waiting time when the candidate instance communicates with other service instances other than the faulty node.

5. The method according to claim 4, characterized in that, After reducing the timeout period of the candidate instance from the current third value to the fourth value, the method further includes: Returning to the step of obtaining the status data of each service instance by pre-setting the instrumentation points of each service instance in the call chain, based on the latest obtained status data, it is determined whether the faulty service instance has recovered to normal operation; If normal operation is restored, the connection pool size of the faulty instance will be restored from the second value to the first value, and the timeout time of the candidate instance will be restored from the fourth value to the third value. If normal operation is not restored, the circuit breaker will suspend calls from the faulty instance to downstream service nodes in its call chain.

6. The method according to claim 1, characterized in that, The step of determining and executing a fault repair strategy for the fault instance based on the fault type includes: Determine the operational status of the faulty instance; Based on the fault type and the operating status, a fault repair strategy for the fault instance is determined and executed.

7. The method according to claim 6, characterized in that, Determining the fault type of the fault instance includes: Obtain the interface success rate of the faulty instance; If the interface success rate is lower than a first threshold, the failure type of the fault instance is determined to be low interface success rate; Determining the operational status of the faulty instance includes: The number of requests to the faulty instance is obtained, and it is determined whether a candidate instance can support a number of calls greater than or equal to the number of requests. The candidate instance is another service instance in the service node where the faulty instance is located, excluding the faulty instance. The step of determining and executing a fault repair strategy for the faulty service instance based on the fault type and the operational status includes: If the faulty instance can be resolved, remove it. If the system cannot support the failure, the system will suspend the calls from the failed instance to the corresponding downstream service instance. After scaling up the corresponding downstream service instance, the system will resume the calls from the failed instance to the corresponding downstream service instance.

8. A fault handling system, characterized in that, include: The monitoring center is used to obtain the status data of each service instance by pre-setting the instrumentation points of each service instance in the call chain, which includes multiple service instances; The scheduling center is used to synchronize the status data to the policy center; The strategy center is used to parse the call relationship between the service instances based on the status data, and to perform fault correlation analysis on the call chain based on the status data and the call relationship, and to identify the service instance that has failed as a fault instance. Determine the fault type of the fault instance, and based on the fault type, determine the fault repair strategy for the fault instance; The dispatch center is also used to synchronize the fault repair strategy to the processing center; The processing center is used to execute fault repair strategies for the faulty instances.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the fault handling method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the fault handling method as described in any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the fault handling method according to any one of claims 1 to 7.