Fault processing method and device
By receiving fault reporting events in the microservice environment and matching trigger events, sending notifications to different system services in the same service cluster, actively degrading across systems services is solved, and the problem of failure status cannot be shared simultaneously among multiple applications/systems in the prior art, ensuring the availability of services.
Patent Information
- Application Number
- CN202410039046.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the fuse and downgrade processing of fault services is mainly aimed at scenarios in a single application or in the same system. The fault status cannot be shared simultaneously between multiple applications/systems, resulting in the failure cannot be sensed by the associated services and are affected.
By receiving service fault reporting events, matching pre-configured fault trigger events, and sending fault trigger event notifications to different system services in the same service cluster, actively downgrade processing of cross-system services.
It effectively avoids the impact of faulty services on associated services, realizes fault status synchronization and active downgrade of cross-system services, and ensures availability between services.
Smart Images

Figure CN120295818A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a method and apparatus for fault handling. Background Art
[0002] In the field of microservices technology, when a downstream service fails, in order to ensure the availability of other related services associated with the downstream service, service fusing and service degradation are usually used to shield the impact of the downstream faulty service on other related services.
[0003] In the process of implementing the present invention, the inventors found the following problems in the prior art:
[0004] The current fusing and degradation handling of faulty services is mainly for single applications or scenarios within the same system, and it is impossible to synchronously share the faulty state of services among services in multiple applications / systems, resulting in related services not within the same application / system being affected by the faulty service because they cannot perceive the occurrence of the fault. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method and apparatus for fault handling, which implement fault handling that supports active degradation of cross-system services and effectively avoid the impact of faulty services on related services.
[0006] To achieve the above object, according to one aspect of embodiments of the present invention, a method for fault handling is provided, including:
[0007] In response to receiving a service fault reporting event of a first service, matching the service fault reporting event with a fault triggering event in a pre-configured fault scenario;
[0008] In the case of a successful match, according to the matched fault triggering event, sending a fault triggering event notification to the service cluster where the service fault reporting event is located, so that a second service related to the first service performs degradation processing according to the fault triggering event notification, where the second service and the first service are services of different systems in the same service cluster.
[0009] Optionally, before matching the service fault reporting event with the fault triggering event in the pre-configured fault scenario, the method further includes: determining each fault triggering event, the corresponding fault degradation plan, and the fault recovery strategy in a scripted manner through a visual interface to obtain a fault scenario.
[0010] Optionally, the fault scenario includes a service fault scenario of at least one service; before sending a fault trigger event notification to the service cluster where the service fault reporting event is located, the method further includes: in response to receiving a fault scenario synchronization request initiated by each service in the service cluster, selecting a valid fault scenario synchronization request from the fault scenario synchronization requests initiated by the each service according to the service fault scenario included in the fault scenario; synchronously distributing the fault scenario to the target service corresponding to the valid fault scenario synchronization request with the service fault scenario as the dimension, so that the target service performs degradation processing according to the service fault scenario therein and the received fault trigger event notification.
[0011] Optionally, the second service has a corresponding service fault scenario; the second service related to the first service performs degradation processing according to the fault trigger event notification, including: finding a fault degradation plan that matches the fault trigger event notification from the service fault scenario corresponding to the second service; performing fault degradation processing according to the fault degradation plan.
[0012] Optionally, the method further includes: when the service fault reporting event of the first service reaches a preset time limit, sending a fault recovery probe notification to the second service, so that the second service performs recovery of the degradation processing according to the received fault recovery probe notification.
[0013] Optionally, the sending of the fault trigger event notification and the fault recovery probe notification adopts a message queue mechanism, and the sending status of the fault trigger event notification and the fault recovery probe notification is determined by receiving an acknowledgment message from the second service.
[0014] Optionally, when the service fault reporting event of the first service reaches a preset time limit, sending a fault recovery probe notification to the second service includes: when the service fault reporting event of the first service reaches a preset time limit, sending a first fault recovery probe notification to the second service, and listening to the fault traffic corresponding to the service fault reporting event, dynamically adjusting the probe ratio in the first fault recovery probe notification according to the fault traffic; when the fault traffic reaches a preset recovery threshold, gradually increasing the probe ratio, and based on the increased probe ratio, sending a second fault recovery probe notification to the second service until the fault traffic does not exist.
[0015] Optionally, each service in the service cluster reports a service fault event, receives a fault trigger event notification, and receives a fault recovery probe notification through a software development kit.
[0016] According to a second aspect of an embodiment of the present invention, there is provided a fault processing apparatus, including:
[0017] A fault matching module, configured to, in response to receiving a service fault reporting event of a first service, match the service fault reporting event with a fault triggering event in a pre-configured fault scenario;
[0018] A service degradation module, configured to, when the matching is successful, send a fault triggering event notification to the service cluster where the service fault reporting event is located according to the matched fault triggering event, so that a second service related to the first service performs a degradation process according to the fault triggering event notification, where the second service and the first service are services of different systems in the same service cluster.
[0019] According to a third aspect of an embodiment of the present invention, there is provided an electronic device for fault handling, including:
[0020] One or more processors;
[0021] A storage device, configured to store one or more programs,
[0022] When the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the first aspect of the embodiment of the present invention.
[0023] According to a fourth aspect of an embodiment of the present invention, there is provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method provided in the first aspect of the embodiment of the present invention is implemented.
[0024] One embodiment of the present invention has the following advantages or beneficial effects: By responding to receiving a service fault reporting event of a first service, matching the service fault reporting event with a fault triggering event in a pre-configured fault scenario; when the matching is successful, sending a fault triggering event notification to the service cluster where the service fault reporting event is located according to the matched fault triggering event, so that a second service related to the first service performs a degradation process according to the fault triggering event notification, where the second service and the first service are services of different systems in the same service cluster, a complete fault handling method that supports active degradation of cross-system services is implemented, effectively avoiding the impact of the faulty first service on the associated second service. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0026] Figure 1 is a schematic diagram of the main process of the method for fault handling according to an embodiment of the present invention;
[0027] Figure 2It is a schematic diagram of the event scheduling service and service relationship in the payment scenario of the embodiment of the present invention;
[0028] Figure 3 It is a schematic flowchart of the service degradation processing in the payment scenario of the embodiment of the present invention;
[0029] Figure 4 It is a schematic flowchart of the sending process of the fault trigger event notification in the embodiment of the present invention;
[0030] Figure 5 It is a schematic flowchart of the complete process of fault handling in the embodiment of the present invention;
[0031] Figure 6 It is a schematic flowchart of the fault recovery in the payment scenario of the embodiment of the present invention;
[0032] Figure 7 It is a schematic diagram of the overall architecture of the method for fault handling in the embodiment of the present invention;
[0033] Figure 8 It is a schematic diagram of the main modules of the device for fault handling according to the embodiment of the present invention;
[0034] Figure 9 It is an exemplary system architecture diagram to which the embodiment of the present invention can be applied;
[0035] Figure 10 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiment of the present invention. Detailed implementation manners
[0036] It should be noted that in the technical solutions of the present disclosure, the acquisition, storage, application, etc. of the user's personal information all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0037] The following makes an explanation of the exemplary embodiments of the present invention with reference to the accompanying drawings. Among them, various details of the embodiments of the present invention are included to help understanding, and they should be considered only exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described here without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0038] The current fuse and degradation processing of the faulty service are mainly for single applications or scenarios within the same system, and it is impossible to synchronously share the faulty state of the service among services in multiple applications / systems. As a result, the associated services that are not in the same application / system cannot perceive the occurrence of the fault and are thus affected by the faulty service, which cannot well meet the actual application.
[0039] To solve the above problems existing in the prior art, the present invention proposes a method for fault handling. When a fault trigger event of a first service is received, a fault trigger event notification is sent to a service cluster to enable a second service related to the first service to perceive the fault of the first service, so that the second service can shield the impact of the fault of the first service on it through degradation processing. A fault handling method that supports active degradation of cross-system services is implemented, effectively avoiding the impact of faulty services on associated services.
[0040] Figure 1 It is a schematic diagram of the main process of the fault handling method according to an embodiment of the present invention. As Figure 1 shown, the fault handling method of the embodiment of the present invention includes the following steps S101 to S103.
[0041] Step S101, in response to receiving a service fault reporting event of a first service, match the service fault reporting event with a fault trigger event in a pre-configured fault scenario.
[0042] Specifically, in a microservices business scenario, when a fault occurs in the interaction link of one of the services, the service in which the fault occurs is defined as the first service. The first service collects fault traffic. When the fault traffic meets a preset fuse threshold, a fuse event of the first service is triggered, and a corresponding service fault event is defined, and the service fault event of the first service is reported to the event scheduling service. The event scheduling service is responsible for handling service fault events of all access services. After receiving the service fault reporting event of the first service, the event scheduling service performs event matching between the service fault reporting event and the fault trigger event in the pre-configured fault scenario to find out whether there is such a service fault reporting event in the configured fault scenario.
[0043] According to an embodiment of the present invention, before matching the service fault reporting event with the fault trigger event in the pre-configured fault scenario, the method further includes: through a visual interface, in a scripted manner, determine each fault trigger event and the corresponding fault degradation plan and fault recovery strategy to obtain a fault scenario.
[0044] Specifically, before matching the service failure reporting event with the failure triggering event in the pre-configured failure scenario, the business administrator needs to clarify the failure triggering event that triggers the failure handling, as well as the corresponding failure degradation plan and failure recovery strategy. The failure degradation plan includes the linkage degradation plan of all related services in the failure triggering event; the failure recovery strategy refers to the traffic recovery strategy of each associated service during the service failure recovery stage, specifically including the switching compensation of traffic and the setting of specific indicators such as the recovery speed. The business administrator edits the failure scenario in a scripted manner through a visual interface and translates the edited configuration into an explanatory language for the program to execute, so that the failure scenario can be configured in a friendly and efficient manner. Based on the configured failure scenario, according to the switching of the service failure reporting events of each first service that has a failure, the event scheduling service has the ability to schedule each related associated service to perform automated active degradation and degradation recovery.
[0045] Understandably, in order to clearly define the logic of failure handling, the configuration of the failure scenario can be implemented by the scenario management service. When the scenario management service receives the configured failure scenario, it saves the failure scenario in the database and at the same time pushes it to the event scheduling service, so that the event scheduling service can focus on handling and resolving various failed services.
[0046] Step S102, in the case of a successful match, according to the matched failure triggering event, send a failure triggering event notification to the service cluster where the service failure reporting event is located, so that the second service related to the first service can perform degradation processing according to the failure triggering event notification, where the second service and the first service are services of different systems in the same service cluster.
[0047] Specifically, the services accessed by the event scheduling service form a service cluster. In the microservice scenario, the service cluster usually includes multiple services, and these services may belong to different systems or different applications. Of course, the event scheduling service can also interface with multiple service clusters. When the service failure reporting event matches the failure triggering event in the configured failed service, the event scheduling service sends a failure triggering event notification to the service cluster where the service failure reporting event is located in a broadcast form, so that the second service that belongs to the same service cluster as the first service and is related to the first service receives the failure triggering event notification, perceives that the first service has a failure according to the failure triggering event notification, and actively degrades by executing the corresponding degradation processing to ensure that it is not affected by the failure of the first service.
[0048] Exemplarily, taking the payment scenario as an example, the service cluster includes a payment application, a payment service list, and a payment channel. The payment application is responsible for interacting with users and providing payment service list query and payment functions; the payment service list returns the list of bank services available to the current user according to the user's bank binding situation; the payment channel is a channel agent that maintains the interaction link with the specific bank payment service. When a payment channel of a certain bank fails, the payment channel will actively fuse the call to the faulty channel and at the same time send a service fault reporting event of the fuse to the event scheduling service. After receiving the service fault reporting event, the event scheduling service matches the payment channel fuse fault trigger event and sends a payment channel fuse fault trigger event notification to the service cluster. In this way, the second services (payment service list and payment application) in the cluster that have subscribed to the fault trigger event sense the fault of a certain bank in the payment channel, and then remove the service items of that bank in this service, avoiding the problem that users cannot complete the payment because they use the faulty service.
[0049] Correspondingly, for the payment scenario in the embodiment of the present invention, the fault trigger event is that the bank link of the payment channel is fused; the corresponding fault degradation plan is that the payment service list suspends the display of the faulty bank in the service list, and the payment application prompts the user with the words of service exception of the faulty bank; the fault recovery strategy is that the payment service list releases the list display of the faulty bank in proportion, and the payment application opens the payment function of the faulty bank in proportion. According to an embodiment of the present invention, the fault scenario includes a service fault scenario of at least one service; before sending a fault trigger event notification to the service cluster where the service fault reporting event is located, the method further includes: in response to receiving a fault scenario synchronization request initiated by each service in the service cluster, selecting a valid fault scenario synchronization request from the fault scenario synchronization requests initiated by each service according to the service fault scenario included in the fault scenario; synchronously sending the fault scenario to the target service corresponding to the valid fault scenario synchronization request in terms of the service fault scenario, so that the target service performs degradation processing according to the service fault scenario therein and the received fault trigger event notification.
[0050] According to an embodiment of the present invention, the fault scenario includes a service fault scenario of at least one service; before sending a fault trigger event notification to the service cluster where the service fault reporting event is located, the method further includes: in response to receiving a fault scenario synchronization request initiated by each service in the service cluster, selecting a valid fault scenario synchronization request from the fault scenario synchronization requests initiated by each service according to the service fault scenario included in the fault scenario; synchronously distributing the fault scenario to the target service corresponding to the valid fault scenario synchronization request in terms of the service fault scenario, so that the target service can perform a downgrade process according to the service fault scenario therein and the received fault trigger event notification.
[0051] Specifically, the fault scenario in the embodiment of the present invention is usually a fault downgrade plan and a fault recovery process for multiple associated services in a linkage. Therefore, when the fault scenario is finely divided in terms of services, the fault scenario includes a service fault scenario of at least one service. Considering that the service fault scenarios of each service may be updated in real time, in order to maintain the synchronization of the service fault scenarios, each service in the service cluster maintains a long connection with the heartbeat communication signal and the event scheduling service. Before sending a fault trigger event notification to the service cluster where the service fault reporting event is located, the event scheduling service receives the fault scenario synchronization requests initiated by each service in the service cluster, that is, the heartbeat communication signals, selects valid fault scenario synchronization requests from the service cluster according to the service fault scenarios involved in the fault scenario; and then synchronously distributes the fault scenario to the corresponding target services in terms of the service fault scenario, realizing that the services in the service cluster have the latest and valid service fault scenarios, so that subsequent active downgrade processing of faults can be performed according to the received fault trigger event notification and their own service fault scenarios.
[0052] According to another embodiment of the present invention, the second service has a corresponding service fault scenario; the second service related to the first service performs a downgrade process according to the fault trigger event notification, including: searching for a fault downgrade plan that matches the fault trigger event notification from the service fault scenario corresponding to the second service; performing a fault downgrade process according to the fault downgrade plan.
[0053] Specifically, the second service in the service cluster receives the corresponding fault trigger event notification by subscribing to the fault trigger event, searches for a fault downgrade plan that matches the fault trigger event of the fault trigger event notification from the local service fault scenario, and actively performs a downgrade process for related faults according to the found fault downgrade plan.
[0054] According to still another embodiment of the present invention, each service in the service cluster reports a service fault event through a software development kit and receives a fault trigger event notification.
[0055] Specifically, each service in the service cluster of the embodiments of the present invention reports service failure events to the event scheduling service by accessing a dedicated SDK (Software Development Kit), and receives failure trigger event notifications from the event scheduling service. The SDK not only provides the logic for communicating with the event scheduling service, but also can update and expand the logic therein according to the actual needs of the business to better meet the actual needs.
[0056] Figure 2 It is a schematic diagram of the relationship between the event scheduling service and services in the payment scenario of the embodiments of the present invention. The payment application service, payment service list service, and payment channel service in the service cluster establish connections with the event scheduling service through their respective SDKs and access the event scheduling service. The solid lines in the figure are operation flows related to the business, and the dashed lines are information flows related to failures. The user obtains the payment service list service through the payment application service, and selects the payment channel service of the corresponding bank from the payment service list service for payment; the administrator configures the failure scenario through the scenario management service and distributes it to the event scheduling service. The event scheduling service synchronizes the failure status through the SDKs of each service, and each service pulls the corresponding service failure scenario from the event scheduling service through its own SDK.
[0057] Figure 3 It is a schematic flow diagram of service degradation processing in the payment scenario of the embodiments of the present invention. In the figure, a failure occurs that the payment interface of Bank A is unavailable. The payment channel service monitors the abnormal availability rate, triggers the circuit breaker degradation of the Bank A interface of the payment channel service, and reports the circuit breaker event of the Bank A interface of the payment channel service as a service failure event to the event scheduling service. The event scheduling service finds a matching failure trigger event in the failure scenario and sends a failure trigger event notification to the relevant second services (payment application service and payment service list); the payment application service sends a message about the abnormal payment of Bank A to the user according to the corresponding failure degradation plan therein, and the payment service list deletes Bank A from the list according to the corresponding failure degradation plan therein.
[0058] According to another embodiment of the present invention, the sending of the failure trigger event notification adopts a message queue mechanism, and the sending status of the failure trigger event notification is determined by receiving the response message of the second service.
[0059] Specifically, to prevent the loss of fault trigger event notifications during transmission, the fault trigger event notifications sent by the event scheduling service in the embodiments of the present invention use the message queue MQ mechanism. When the second service subscribing to the fault trigger event receives the message of the fault trigger event notification, it will send an acknowledgement message to the event scheduling service to determine that the fault trigger event notification has been successfully sent. In addition, in the embodiments of the present invention, the entire process of the degradation processing and degradation recovery of each second service triggered by the fault reporting event in the service cluster is recorded in the form of a fault work order. When the acknowledgement message of the trigger event notification sent by the second service is received by the event scheduling service, the status of the corresponding item in the fault work order will also be updated in real time.
[0060] Figure 4 It is a schematic diagram of the sending process of the fault trigger event notification in the embodiments of the present invention. The first service (Service A) reports a service fault event. When the event scheduling service matches the fault trigger event, it sends the fault trigger event notification to each service in the service cluster in a broadcast manner through the message queue MQ. The second service (Service B) subscribes to the fault trigger event, so it receives the fault trigger event notification and sends the received acknowledgement message to the event scheduling service, so that the event scheduling service can update the corresponding fault work order according to the received acknowledgement message.
[0061] According to another embodiment of the present invention, the method further includes: when the service fault reporting event of the first service reaches a preset time limit, sending a fault recovery probe notification to the second service, so that the second service can perform the recovery of the degradation processing according to the received fault recovery probe notification.
[0062] Specifically, in the embodiments of the present invention, a silent period for service faults, that is, a time limit, is set. When the service fault reporting time reaches the preset time limit, a probe notification for attempting fault recovery is sent to the second service. After receiving the fault recovery probe notification, the second service determines specific recovery actions according to the fault recovery probe notification and the corresponding fault recovery strategy in the fault scenario, and performs the recovery of the degradation processing, so as to achieve that when the faulty service is restored, the second service also adaptively performs the recovery of the degradation processing. It can be understood that the event scheduling service can also send the fault recovery probe notification to the service cluster where the service fault reporting event is located in a broadcast manner, so that the second service can obtain the fault recovery probe notification through subscribing to the fault trigger event to perform the recovery of the degradation processing.
[0063] According to an embodiment of the present invention, when the service failure reporting event of the first service reaches a preset time limit, a failure recovery probe notification is sent to the second service, including: when the service failure reporting event of the first service reaches the preset time limit, a first failure recovery probe notification is sent to the second service, and the failure traffic corresponding to the service failure reporting event is monitored, and the probe ratio in the first failure recovery probe notification is dynamically adjusted according to the failure traffic; when the failure traffic reaches a preset recovery threshold, the probe ratio is gradually increased, and based on the increased probe ratio, a second failure recovery probe notification is sent to the second service until the failure traffic does not exist.
[0064] Specifically, when the duration of the service failure reporting event reaches the preset time limit, the event scheduling service sends a first failure recovery probe notification to the above-mentioned second service, allowing the second service to start attempting to perform the recovery of the degradation process. When the second service receives the first failure recovery probe notification, it determines the recovery action according to the failure recovery strategy in the failure scenario, and then combines the probe ratio in the first failure recovery probe notification to tentatively perform the failure recovery. For example, the second service in the payment scenario mentioned above: the payment service list. When the payment service list receives the first failure recovery probe notification, according to the failure recovery strategy and the probe ratio of 5% in the first failure recovery probe notification, it determines to add the bank that had a failure before to the payment service list for 5% of the users. At the same time, the failure traffic corresponding to the first service where the service failure reporting event is located is monitored, and the probe ratio in the first failure recovery probe notification is dynamically adjusted according to the change state of the failure traffic. For example, if the failure traffic is decreasing, it means that the failing first service is gradually recovering. At this time, the probe ratio in the first failure recovery probe notification can be increased, so that the second service can adaptively and dynamically perform the degradation recovery.
[0065] Furthermore, when the failure traffic reaches the preset recovery threshold, it means that the failing service is about to return to normal. At this time, the probe ratio is increased, and based on the increased probe ratio, a second failure recovery probe notification is sent to the second service to further perform the recovery of the degradation process. According to the above recovery process, the probe ratio is gradually increased until the failure traffic no longer exists, indicating that the failing service has returned to normal. At this time, the probe ratio is also increased to 100% accordingly, that is, the second service has returned to the normal call state, and thus a complete set of fault handling processes for active service degradation is completed.
[0066] Figure 5It is a schematic diagram of the complete process for fault handling in an embodiment of the present invention. When a circuit breaker fault event occurs in the first service (Service A), the circuit breaker service fault event is uploaded to the event scheduling service. When the event scheduling service matches the corresponding fault trigger event, it sends a fault trigger event notification. The second service (Service B) subscribes to this fault trigger event, receives the fault trigger event notification, and starts the degradation process. When the service fault reporting event reaches the preset time limit, the event scheduling service sends a first fault recovery probe notification. Service B starts the tentative degradation recovery based on the received first fault recovery probe notification and monitors the fault traffic corresponding to Service A where the circuit breaker fault event is located. When the fault traffic reaches the preset recovery threshold, it indicates that the first service (Service A) is about to recover from the fault. At this time, a second fault recovery probe notification is sent to gradually increase the probe ratio until the fault traffic disappears and the fault of Service A is completely recovered. Correspondingly, the probe ratio of the second service (Service B) also reaches 100%, and the call to Service A is adaptively restored.
[0067] According to another embodiment of the present invention, each service in the service cluster receives the fault recovery probe notification through a software development kit.
[0068] Specifically, each service in the service cluster of the embodiment of the present invention receives the fault recovery probe notification by accessing a dedicated SDK software development kit. The SDK software development kit not only provides the logic for communicating with the event scheduling service, but also can update and expand the logic therein according to the actual needs of the business to better meet the actual needs.
[0069] According to another embodiment of the present invention, the sending of the fault recovery probe notification adopts a message queue mechanism, and the sending status of the fault recovery probe notification is determined by receiving the response message of the second service.
[0070] Specifically, the sending mechanism of the fault recovery probe notification is similar to that of the above-mentioned fault trigger event, and both adopt a message queue mechanism, which can effectively avoid the loss of the fault trigger event notification during sending.
[0071] Figure 6It is a schematic flowchart of the fault recovery in the payment scenario of the embodiment of the present invention. When the service fault reporting event reaches the preset time limit, a first fault recovery probe notification is sent to the second service (payment service list and payment application service). After receiving it, the payment service list adds Bank A to the list according to the probe ratio. The payment application service displays the Bank A option according to the probe ratio and monitors the fault traffic reported by Bank A in the first service (payment channel service) in real time. When the fault traffic reaches the preset recovery threshold, the probe ratio is gradually increased. Based on the increased probe ratio, a second fault recovery probe notification is sent to the second service. The payment service list and the payment application service also continue to increase the probe ratio until the fault traffic disappears, the Bank A fault is recovered, and the first service fault is recovered. At this time, the payment service list and the payment application service also correspondingly return to the normal state of calling the payment channel of Bank A.
[0072] Figure 7 It is a schematic diagram of the overall architecture of the fault handling method in the embodiment of the present invention. The business administrator configures the fault scenario through the scenario management service, stores the fault scenario in the database, and distributes it to the event scheduling service. The event scheduling service obtains the service fault events through the SDKs of each service in the connected service cluster. The fault trigger event notification and the fault recovery probe notification are synchronously sent through the synchronous message queue. Each service obtains the corresponding fault trigger event notification and fault recovery probe notification of this service through event subscription to implement the active degradation processing and adaptive degradation recovery of the service.
[0073] The event scheduling service in the embodiment of the present invention supports the access of services from different systems / applications, that is, the first service and the second service are not in the same system / application. The event scheduling service monitors the status of each service and schedules the relevant second service to perform active degradation processing according to the monitored fault trigger event notification of the first service. When the duration of the fault reaches the preset time limit, the degradation of the second service is adaptively recovered by sending a fault recovery probe notification to the second service, realizing a fault handling method for active degradation of services across systems / applications.
[0074] Figure 8 It is a schematic diagram of the main modules of the fault handling device according to the embodiment of the present invention. As Figure 8 shown, the fault handling device 800 mainly includes a fault matching module 801 and a service degradation module 802.
[0075] The fault matching module 801 is configured to match the service fault reporting event with the fault trigger event in the pre-configured fault scenario in response to receiving the service fault reporting event of the first service;
[0076] A service degradation module 802, which is used to, in case of successful matching, send a fault trigger event notification to the service cluster where the service fault reporting event is located according to the matched fault trigger event, so that a second service related to the first service performs degradation processing according to the fault trigger event notification, where the second service and the first service are services of different systems in the same service cluster.
[0077] According to an embodiment of the present invention, the fault handling device 800 further includes a fault scenario construction module (not shown in the figure), which is used to: before matching the service fault reporting event with the fault trigger events in the pre-configured fault scenarios, determine each fault trigger event, the corresponding fault degradation plan and fault recovery strategy in a scripted manner through a visual interface to obtain a fault scenario.
[0078] According to another embodiment of the present invention, the fault scenario includes service fault scenarios of at least one service; the fault handling device 800 further includes a fault scenario distribution module (not shown in the figure), which is used to: before sending a fault trigger event notification to the service cluster where the service fault reporting event is located, in response to receiving fault scenario synchronization requests initiated by each service in the service cluster, select valid fault scenario synchronization requests from the fault scenario synchronization requests initiated by each service according to the service fault scenarios included in the fault scenario; synchronously distribute the fault scenario to the target service corresponding to the valid fault scenario synchronization request in terms of service fault scenarios, so that the target service performs degradation processing according to the service fault scenarios therein and the received fault trigger event notification.
[0079] According to still another embodiment of the present invention, the second service has a corresponding service fault scenario; the service degradation module 802 is further used to: find a fault degradation plan that matches the fault trigger event notification from the service fault scenario corresponding to the second service; perform fault degradation processing according to the fault degradation plan.
[0080] According to yet another embodiment of the present invention, the fault handling device 800 further includes a fault recovery module (not shown in the figure), which is used to: in case the service fault reporting event of the first service reaches a preset time limit, send a fault recovery probe notification to the second service, so that the second service performs recovery from the degradation processing according to the received fault recovery probe notification.
[0081] According to another embodiment of the present invention, the sending of the fault trigger event notification and the fault recovery probe notification adopts a message queue mechanism, and the sending status of the fault trigger event notification and the fault recovery probe notification is determined by receiving the response message of the second service.
[0082] According to another embodiment of the present invention, the fault recovery module (not shown in the figure) is further configured to: when the service fault reporting event of the first service reaches a preset time limit, send a first fault recovery probe notice to the associated service, and monitor the fault traffic corresponding to the service fault reporting event, and dynamically adjust the probe ratio in the first fault recovery probe notice according to the fault traffic; when the fault traffic reaches a preset recovery threshold, gradually increase the probe ratio, and based on the increased probe ratio, send a second fault recovery probe notice to the second service until the fault traffic does not exist.
[0083] According to another embodiment of the present invention, each service in the service cluster reports service fault events through a software development kit, receives fault trigger event notifications, and receives fault recovery probe notices.
[0084] Figure 9 It is an exemplary system architecture diagram to which the embodiments of the present invention can be applied.
[0085] As Figure 9 shown, the system architecture 900 may include terminal devices 901, 902, 903, a network 904, and a server 905. The network 904 is used to provide a medium for communication links between the terminal devices 901, 902, 903 and the server 905. The network 904 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0086] Users can use the terminal devices 901, 902, 903 to interact with the server 905 through the network 904 to receive or send messages, etc. Various communication client applications, such as a fault handling application (only for example), may be installed on the terminal devices 901, 902, 903.
[0087] The terminal devices 901, 902, 903 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0088] The server 905 can be a server that provides various services, such as a background management server (merely an example) that supports the fault handling performed by the user using the terminal devices 901, 902, and 903. The background management server can match the service fault reporting event with the fault triggering event in the pre-configured fault scenario in response to receiving the service fault reporting event of the first service; in the case of successful matching, according to the matched fault triggering event, send a fault triggering event notification to the service cluster where the service fault reporting event is located, so that the second service related to the first service can perform downgrading processing according to the fault triggering event notification, where the second service and the first service are services of different systems in the same service cluster and other processing, and feedback the processing result to the terminal device.
[0089] It should be noted that the method for fault handling provided in the embodiments of the present invention is generally executed by the server 905. Correspondingly, the device for fault handling is generally arranged in the server 905.
[0090] It should be understood that Figure 9 the numbers of the terminal devices, networks, and servers in
[0091] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 10 is a schematic structural diagram of a computer system suitable for implementing the terminal device or server of the embodiments of the present invention. Figure 10 The terminal device or server shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0092] As Figure 10 shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage section 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the system 1000 are also stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0093] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as required. A removable medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1010 as required so that a computer program read therefrom is installed into the storage section 1008 as required.
[0094] Specifically, according to the embodiments disclosed by the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed by the present invention include a computer program product which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by a central processing unit (CPU) 1001, the above-described functions defined in the system of the present invention are executed.
[0095] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0096] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0097] The units involved in the embodiments of the present invention can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a fault matching module and a service degradation module.
[0098] Among them, the names of these modules do not constitute a limitation on the modules themselves in some cases. For example, the fault matching module can also be described as "a module for matching the service fault reporting event with the fault triggering event in a pre-configured fault scenario in response to receiving a service fault reporting event of a first service".
[0099] On the other hand, the present invention also provides a computer-readable medium, which can be included in the device described in the embodiments; or it can exist alone without being assembled into the device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: matching the service fault reporting event with the fault triggering event in a pre-configured fault scenario in response to receiving a service fault reporting event of a first service; in the case of successful matching, sending a fault triggering event notification to the service cluster where the service fault reporting event is located according to the matched fault triggering event, so that a second service related to the first service performs degradation processing according to the fault triggering event notification, where the second service and the first service are services of different systems in the same service cluster. According to the technical solution of the embodiments of the present invention, the following advantages or beneficial effects are achieved: by matching the service fault reporting event with the fault triggering event in a pre-configured fault scenario in response to receiving a service fault reporting event of a first service; in the case of successful matching, sending a fault triggering event notification to the service cluster where the service fault reporting event is located according to the matched fault triggering event, so that a second service related to the first service performs degradation processing according to the fault triggering event notification, where the second service and the first service are services of different systems in the same service cluster, a complete fault handling method for actively degrading cross-system services is realized, effectively avoiding the impact of the faulty first service on the associated second service.
[0100] The specific implementation manners do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for fault handling, characterized in that, Including: In response to receiving a service failure reporting event of a first service, matching the service failure reporting event with a failure triggering event in a pre-configured failure scenario; In the case of successful matching, according to the matched failure triggering event, sending a failure triggering event notification to the service cluster where the service failure reporting event is located, so that a second service related to the first service performs a downgrade process according to the failure triggering event notification, where the second service and the first service are services of different systems in the same service cluster.
2. The method according to claim 1, wherein Before matching the service failure reporting event with the failure triggering event in the pre-configured failure scenario, the method further includes: Through a visual interface, in a scripted manner, determining each failure triggering event and the corresponding failure downgrade plan and failure recovery strategy, to obtain a failure scenario.
3. The method according to claim 1, wherein The failure scenario includes service failure scenarios of at least one service; Before sending the failure triggering event notification to the service cluster where the service failure reporting event is located, the method further includes: In response to receiving failure scenario synchronization requests initiated by each service in the service cluster, according to the service failure scenarios included in the failure scenario, selecting valid failure scenario synchronization requests from the failure scenario synchronization requests initiated by each service; Synchronously distributing the failure scenario to the target service corresponding to the valid failure scenario synchronization request with the service failure scenario as the dimension, so that the target service performs a downgrade process according to the service failure scenario therein and the received failure triggering event notification.
4. The method according to claim 3, characterized in that, The second service has a corresponding service failure scenario; The second service related to the first service performing a downgrade process according to the failure triggering event notification includes: Searching for a failure downgrade plan that matches the failure triggering event notification from the service failure scenario corresponding to the second service; Performing a failure downgrade process according to the failure downgrade plan.
5. The method according to claim 1, wherein The method further includes: In the case where the service failure reporting event of the first service reaches a preset time limit, sending a failure recovery probe notification to the second service, so that the second service performs a recovery of the downgrade process according to the received failure recovery probe notification.
6. The method according to claim 5, wherein The sending of the failure triggering event notification and the failure recovery probe notification adopts a message queue mechanism, and determines the sending status of the failure triggering event notification and the failure recovery probe notification by receiving the response message of the second service.
7. The method according to claim 5, characterized in that, In the case where the service failure reporting event of the first service reaches a preset time limit, sending a failure recovery probe notification to the second service includes: In the case where the service failure reporting event of the first service reaches a preset time limit, sending a first failure recovery probe notification to the second service, and monitoring the failure traffic corresponding to the service failure reporting event, and dynamically adjusting the probe ratio in the first failure recovery probe notification according to the failure traffic; In the case where the failure traffic reaches a preset recovery threshold, gradually increasing the probe ratio, and based on the increased probe ratio, sending a second failure recovery probe notification to the second service until the failure traffic does not exist.
8. The method according to claim 5, characterized in that Each service in the service cluster reports service fault events through a software development kit, receives fault trigger event notifications, and receives fault recovery probe notifications.
9. A device for fault handling, characterized in that, It includes: A fault matching module, configured to, in response to receiving a service fault reporting event of a first service, match the service fault reporting event with a fault trigger event in a pre-configured fault scenario; A service degradation module, configured to, in case of successful matching, send a fault trigger event notification to the service cluster where the service fault reporting event is located according to the matched fault trigger event, so that a second service related to the first service performs degradation processing according to the fault trigger event notification, where the second service and the first service are services of different systems in the same service cluster.
10. A mobile electronic device terminal, characterized in that, It includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-8.
11. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1-8.