Methods and systems for automated intelligent transaction monitoring, circuit breaker triggering, and rapid recovery.
By using the dynamic threshold adjustment and automated recovery mechanism of the intelligent transaction monitoring system, the problems of misjudgment and recovery lag in the order transaction system under high concurrency environment are solved, achieving more efficient system stability and operation and maintenance efficiency.
Patent Information
- Application Number
- CN202511254747.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-04
AI Technical Summary
Existing order transaction systems are prone to misjudgment or delayed response in high-concurrency environments, lack automated recovery mechanisms, and have insufficient monitoring granularity, resulting in low system stability and low operational efficiency.
An intelligent transaction monitoring system based on dynamic thresholds is adopted, including data acquisition, intelligent anomaly detection, and dynamic circuit breaker recovery decision-making modules. By collecting transaction data in real time, the system dynamically adjusts indicator weights and health scores to achieve adaptive circuit breaker and automated recovery.
It improved the stability and operational efficiency of the order transaction system, reduced misjudgments and recovery time, and enhanced the system's self-healing capabilities and monitoring accuracy.
Smart Images

Figure CN120743612B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic transactions, and in particular to a distributed transaction system protection scheme based on automated intelligent monitoring, dynamic circuit breaker mechanism and fast recovery strategy, which is suitable for order transaction scenarios with high concurrency and high availability requirements, such as offline order collection and online transactions. Background Technology
[0002] In modern order trading systems, data services typically employ a distributed architecture to handle massive transaction requests from various locations. These systems are characterized by high concurrency, low latency, and multi-node collaborative operation, enabling the rapid completion of tens of thousands of orders per second. However, this complex system structure also presents new challenges: if a critical service malfunctions or becomes overloaded, it can easily trigger a chain reaction of failures, leading to performance degradation or even system paralysis. To address this, the industry commonly employs a "circuit breaker mechanism" as a risk control measure. Similar to a fuse in a circuit, which automatically cuts off power to prevent equipment damage when current is too high, the circuit breaker mechanism in order trading systems temporarily freezes the trading platform when order processing speed, price fluctuations, or transaction volume exceed safe limits. This prevents data errors or cascading crashes and allows regulators to intervene. Currently, most common circuit breaker methods are based on fixed thresholds, such as triggering a circuit breaker when a certain indicator (e.g., order failure rate, response time, CPU utilization) exceeds a preset value. While this approach is simple to implement, it has several significant drawbacks in practical applications:
[0003] 1. Static threshold limitation: unable to adapt to real-time changes, prone to misjudgment or delayed response.
[0004] Traditional circuit breaker mechanisms rely on pre-set fixed values as the basis for judgment, such as "circuit breaker is triggered if more than 50 orders fail in a single minute". However, the volatility of trading orders is very large, and the load difference between peak and off-peak hours is significant. Fixed thresholds often cannot accurately reflect the true risk status.
[0005] - When business volume surges, normal transactions may be misjudged as abnormal, causing unnecessary circuit breakers;
[0006] - When a real anomaly occurs, the system may miss the best response time and delay processing because it has not reached the set threshold.
[0007] Therefore, this static strategy lacks flexibility and is difficult to adapt to complex and ever-changing trading environments.
[0008] 2. Lack of Automatic Recovery Mechanism: After a failure in the order transaction system, timely recovery is necessary. The recovery process after an order system failure mainly includes steps such as fault detection, status restoration, fault-tolerant switching, and post-event optimization. However, traditional recovery mechanisms rely on manual operation, affecting transaction continuity. In other words, most existing solutions lack a corresponding automated recovery mechanism after a circuit breaker is triggered. This means that after the system stops, manual intervention is required to troubleshoot the problem, confirm the risk has been eliminated, and then manually restore service.
[0009] This manual recovery method has several problems:
[0010] - The recovery period is long, which may result in prolonged transaction interruptions;
[0011] - It relies heavily on the experience of maintenance personnel and is prone to errors;
[0012] - It is not conducive to building a highly available and self-healing intelligent trading system.
[0013] Especially in a transaction environment where every second counts, even a few minutes of pause can lead to huge losses.
[0014] 3. Insufficient monitoring granularity: The following problems still exist in the process of monitoring the order transaction system:
[0015] The lack of granular tracking across the entire transaction chain means that many current systems monitor only a single metric (such as interface response time or server resource usage), lacking end-to-end visual monitoring of the entire transaction process. This implies:
[0016] - It's difficult to quickly pinpoint which part of the process went wrong;
[0017] - When a fault occurs, the troubleshooting efficiency is low, which affects the repair speed;
[0018] - The lack of in-depth analysis of user transaction behavior and order flow path makes it difficult to provide accurate early warnings and interventions.
[0019] Therefore, there is a demand for a new generation of automated intelligent transaction monitoring circuit breaker and rapid recovery solution that is more intelligent, adaptive, and has automated recovery capabilities, in order to improve the stability, security, and operational efficiency of the trading system. Summary of the Invention
[0020] This application provides a new generation of circuit breaker and recovery control system that is more intelligent, adaptive, and has automated recovery capabilities, in order to improve the stability, security, and operational efficiency of the trading system.
[0021] According to a first aspect of this application, a system for automated intelligent transaction monitoring, circuit breaker triggering, and rapid recovery is provided, comprising:
[0022] The data acquisition module is configured to collect transaction data in real time, including various indicator data.
[0023] An intelligent anomaly detection module is configured to evaluate the health score of the order transaction system based on the collected indicator data and predefined rules; and
[0024] The dynamic circuit breaker and recovery decision module is configured to automatically trigger different levels of circuit breaker mechanisms based on the severity of the detected anomalies and the circuit breaker strategy set according to the scenario, while also providing a corresponding gradual recovery mechanism.
[0025] According to a second aspect of this application, a method for automated intelligent transaction monitoring, circuit breaker triggering, and rapid recovery is provided, comprising:
[0026] In the initial stage:
[0027] Real-time collection of various transaction data, including indicator data;
[0028] The weights of the indicator data are dynamically adjusted based on the indicator data and the historical performance of each indicator data.
[0029] The health score of the order transaction system is calculated based on the adjusted weights and the indicator data;
[0030] The threshold used to determine whether to trigger the circuit breaker mechanism is dynamically calculated based on the basic threshold and the health score.
[0031] During the monitoring phase:
[0032] Whether to trigger a circuit breaker is determined by comparing the current health score of the order transaction system with the threshold:
[0033] If the current health score of the order transaction system does not exceed the threshold, the order transaction system continues to operate normally.
[0034] If the current health score of the order transaction system exceeds the threshold, a circuit breaker will be triggered according to the circuit breaker strategy associated with the current scenario; and after a preset cooling-off period from the occurrence of the circuit breaker, a gradual recovery phase will begin.
[0035] This overview is provided to introduce, in a simplified form, some of the concepts further described in the detailed description below. This overview is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0036] To describe how the above and other advantages and features of the invention are obtained, a more detailed description of the invention, which has been briefly described above, will be presented with reference to specific embodiments of the invention shown in the accompanying drawings. It will be understood that these drawings depict only exemplary embodiments of the invention and are therefore not intended to limit its scope. The invention will be described and explained using the drawings and with the aid of additional features and details, in which:
[0037] Figure 1 A schematic system architecture diagram of an automated intelligent transaction monitoring circuit breaker and rapid recovery system according to an embodiment of this application is shown.
[0038] Figure 2 A schematic flowchart of an automated intelligent transaction monitoring circuit breaker and rapid recovery method according to an embodiment of this application is shown. Detailed Implementation
[0039] As mentioned earlier, existing order transaction system circuit breaker and recovery mechanisms have gradually revealed problems such as untimely response, high misjudgment rate, and slow recovery when facing high concurrency and complex and ever-changing order transaction scenarios. The market urgently needs a new generation of circuit breaker and recovery control system that is more intelligent, adaptive, and has automated recovery capabilities to improve the stability, security, and operational efficiency of transaction systems. This is precisely the core problem that this invention aims to solve.
[0040] Therefore, the solution proposed in this application overcomes the various defects of the traditional technologies mentioned above by introducing dynamic threshold adjustment, full-link monitoring and automated recovery mechanism, thereby creating a smarter, more efficient and more flexible order transaction circuit breaker and recovery solution.
[0041] First of all, Figure 1 The diagram illustrates a schematic system architecture of an automated intelligent transaction monitoring, circuit breaker, and rapid recovery system according to an embodiment of this application. The architecture visually demonstrates the interaction between data monitoring, circuit breaker decision-making, execution, and recovery.
[0042] As shown in the figure, the system described in this application mainly includes six modules: real-time data acquisition, intelligent anomaly detection based on rule matching and evaluation, dynamic circuit breaker and recovery decision-making, visualization and alarm notification, and logging and auditing. Among these, the data acquisition module 110, the intelligent anomaly detection module 120, and the dynamic circuit breaker and recovery decision-making module 130 are the core modules for implementing the solution of this application. The visualization and alarm notification module 140 and the logging and auditing module 150 are optional modules added to further improve the solution.
[0043] These modules exchange data with each other using various wireless or wired communication technologies. The following sections, with reference to the accompanying diagrams, will describe each module in detail.
[0044] As shown in the figure, the data acquisition module 110 is configured to collect transaction data in real time.
[0045] Common transaction data can include: product data (name, price, inventory, category), order data (number, status, order time, shipping address), logistics data (courier company, tracking number, status), user data (ID, name, contact information, address), payment data (method, amount, status), marketing campaign data (promotion type, coupons), statistical data (transaction volume, transaction amount, store rating), request response time, error rate, throughput, success rate, etc.
[0046] In this application, the primary purpose is to monitor abnormal situations in transactions, rather than to collect transaction data. Therefore, the data acquisition module 110 mainly collects various indicator data related to transaction service quality, including but not limited to:
[0047] - Request response time T (Latency), also known as "request timeout": the time required to process each request. Where latencySum is the total latency, and totalRequests is the total number of requests.
[0048] Error Rate (E): The percentage of failed requests. Here, failRequests is the number of failed requests, and totalRequests is the total number of requests.
[0049] - Throughput Q: The number of requests successfully processed per unit of time. Where totalRequest is the total number of requests and windowsSeconds is the window duration.
[0050] - Success Rate (S): The percentage of successful requests. Where successRequests is the number of successful requests and totalRequests is the total number of requests.
[0051] It should be understood that although the embodiments of this application are discussed using the above four dimensions of indicator data, in practical applications, it is also feasible to implement the solution of this application using more or fewer dimensions of data. For example, in order to improve the accuracy of circuit breaking and recovery of this solution, hardware and network indicators such as CPU load, memory load, IO load, and network load, and auxiliary indicators such as latency volatility, dependent service status, and abnormal user behavior rate can be introduced, thereby implementing the solution of this application from more dimensions, all of which are within the scope of protection of this application.
[0052] By monitoring the trading system and obtaining these indicator data from the trading data, and then providing them to the intelligent anomaly detection module 120 based on rule matching and evaluation for analysis, it is possible to determine whether the trading system has experienced any abnormalities. The intelligent anomaly detection module 120 is configured to evaluate the health score of the order trading system based on the collected indicator data and predefined rules.
[0053] In traditional order trading systems, circuit breaker rules are based on a static threshold limit, meaning that a circuit breaker is triggered when a monitored indicator exceeds a fixed threshold. While this mechanism is simple to implement, it often fails to accurately reflect the true risk situation in practice.
[0054] In real-world production environments, traffic patterns can change significantly due to factors such as business cycles, holidays, and promotional activities. For example, on certain dates, order volume may surge by tens or even hundreds of times compared to usual, which does not actually constitute an anomaly in transactions. If the original fixed threshold is still used, it will trigger unnecessary circuit breakers.
[0055] To make the circuit breaker mechanism more intelligent, stable, and adaptable, this application incorporates the concept of a "health score." By dynamically adjusting the weights and warning thresholds of various indicator data within the health score based on the current system state, a health score that accurately reflects the health status of the order transaction system can be provided. Based on a comparison of the health score with preset status thresholds, it can be determined whether the system is in an abnormal state.
[0056] Note: For ease of understanding, the calculation formulas and parameters and contents involving specific values given below are all examples or default parameters. In actual use, it is necessary to obtain historical data in the specific application scenario, and they all support dynamic configuration and adjustment.
[0057] As shown in the figure, the intelligent anomaly detection module 120 based on rule matching and evaluation mainly includes three major modules (functions): dynamic weight adjustment, health score, and threshold adjustment.
[0058] 1.1 Dynamic Weight Adjustment (Module):
[0059] As mentioned above, the intelligent anomaly detection module 120 can obtain various indicator data related to the quality of transaction services, such as request response time, error rate, throughput and success rate, from the data acquisition module 110.
[0060] Upon receiving this indicator data, the intelligent anomaly detection module 120 dynamically adjusts the proportion of each indicator data in the health score using a sliding time window, based on the importance of different indicator data to the health score. In other words, by statistically analyzing the historical performance of each indicator data (such as volatility, number of anomalies, deviation ratio, etc.) through a sliding time window, the weight of each indicator data in the health score is dynamically adjusted based on these statistical results.
[0061] Explanation: A sliding time window is a dynamic statistical method that slides a fixed-length time window along a continuous time axis to calculate statistical values (such as mean, variance, maximum value, number of outliers, etc.) for a recent period. The sliding time window can be divided into the following two steps:
[0062] 1) Time Window Segmentation: The fixed period of the sliding window is divided into multiple smaller time windows, and each smaller window independently records the number of requests. For example, suppose that when monitoring a business service, the sliding window has a fixed period of ten minutes, and each period is divided into ten smaller time windows. This means that the success rate is recorded once per minute.
[0063] 2) Sliding Statistics: As time progresses, the sliding window dynamically moves, always covering the most recent small windows for statistical analysis. Taking the example above, when one minute has passed, the sliding window will cover the success rate recorded in the eleventh minute, while excluding the success rate data from the first minute.
[0064] Taking the dynamic calculation of deviation ratio dimensions as an example, the specific calculation process can be represented as follows:
[0065] Step 1: Input parameters, for example, obtain the following parameters through data acquisition module 110, including:
[0066] - Current request response time: T = currentLatency, in milliseconds (ms), the lower the value, the better;
[0067] - Current throughput: Q = currentThroughput, which is the number of requests processed per unit of time. The higher the value, the better.
[0068] - Current error rate: E = currentErrorRate, in percentage (%) form, the lower the value, the better;
[0069] - Current success rate: S = currentSuccessRate, in percentage (%) form, the higher the value, the better.
[0070] Step 2: Determine the baseline values for the parameters, including:
[0071] - Request response time baseline value T0 = baseLatency;
[0072] - Throughput baseline value Q0 = baseThroughput;
[0073] - Error rate baseline E0 = baseErrorRate; and
[0074] - The success rate baseline value S0 = baseSuccessRate.
[0075] Different application scenarios may require different baseline value selection strategies, which may include:
[0076] a) Strategy based on historical baselines:
[0077] - Average value: Uses the average value over a past period as a benchmark. Suitable for stable services with low cyclicality.
[0078] - Moving average: Using the average value within a sliding window can better reflect recent trends.
[0079] - Percentiles (e.g., 95th percentile): Suitable for situations where the impact of extreme values needs to be considered, and can better capture the performance level of the service in most cases.
[0080] b) Target value-based strategy:
[0081] - Business Objectives: Ideal values set based on business needs. For example, during e-commerce promotions, set high target values for the success rate of certain key interfaces.
[0082] - Service Level Agreement (SLA): Sets a baseline value based on the service quality commitment signed with the customer.
[0083] c) Strategies based on scenario-based baselines:
[0084] - Holiday / Promotional Period Special Exception: Specific baseline values are set to take into account the different traffic patterns during special time periods.
[0085] - Different user groups: Set different baseline values for different user groups (such as VIP users and regular users).
[0086] Step 3: Calculate the deviation ratio between the input parameter value and the baseline value, including:
[0087] - Request response time deviation ratio ;
[0088] - Throughput deviation ratio, ;
[0089] - Error rate deviation ratio ;
[0090] The success rate deviation ratio is 1 - request response time deviation ratio - throughput deviation ratio - error rate deviation ratio.
[0091] Step 4: Determine the dynamic weights, as shown in the following formula:
[0092] - Request response time weight W T : .
[0093] Response time directly impacts user experience and system load. Considering its sensitivity in high-concurrency scenarios and its direct impact on user experience, its initial weight is set to 0.4. This means that in the health scoring model, request response time has the largest influence in the initial stage, ensuring a rapid response when latency fluctuations occur. As latency increases, the weight decreases accordingly; however, it can only decrease to a maximum of 0.4 × (1 - 1.5) = 0.2. When the current request response time T exceeds 1.5 times the baseline value T0, the request response time weight no longer decreases.
[0094] - Error rate weight W E : .
[0095] Error rate directly reflects the abnormality of an interface or service. Especially when system load increases or dependent services malfunction, the error rate is often one of the first indicators to show significant changes. Here, its initial weight is set to 0.35, indicating that the error rate has a significant impact in the initial stage of the health scoring model, so that circuit breakers or alarm mechanisms can be triggered promptly in the event of abnormal fluctuations. When the current error rate exceeds the baseline value, the weight is increased for each point exceeding it, but the maximum increase is 0.2.
[0096] - Throughput weight W Q : .
[0097] Throughput reflects a system's processing and load capacity, and is particularly valuable during peak traffic periods. Therefore, its initial weight is set to 0.15 in the model to enhance the system's ability to detect performance bottlenecks. When the current throughput is below half the baseline value, the weight remains at 0.075; otherwise, it increases as throughput increases.
[0098] - Success rate weight W S : .
[0099] Since the sum of the dynamic weight coefficients of the four dimensions should be 1, the success rate weight is calculated using the formula above.
[0100] Therefore, the final output weight vector is:
[0101]
[0102] It should be understood that the above dynamic weight calculation based on the deviation ratio dimension is merely given as an example and is not intended to limit the scope. The indicator data and values in the examples shown in the calculation process are based on common conventions and practical experience, and are not restrictive requirements.
[0103] 1.2 Health Score Calculation (Module):
[0104] Based on the four dimensions of indicator data (such as request response time, error rate, throughput, and success rate mentioned above), combined with the weight coefficients calculated from the dynamic weight module, and after normalization and weighted summation, the final health score HealthScoreFin can be obtained.
[0105] The example calculation process can be represented as follows:
[0106] Step 1: Input parameters. As mentioned above, the following parameters can be obtained through the data acquisition module 110 (see the previous description for details):
[0107] - T: Request response time (ms) ,
[0108] - E: Error rate ,
[0109] - Q: Throughput (QPS) ,
[0110] - S: Success rate .
[0111] Step 2: Perform normalization:
[0112] To ensure direct comparison of different indicator data, each indicator data needs to be normalized. Linear normalization is used here:
[0113] Linear normalization (also known as max-min normalization or Min-Max Scaling) is a commonly used data preprocessing method. Its basic formula is:
[0114] - For metrics where "the smaller the better" (such as response time and error rate), use the following formula:
[0115]
[0116] - For metrics where "the higher the better" (such as throughput and success rate), use the following formula:
[0117]
[0118] Where N max and N min These are the minimum and maximum values of the indicator data over a certain period of time. The selection and determination of these values can be based on the following methods.
[0119] a) Historical data analysis
[0120] Definition: Using statistical data over a past period to determine the maximum and minimum values.
[0121] Example:
[0122] - Regarding response time (RT), assuming the longest recorded time in the past month is 1500ms, then T can be... max = 1500ms; the shortest time is 200ms, then T min =200ms.
[0123] - For the error rate (ER), if the highest error rate in the past month was 1%, then E can be set as... max = 0.01; the minimum error rate is 0.1%, then E min =0.001.
[0124] b) Business Requirements and Service Level Agreement (SLA)
[0125] Definition: Set the maximum and minimum values according to business needs or SLA requirements.
[0126] Example:
[0127] - If the service level agreement stipulates that the response time must not exceed 150 milliseconds - 2 seconds (150ms - 2000ms), then even if this value is not reached in historical data, T can still be used. max = 2000ms as the upper limit, T min =150ms as the lower limit.
[0128] - Regarding the error rate, if the SLA requires the error rate to be no more than 0.1%-0.5%, then E can be set. max = 0.005, E min = 0.001.
[0129] c) Experience and expert judgment
[0130] Definition: To set reasonable maximum and minimum values by combining industry experience or domain expert knowledge.
[0131] Example:
[0132] In some high-performance computing environments, a response time exceeding 1 second (1000ms) may be considered unacceptable, therefore a timeout of T can be set. max = 1000ms, while based on historical data and experience, the minimum time is estimated to be 150ms, so T can be set. min =150ms.
[0133] Regarding the error rate, considering user satisfaction, it is generally believed that an error rate exceeding 1% will negatively impact user experience; therefore, let E be... max = 0.01, while the error rate of similar systems is 0.5%, so E can be set min =0.005.
[0134] Therefore, we can conclude that:
[0135] - Request response time normalization: ,
[0136] - Error rate normalization: ,
[0137] - Throughput normalization: ,
[0138] - Success rate normalization: ,
[0139] It should be understood that the normalization may include other types of normalization processing besides linear normalization, such as maximum value normalization, robust normalization, etc., and is not limited to the examples.
[0140] Step 3: Calculate the health score:
[0141] Combined with the corresponding weight vector obtained in 1.1 dynamic weight adjustment In the example, it is Given (0.4, 0.35, 0.15, 0.1), the corresponding health score can be calculated using the health score formula shown in the example below:
[0142]
[0143] Among them, due to N T With N E The values for these two indicators are "the smaller the better," so when calculating the health score, these two indicators were reversed. This ensures that the final calculated score follows the principle of "the higher the better".
[0144] It should be understood that the final result of the health score should be limited to the range [0, 1], that is:
[0145]
[0146] A health score of 1 indicates that the system is functioning well.
[0147] A health score of 0 indicates that the system is completely unusable.
[0148] A health score in the middle indicates varying degrees of abnormality or deterioration.
[0149] The health score plays a key role in the circuit breaker mechanism in the following ways:
[0150]
[0151] 1.3 Threshold Adjustment (Module):
[0152] In addition to dynamically adjusting the weights of the indicator data in this application, a threshold containing multiple indicator data (also known as a "comprehensive monitoring threshold") can be dynamically calculated based on the basic threshold and the current system health score value to determine whether to trigger the circuit breaker mechanism.
[0153] The threshold plays a key role in the circuit breaker mechanism in the following ways:
[0154]
[0155] The example calculation formula for the threshold is as follows:
[0156] a) For metrics that are "the smaller the better" (request response time, error rate)
[0157] Objective: When health scores decline, appropriately relax the upper limit (allowing for slower speeds or more errors).
[0158]
[0159] Here, Threshold represents the threshold, HealthScoreFin is the health score calculated above, and baseThreshold is the boundary of indicators that the system can tolerate in a healthy state. It is the starting point for the calculation of dynamic monitoring thresholds. Its setting should be combined with historical data, business needs, SLA and expert experience, and can be continuously optimized through dynamic mechanisms.
[0160] For example:
[0161]
[0162] Where k is an adjustment factor, reflecting the degree of influence of the health score on the final circuit breaker threshold. Based on current usage scenario data, the following suggestions can be given:
[0163] - Conservative Strategy (High Sensitivity): Suitable for situations with extremely high service quality requirements and where no performance degradation is permissible. Recommended adjustment factor range: k=0.8 to k=1.0.
[0164] - Balanced Sensitivity: Ideal for most web applications and services, it allows for rapid response to unexpected problems while tolerating a certain degree of short-term volatility. Recommended adjustment factor range: k=0.4 to k=0.7.
[0165] - Low Sensitivity Strategy: Suitable for applications that are less sensitive to latency and prioritize long-term continuous service capabilities. Recommended adjustment factor range: k=0.1 to k=0.3.
[0166] Therefore, a lower health score leads to a larger (1 − HealthScoreFin) threshold, which in turn allows for a higher response time or error rate. Conversely, a higher health score results in a tighter threshold and a more sensitive system.
[0167] Example 1 (Request Response Time):
[0168] - baseThreshold = 500ms
[0169] - k = 1
[0170] - HealthScoreFin = 0.6
[0171] - Threshold = 500 + 200 × (1 - 0.6) = 500 + 80 = 580ms
[0172] If the current request response time is > 580ms → trigger an alarm or circuit breaker.
[0173] Example 2 (Error Rate):
[0174] - baseThreshold = 0.01 (1%)
[0175] - k = 0.005
[0176] - HealthScoreFin = 0.7
[0177] - Threshold = 0.01 + 0.005 × (1 - 0.7) = 0.01 + 0.0015 = 0.0115
[0178] If the current error rate is > 1.15%, trigger an exception check.
[0179] b) For metrics that are considered "the bigger the better" (throughput, success rate)
[0180] Objective: When the health score declines, appropriately lower the lower limit (allowing for lower throughput or success rate).
[0181]
[0182] Specifically: a lower health score leads to a higher (1−HealthScoreFin) threshold, which in turn lowers the threshold and allows for lower throughput or success rates. Conversely, a higher health score increases the threshold and places higher demands on the system.
[0183] Example 3 (throughput):
[0184] - baseThreshold = 800 TPS
[0185] - k = 1
[0186] - HealthScoreFin = 0.6
[0187] - Threshold = 800 - 100 × (1 - 0.6) = 800 - 40 = 760 TPS
[0188] If the current throughput is less than 760, a performance alert will be triggered.
[0189] Example 4 (Success Rate):
[0190] - baseThreshold = 0.99 (99%)
[0191] - k = 0.005
[0192] - HealthScoreFin = 0.7
[0193] - Threshold = 0.99 - 0.005 × (1 - 0.7) = 0.99 - 0.0015 = 0.9885
[0194] If the current success rate is less than 98.85%, trigger anomaly detection.
[0195] Based on the above formula, the warning thresholds for the four indicators—request response time, error rate, throughput, and success rate—can be calculated as follows:
[0196] - Request response time threshold: T Threshold =T0+k×(1-HealthScoreFin)
[0197] - Error rate threshold: E Threshold =E0+k×(1-HealthScoreFin)
[0198] - Throughput threshold: Q Threshold =Q0-k×(1-HealthScoreFin)
[0199] - Success rate threshold: S Threshold =S0-k×(1-HealthScoreFin)
[0200] Among them, the request response time threshold is: T Threshold Error rate threshold: E Threshold Throughput threshold: Q Threshold Success rate threshold: S Threshold These four thresholds will be used by the circuit breaker module to determine whether the current system is normal, i.e., whether to trigger the circuit breaker. Alternatively, other handling measures can be adopted based on pre-configured circuit breaker methods such as the business system level and user type. For example, "trigger performance alarm", "trigger anomaly detection", and "trigger anomaly judgment". These measures are all specific methods of circuit breaking. In fact, the logic after each indicator reaches the threshold is basically the same, which is roughly: execute the circuit breaker (the circuit breaker here includes alarms, anomaly detection, interface degradation, etc.), while increasing the detection frequency to check whether the system meets the recovery conditions, and then gradually restore it.
[0201] The dynamic circuit breaker and recovery decision module 130 is configured to automatically trigger different levels of circuit breaker mechanisms (such as request degradation, limiting request frequency, etc.) based on the severity of the detected anomaly (health score) and the circuit breaker strategy (risk level strategy) set according to the scenario, in order to prevent the spread of risk, and also provides a corresponding gradual recovery mechanism.
[0202] As mentioned earlier, traditional circuit breaker mechanisms use static thresholds (such as triggering a circuit breaker when the failure rate exceeds 50%). This mechanism is prone to the following problems in the following scenarios:
[0203] - During peak holiday periods, trading volume surges, and normal trading behavior is mistakenly identified as abnormal.
[0204] - User behavior patterns change drastically during promotional activities, and fixed thresholds cannot adapt to this.
[0205] - Delayed response during sudden attacks or system failures, missing the best time to trip the circuit breaker;
[0206] - The threshold for business downturns is too high, making it impossible to detect potential risks in a timely manner.
[0207] Therefore, the "dynamic circuit breaker" in this application aims to enable the system to "dynamically adjust the circuit breaker strategy and threshold based on multiple factors such as time period, business characteristics, risk level, and historical behavior, so as to achieve refined circuit breaker control." In layman's terms, it means "looking at the time, the scenario, and the trend" to make circuit breaker decisions that are "tailored to local conditions."
[0208] Therefore, the core objective of this application in formulating the circuit breaker strategy is:
[0209]
[0210] To achieve these objectives, the dynamic circuit breaker and recovery decision module constructs a circuit breaker mechanism based on a standard three-state model. The table below shows the three states of the circuit breaker:
[0211]
[0212] Detailed explanation:
[0213] CLOSED: The request has been approved. Health scores and various indicators will continue to be monitored.
[0214] OPEN: Enters after the circuit breaker is triggered, all requests fail quickly (fallback).
[0215] HALF-OPEN: After the cooldown period ends, attempt to resume operations, allowing only a small number of requests for probing.
[0216] The transition conditions are achieved by comparing health scores with dynamic thresholds.
[0217] The logic for transitioning between the three states is as follows:
[0218]
[0219] For example, when a system transitions from the CLOSED state to the OPEN state due to a service anomaly, it will reject all requests to the target service to prevent the failure from spreading. Subsequently, after a preset cooling-off period (e.g., 30 seconds), the system enters the HALF-OPEN state, signifying that the system begins to attempt to gradually restore traffic (i.e., through "small-volume probing" and evaluating the probing results), rather than allowing all requests at once.
[0220] Depending on the application scenario, various circuit breaker strategies can be formulated to enter the OPEN state. The table below lists an example circuit breaker strategy:
[0221]
[0222] Proportional circuit breaking refers to determining whether to trigger a circuit breaker based on the success rate of requests over a period of time. It is a dynamically adjusted strategy that can flexibly respond to changes in service status under different loads. It is suitable for services with large traffic fluctuations and occasional momentary failures, such as recommendation systems and push advertising.
[0223] Example: In a news application, the recommendation system's API may become unstable due to sudden surges in traffic. If the failure rate of the recommendation service reaches 40% within five minutes, the system will trigger a proportional circuit breaker, suspending requests to the service and displaying default content to the user until the service returns to normal.
[0224] Tiered circuit breaking refers to triggering circuit breakers or downgrades based on business level and / or user type. Therefore, this tiered circuit breaking strategy fully considers the specific business type and user identity before deciding which circuit breaker or downgrade to execute when triggering a circuit breaker.
[0225] The following table shows an example of a tiered circuit breaker rule table:
[0226]
[0227] Dependency circuit breaking is specifically designed for service call chains with explicit dependencies. When a dependent service encounters a problem, the dependency circuit breaking mechanism can promptly sever the connection, preventing upstream services from being affected. It is primarily used in inter-service call scenarios in distributed systems, especially for applications that heavily rely on external services.
[0228] Example: In an aggregated payment scenario, the POS terminal supports two payment methods: QR code scanning and card swiping. If a card swiping transaction channel is under maintenance and cannot function properly, the circuit breaker mechanism will prevent the POS from continuing to attempt to send a request to that channel. Instead, it will return a message saying "Card swiping is temporarily unavailable" and suggest that the user try again later or use another payment method to complete the transaction.
[0229] After a circuit breaker is triggered, corresponding processing actions (degradation strategies) are usually required to provide feedback. Example degradation strategies are shown in the table below:
[0230]
[0231] After entering the HALF-OPEN state, the system begins to attempt to gradually restore traffic. Gradual restoration means that instead of pursuing a "one-time restoration," it uses a closed-loop mechanism of "small-volume probing, followed by evaluation of the feedback results, and then, if the probing is successful, gradually increasing the traffic until full access is granted; if the probing fails, it returns to the OPEN state," thus achieving a safe and smooth service recovery process.
[0232] For example, in HALF-OPEN mode, a low-flow detection phase is initiated first:
[0233] 1) The dynamic circuit breaker and recovery decision module 130 only allows a very small number of requests to pass (e.g., 1-2 per second), while the remaining requests still follow the degradation logic. Among them, the probe requests cover typical business scenarios (such as queries, order placement, etc.) to ensure that the evaluation is representative.
[0234] 2) Real-time statistics of detection results: success rate, response time, whether dynamic thresholds are triggered, changes in health scores, etc.
[0235] 3) Based on the results, determine whether the goals of this stage have been achieved to verify whether the service has basic availability. If the service has basic availability,
[0236] Based on the judgment result, execute the corresponding recovery decision (two paths):
[0237] a) Path 1: Recovery failed (service is not yet available) → Return to OPEN.
[0238] Example: If any of the following conditions are present in the probe request:
[0239] Success rate < 80%;
[0240] Response time > dynamic threshold;
[0241] Health score < 0.6;
[0242] If the service is still unstable, the system will re-enter the OPEN state.
[0243] Optional strategy: Extend the cooling-off period (e.g., use exponential retreat: 30s → 60s → 120s) to avoid frequent testing that could cause pressure fluctuations.
[0244] This path, the transition from HALF-OPEN to OPEN, is a safety net within the "gradual recovery" mechanism, preventing the system from being overwhelmed again before it is ready.
[0245] b) Path 2: Recovery successful (service has basic availability) → Enter the "gradual rollout" phase.
[0246] Example of a gradual scaling strategy:
[0247]
[0248] As can be seen from the table above, the proposed solution employs a "gradual recovery" approach. Instead of immediately entering CLOSED (full request recovery) after HALF-OPEN, it gradually increases the allowed traffic in stages, evaluating the metrics at the end of each stage. If any abnormalities are found, it immediately reverts to the OPEN state. This avoids false recovery phenomena caused by fluctuations in metric data.
[0249] The following section provides several specific application scenarios based on the above-mentioned circuit breaking and recovery methods, to help those skilled in the art better understand this invention.
[0250] In the example described, the basic threshold / parameter configuration shown in the table below is used, which is also the recommended configuration. Of course, the configuration can be modified accordingly based on specific application scenarios and needs, and all such modifications fall within the scope of protection of this application.
[0251] Practical application scenarios of circuit breaker mechanisms (classified by type):
[0252] a) Full circuit breaker: Isolation of sudden failures in the payment system
[0253] Scene description:
[0254] During the "Double 11" shopping festival, an e-commerce platform used third-party payment gateways (such as Alipay / WeChat Pay) to complete payments after users placed orders. Due to the surge in traffic, the payment gateway's response latency increased from 200ms to 3s, and it began returning timeout errors.
[0255] Consequences of not triggering a circuit breaker:
[0256] - Payment requests are piling up, and the thread pool is exhausted;
[0257] - Users are stuck on the "Payment in progress" page for an extended period of time;
[0258] - Order processing was slowed down, affecting the entire order placement process.
[0259] Full circuit breaker response strategy:
[0260] - Configure circuit breaker rules: If ≥6 out of 10 consecutive calls fail (error rate ≥60%), immediately enter the OPEN state;
[0261] - All subsequent payment requests failed directly, returning the message: "The current payment channel is busy. Please try again later or choose another method."
[0262] - Simultaneously log and issue alerts to notify operations and maintenance personnel to investigate.
[0263] Effect:
[0264] To prevent payment service collapse and protect the core order process; although users cannot pay immediately, the system as a whole remains available.
[0265] b) Proportional Circuit Breaker: Automatic degradation of the recommendation system due to abnormal fluctuations.
[0266] Scene description:
[0267] The "personalized recommendation" module on the homepage of a news app relies on an AI model service, which normally has a 99% success rate. However, in one instance, a memory leak caused some requests to fail, resulting in 180 failures out of 500 calls within one minute (a failure rate of 36%).
[0268] Consequences of not triggering a circuit breaker:
[0269] - Many users see a blank recommendation section or a loading screen;
[0270] - Client-side retries increase backend pressure;
[0271] - A decline in user experience may lead to user churn.
[0272] Proportional circuit breaker response strategy:
[0273] - Set a sliding window to count requests within the last 60 seconds;
[0274] - Triggering conditions: Failure rate > 30% and total number of requests ≥ 50;
[0275] - Triggered action: Pause the recommendation service for 30 seconds and instead display a static list of "Hot News".
[0276] Effect:
[0277] The system automatically identifies abnormal fluctuations and quickly switches to the backup plan to ensure uninterrupted display of homepage content.
[0278] c) Tiered Circuit Breaker: Multi-level service protection for e-commerce platforms
[0279] Scene description:
[0280] An e-commerce system comprises multiple business modules. During peak sales periods, resources are scarce, and priority must be given to ensuring the core transaction process.
[0281]
[0282] Based on the actual response to the above-mentioned tiered circuit breaker strategy:
[0283] - The system dynamically adjusts the circuit breaker threshold of each module based on the current load;
[0284] - When the overall health score is below 0.7, proactively shut down L3 / L4 services to free up resources for L1 / L2;
[0285] Effect:
[0286] Under high load, the system prioritizes core functions while sacrificing peripheral ones, ensuring users can complete orders and payments.
[0287] d) Dependency Circuit Breaker: Service Chain Protection in Microservice Architecture
[0288] Scene description:
[0289] In a bank's core system, the transfer function relies on multiple downstream services:
[0290] [Fund Transfer Service] → [Account Service] → [Risk Control Service] → [Bookkeeping Service]
[0291] One day, the risk control service experienced a sudden spike in response time to 5 seconds due to database table locking.
[0292] The consequences of not relying on circuit breakers:
[0293] - All transfer requests were stuck at the risk control stage;
[0294] - The account service thread pool is full and cannot process other requests;
[0295] - This ultimately caused the entire transfer function to malfunction.
[0296] Reliance on circuit breaker response strategies:
[0297] - Set up a dependency circuit breaker mechanism at the entry point where the transfer service calls the risk control service;
[0298] - Configuration rule: Trigger circuit breaker if timeout is 800ms or error rate > 40%;
[0299] - Post-circuit breaker action: Skip risk control checks (go through the fast track), record asynchronous audit logs, and conduct supplementary audits afterward;
[0300] - Downgrade notification: "Transaction has been accepted, risk control will be completed in the background."
[0301] Effect:
[0302] Even if the risk control service malfunctions, the transfer can still continue, achieving "fault isolation + business continuity".
[0303] e) Comprehensive Case Study: Application of Circuit Breaker Combinations in the "Spring Festival Gala Red Envelope Campaign" by Live Streaming Platforms
[0304] Background: A live streaming platform held a "red envelope giveaway" event during the Spring Festival, with a peak QPS expected to reach 500,000.
[0305] Facing challenges:
[0306] - The red envelope service relies on multiple systems such as the user center, wallet, and push notifications;
[0307] - Traffic is highly concentrated, and some services may be temporarily unavailable;
[0308] - The core "red envelope grabbing" function must be kept stable.
[0309] Combination of circuit breaker strategies:
[0310]
[0311] Final result:
[0312] During the event, the availability of the red envelope service reached 99.95%. The failure of a few dependent services did not affect the user experience, and the system smoothly weathered the traffic surge.
[0313] Based on the above implementation scenarios, the applicable scenarios and advantages of selecting and combining circuit breaker strategies can be summarized in the following table:
[0314]
[0315] Best practice recommendations:
[0316] In real systems, multiple circuit breaker strategies need to be combined according to business characteristics, along with rate limiting, degradation, retries, caching, and other means, to build a complete high-availability protection system.
[0317] Execution module
[0318] a) Circuit Breaker: Based on the circuit breaker decision, a circuit breaker instruction is sent to the transaction core system via interface call or message queue to immediately block the operation privileges of abnormal request services.
[0319] b) Health check and recovery: After the abnormal behavior is confirmed to be eliminated or repaired, transaction permissions are gradually restored through a health check process to ensure that the system restart process is smooth and controllable and to avoid secondary impact.
[0320] Execution of circuit breaking and recovery:
[0321] Circuit Breaker: Based on the circuit breaker decision, a circuit breaker command is sent to the transaction core system via interface call or message queue to immediately block the operation privileges of abnormal request services.
[0322] Health check and recovery: After the abnormal behavior is confirmed to be eliminated or repaired, transaction permissions are gradually restored through a health check process to ensure that the system restart process is smooth and controllable and to avoid secondary impact.
[0323] The visualization and alarm notification module 140 is configured to provide a graphical interface to display transaction status, circuit breaker events and system health, and to send alarm information to relevant personnel via SMS, email or API interface.
[0324] The logging and auditing module 150 is configured to record the time, cause, scope of impact, and operation logs for each circuit breaker event, for subsequent regulatory review and system optimization.
[0325] This concludes the introduction of the schematic system architecture of the automated intelligent transaction monitoring circuit breaker and rapid recovery system.
[0326] Based on the above schematic system architecture, combined with Figure 2 The following is a schematic flowchart illustrating an embodiment of an automated intelligent transaction monitoring circuit breaker and rapid recovery method according to this application.
[0327] As shown in the figure, firstly, in the initial stage:
[0328] In step 202, transaction data is collected in real time. As mentioned above, some of this data can be used as indicator data, including but not limited to: request response time T, error rate E, throughput Q, and success rate S.
[0329] Subsequently, in step 204, the corresponding weights of these indicator data are dynamically adjusted based on the indicator data and the historical performance of each indicator data (see 1.1 Dynamic Weight Adjustment).
[0330] Next, in step 206, the health score of the system is calculated based on the adjusted weights and the indicator data (see 1.2 Health Score Calculation).
[0331] Then, in step 208, a threshold for determining whether to trigger the circuit breaker mechanism is dynamically calculated based on the base threshold and the health score (see 1.3 Threshold Adjustment).
[0332] After these preparations are completed, the process enters the monitoring phase:
[0333] In step 210, the system continuously monitors the system's status, that is, it determines whether to trigger a circuit breaker by comparing whether the system's current health score exceeds the threshold.
[0334] If the system's current health score does not exceed the threshold, the process proceeds to step 212, and the system continues to operate normally.
[0335] If the system’s current health score exceeds the threshold, the process proceeds to step 214, where the circuit breaker is executed according to the circuit breaker policy associated with the current scenario.
[0336] As mentioned above, this application provides various circuit breaker strategies for different practical application scenarios, including but not limited to: full circuit breaker, proportional circuit breaker, tiered circuit breaker, and dependent circuit breaker. These circuit breaker strategies can be used individually or in combination according to the needs of the scenario to automatically adjust the corresponding strategies under different load conditions and maintain service continuity and quality.
[0337] After the circuit breaker is triggered, the aforementioned degradation strategy can also be executed.
[0338] Subsequently, in step 216, the circuit breaker event is logged and alarm information is displayed and sent to relevant personnel via a graphical interface, SMS, email, or API interface.
[0339] After a preset cooling period (e.g., 30 seconds) following the occurrence of the meltdown, the gradual recovery phase begins in step 218:
[0340] As mentioned earlier, after the system enters the HALF-OPEN state, it begins to attempt to gradually restore traffic.
[0341] In HALF-OPEN mode, first enter the low-flow detection phase:
[0342] 1) The dynamic circuit breaker and recovery decision module 130 only allows a very small number of requests to pass (e.g., 1-2 per second), while the remaining requests still follow the degradation logic. Among them, the probe requests cover typical business scenarios (such as queries, order placement, etc.) to ensure that the evaluation is representative.
[0343] 2) Real-time statistics of detection results: success rate, response time, whether dynamic thresholds are triggered, changes in health scores, etc.
[0344] 3) Determine whether the goal of this stage has been achieved based on the results to verify whether the service has basic availability.
[0345] Based on the results of the low-flow detection, execute the corresponding recovery decision (two paths):
[0346] a) Path 1: Recovery failed (service is not yet available) → Return to OPEN (circuit breaker state).
[0347] This path, the transition from HALF-OPEN to OPEN, is a safety net within the "gradual recovery" mechanism, preventing the system from being overwhelmed again before it is ready.
[0348] Optional strategy: Extend the cooling-off period (e.g., use exponential retreat: 30s → 60s → 120s) to avoid frequent testing that could cause pressure fluctuations.
[0349] b) Path 2: Small-scale recovery successful (service has basic availability) → Enter the "gradual ramp-up" phase. That is, continue with small-scale recovery → medium-scale recovery → full recovery.
[0350] Once the full recovery is complete, the system returns to the previous monitoring phase and begins monitoring for the next anomaly.
[0351] In some preferred embodiments, machine learning technology can be introduced into the rule-based matching and evaluation-based intelligent anomaly detection module 120. Machine learning models (such as LSTM, isolated forest, deep neural networks, etc.) are used to train historical transaction data to build a normal transaction behavior model. During real-time transactions, the difference between the current behavior and the model is compared to identify abnormal transaction behavior, providing pre-judgment capabilities and reducing the probability of false circuit breakers.
[0352] In summary, the automated intelligent transaction monitoring circuit breaker and rapid recovery solution of this application has the following advantages:
[0353] 1) Multiple performance metrics (such as latency, error rate, and throughput) are introduced to assess the system's health status. Each metric is normalized to ensure that data of different magnitudes can be compared on the same scale, improving the accuracy of system status assessment, providing a more comprehensive reflection of the system's actual operating condition, reducing the risk of misjudgment due to fluctuations in a single metric, and enhancing the system's stability and reliability.
[0354] 2) The weights of each indicator and the threshold for triggering circuit breakers are dynamically adjusted based on real-time data. Historical data analysis is used to optimize weight configuration and threshold settings, which enhances the system's adaptability and enables it to automatically adjust strategies under different load conditions to maintain service continuity and quality.
[0355] 3) Multi-level recovery mechanism: Supports second-level circuit breaker response and minute-level service recovery, reducing the impact of service interruption through a closed loop of "circuit breaker-assessment-recovery";
[0356] Although the techniques have been described using language specific to structural features and / or methodological actions, it should be understood that the appended claims are not necessarily limited to the described features or actions. Rather, these features and actions are described as exemplary forms of implementing these techniques.
[0357] The operations of the example processes are shown in separate boxes and are summarized with reference to these boxes. These processes are shown as a flow of logical boxes, each of which may represent one or more operations that can be implemented using hardware, software, or a combination thereof. In the context of software, these operations represent computer-executable instructions stored on one or more computer-readable media that, when executed by one or more processors, cause one or more processors to perform a given operation. Generally, computer-executable instructions include routines, programs, objects, modules, components, data structures, etc., that perform a particular function or implement a particular abstract data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the operations may be executed in any order, combined in any order, subdivided into multiple sub-operations, and / or executed in parallel to implement the described process. The described process may be executed by resources associated with one or more computing devices, such as one or more internal or external CPUs or GPUs, and / or one or more pieces of hardware logic, such as FPGAs, DSPs, or other types of accelerators.
[0358] All of the methods and processes described above can be embodied in software code modules executed by one or more general-purpose computers or processors, and can be fully automated via these software code modules. These code modules can be stored on any type of computer-executable storage medium or other computer storage device. This code can also be packaged into corresponding computer program products. Some or all of these methods can alternatively be embodied in dedicated computer hardware.
[0359] Any routine description, element, or box in the flowcharts described herein and / or in the accompanying drawings should be understood as potentially representing a module, segment, or portion of code comprising one or more executable instructions for implementing a specific logical function or element in that routine. Alternative implementations are included within the scope of the examples described herein, wherein elements or functions may be removed or performed inconsistently with the order shown or discussed, including substantially synchronous or reverse order execution, depending on the functionality involved, as will be understood by those skilled in the art.
[0360] While different embodiments have been described above, it should be understood that they are merely examples and not limitations. Those skilled in the art will appreciate that various modifications in form and detail may be made without departing from the spirit and scope of the invention as defined in the appended claims. Therefore, the breadth and scope of the invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined solely by the appended claims and their equivalents.
Claims
1. A system for automated intelligent transaction monitoring, circuit breaker triggering, and rapid recovery, comprising: The data acquisition module is configured to collect transaction data in real time, including various indicator data. The intelligent anomaly detection module is configured to evaluate the health score of the order transaction system based on the collected indicator data and predefined rules. as well as The dynamic circuit breaker and recovery decision module is configured to automatically trigger different levels of circuit breaker mechanisms based on the severity of detected anomalies and the circuit breaker strategy set according to the scenario, while also providing a corresponding gradual recovery mechanism; The intelligent anomaly detection module is configured to perform the following processes: dynamic weight adjustment, health score calculation, and threshold adjustment. The dynamic weight adjustment includes: dynamically adjusting the corresponding weights of the indicator data based on the indicator data and the historical performance of each indicator data; The health score calculation includes: calculating the health score of the order transaction system based on the adjusted weights and the indicator data; The threshold adjustment includes: dynamically calculating a threshold based on a base threshold and the health score to determine whether to trigger the circuit breaker mechanism.
2. The system as described in claim 1, characterized in that, The metrics include request response time, error rate, throughput, and success rate.
3. The system as described in claim 1, characterized in that, The dynamic circuit breaker and recovery decision module is built based on a standard three-state model, which includes: CLOSED status: The request has been approved. Health scores and various indicators will be continuously monitored. OPEN state: Entered after the circuit breaker is triggered, all requests fail rapidly; HALF-OPEN state: Attempt to resume after the cooldown period ends, allowing only a small number of requests for probing. When the system enters the OPEN state from the CLOSED state due to a service anomaly, the system will suspend all requests to the target service and, after a preset cooling-off period, enter the HALF-OPEN state to attempt to gradually restore traffic.
4. The system as described in claim 3, characterized in that, The circuit breaker strategies include: full circuit breaker that rejects all requests, proportional circuit breaker that rejects some requests, tiered circuit breaker that breaks circuits based on user level / business importance, and dependency circuit breaker that breaks circuits separately for dependent services.
5. The system as described in claim 1, characterized in that, After executing the circuit breaker operation, the dynamic circuit breaker and recovery decision module also executes one of the following degradation strategies for feedback: Returns the default value; Invoke local cache; Forward to backup interface; Asynchronous processing.
6. The system as described in claim 3, characterized in that, When the system enters the HALF-OPEN state, the following operations are performed: During the low-flow detection phase: 1) Only a very small number of requests are allowed; 2) Real-time statistical analysis of detection results; 3) Determine whether the objective of this stage has been achieved based on the detection results to verify whether the service has basic availability; Based on the assessment results, the following recovery decision will be implemented: If the determination result is that recovery has failed, the system returns to the OPEN state; If the recovery is successful, the system enters the gradual release phase, in which the released traffic is gradually increased in stages, and the indicator data is evaluated after each stage to determine if it is normal. If it is abnormal, the system immediately reverts to the OPEN state.
7. The system as described in claim 1, characterized in that, The system also includes: The visualization and alarm notification module is configured to provide a graphical interface to display transaction status, circuit breaker events and the health of the order transaction system, and to send alarm information to relevant personnel via SMS, email or API interface; The logging and auditing module is configured to record the time, cause, scope of impact, and operation logs for each circuit breaker event, for subsequent regulatory review and system optimization.
8. A method for automated intelligent transaction monitoring, circuit breaker triggering, and rapid recovery, comprising: In the initial stage: Real-time collection of various transaction data, including indicator data; The weights of the indicator data are dynamically adjusted based on the indicator data and the historical performance of each indicator data. A health score for the order transaction system is calculated based on the adjusted weights and the aforementioned indicator data; The threshold used to determine whether to trigger the circuit breaker mechanism is dynamically calculated based on the basic threshold and the health score. During the monitoring phase: Whether to trigger a circuit breaker is determined by comparing the current health score of the order transaction system with the threshold: If the current health score of the order transaction system does not exceed the threshold, the order transaction system continues to operate normally. If the current health score of the order transaction system exceeds the threshold, a circuit breaker will be triggered according to the circuit breaker strategy associated with the current scenario; and after a preset cooling-off period from the occurrence of the circuit breaker, a gradual recovery phase will begin.
9. The method as described in claim 8, characterized in that, The circuit breaker mechanism includes: CLOSED status: The request has been approved. Health scores and various indicators will be continuously monitored. OPEN state: Entered after the circuit breaker is triggered, all requests fail rapidly; In the HALF-OPEN state, after the cooling-off period following the system entering the OPEN state, the system enters the HALF-OPEN state and performs the following operations: During the low-flow detection phase: 1) Only a very small number of requests are allowed; 2) Real-time statistical analysis of detection results; 3) Determine whether the objective of this stage has been achieved based on the detection results to verify whether the service has basic availability; Based on the assessment results, the following recovery decision will be implemented: If the determination result is that recovery has failed, the system returns to the OPEN state; If the recovery is successful, the system enters the gradual release phase, in which the released traffic is gradually increased in stages, and the indicator data is evaluated after each stage to determine if it is normal. If it is abnormal, the system immediately reverts to the OPEN state.
Citation Information
Patent Citations
Distributed financial system-based non-functional test method and system
CN118939562A
Aggregate payment system based on intelligent routing
CN120471616A
Cited By
A dual-loop closed-loop control system and method for bulk commodity trading
CN122571684A