Short message channel fault processing method based on dynamic routing adjustment

By deploying the service mesh layer and Sidecar agent in the microservice architecture SMS system, building a multi-stage circuit breaker model and an intelligent chaos experimental system, the problems of troubleshooting complexity and response delay in the microservice architecture SMS system are solved, and the system's high elasticity and fast response capabilities are achieved.

CN120075859AInactive Publication Date: 2025-05-30安徽创瑞技术股份有限公司

Patent Information

Application Number
CN202510542809.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The problem of troubleshooting complexity and response delay in the SMS system in microservice architecture is difficult to effectively deal with in complex fault combination situations, resulting in the system showing unanticipated behavior and service interruptions.

Method used

By deploying the service mesh layer and Sidecar agent, a multi-stage circuit breaker model and an intelligent chaos experimental system are built, dynamic routing adjustment and progressive elastic boundary exploration are realized, Bayesian optimization and multi-objective evolution algorithms are used to automatically search for the most information-worthy fault scenarios, and routing rules and circuit breaker configuration are optimized.

Benefits of technology

It improves the system's elasticity, responds quickly and automatically switches channels, reduces the impact of faults, realizes a more refined fault handling mechanism, and improves the system's reliability and response efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075859A_ABST
    Figure CN120075859A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer network communication, and discloses a short message channel fault processing method based on dynamic routing adjustment, which comprises a Sidecar agent and a centralized control plane which are deployed in association with each micro-service instance, and automatically adopts response strategies of different levels according to fault types and severity through a multi-level circuit breaker model. An intelligent chaos experiment system is constructed in combination with Bayesian optimization and a multi-objective evolutionary algorithm to automatically search a high-value fault scene, progressive elastic boundary exploration is achieved, and routing rules and circuit breaker configuration are optimized through an elastic knowledge automatic learning system. According to the invention, the problems of fault processing complexity and response delay in a micro-service architecture short message system are solved, and the elastic capability of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer network communication, and more specifically, it relates to a method for handling short message channel failures based on dynamic routing adjustment. Background Art

[0002] With the rapid development of the mobile Internet, short message services, as an important communication means, are widely used in various industries. Modern short message systems generally adopt a microservices architecture, and each functional component is decoupled and deployed, making channel management more complex. Traditional short message channel management solutions usually directly embed routing logic into business code, making the business logic tightly coupled with communication control and difficult to flexibly respond to dynamic changes in the microservices environment. When it is necessary to adjust the routing strategy, a large amount of business code often needs to be modified, increasing the system maintenance cost and error risk.

[0003] The design and verification of routing strategies under the existing service mesh architecture mainly rely on static analysis and manual testing in limited scenarios, lacking a comprehensive understanding of the system's behavior in complex fault combination scenarios. This results in the system often showing unexpected behavior and even causing service interruptions when facing sudden and complex faults in the real production environment.

[0004] Most existing chaos testing methods adopt random fault injection or fixed scenario testing based on experience, lacking clear goal orientation and precise experimental design mechanisms, and unable to cover key fault combinations within a limited test time window, reducing the effectiveness of testing.

[0005] In addition, traditional fault handling mechanisms are scattered in different system levels, such as the application layer, network layer, and infrastructure layer, with complex management and low response efficiency. Lack of automated elastic boundary exploration capabilities, making it difficult to accurately grasp the behavior characteristics of the system under extreme conditions, resulting in a lack of strong basis for formulating elastic strategies. Therefore, there is an urgent need for a technical solution that can effectively solve the complexity of fault handling and response delay in the microservices architecture short message system. Summary of the Invention

[0006] The present invention provides a method for handling short message channel failures based on dynamic routing adjustment, which solves the technical problems of complexity of fault handling and response delay in the microservices architecture short message system in the related art.

[0007] The present invention provides a method for handling short message channel failures based on dynamic routing adjustment, including:

[0008] Deploy a service mesh layer in the short message system with a microservices architecture, where the service mesh layer includes Sidecar proxies associated with each microservice instance and a centralized control plane;

[0009] Build a multi - level circuit breaker model based on the service mesh layer. The multi - level circuit breaker model automatically adopts different levels of response strategies according to the fault type and severity. The response strategies include current limiting, partial traffic switching, and complete fuse - and - reroute;

[0010] Based on the response data of the multi - level circuit breaker model, construct an intelligent chaotic experiment system using Bayesian optimization and multi - objective evolutionary algorithms to automatically search for the most informative fault scenarios;

[0011] Based on the fault scenarios generated by the intelligent chaotic experiment system, achieve progressive elastic boundary exploration, and detect the system's elastic boundary by gradually increasing the fault intensity;

[0012] Based on the results of progressive elastic boundary exploration, construct an elastic knowledge automated learning system to extract knowledge from the experimental results and optimize routing rules and circuit breaker configurations;

[0013] Apply the optimized routing rules and circuit breaker configurations of the elastic knowledge automated learning system to the service mesh layer to achieve dynamic routing adjustment.

[0014] Furthermore, the steps of deploying the service mesh layer include:

[0015] Deploy lightweight Sidecar agents beside each microservice instance to form a data plane responsible for inter - service communication and traffic control;

[0016] Build a centralized control plane responsible for global routing policy management and service discovery;

[0017] Implement a secure communication system between the control plane and the data plane to ensure that routing policies can be quickly and accurately distributed to each Sidecar agent.

[0018] Furthermore, the steps of building a multi - level circuit breaker model include:

[0019] Build state - transition logic and use a metric - based finite - state machine model to control the circuit breaker state transition;

[0020] Use a sliding - window algorithm to calculate the error rate and use an exponentially weighted moving average to calculate the request - delay metric;

[0021] Define the trigger conditions and response measures for multi - level circuit - breaker states. The first - level circuit - breaker state triggers current limiting when the error rate exceeds a preset threshold, the second - level circuit - breaker state triggers partial traffic switching, and the third - level circuit - breaker state triggers complete fuse - and - reroute.

[0022] Furthermore, the steps of building an intelligent chaotic experiment system include:

[0023] Implement an experimental configuration engine that combines Bayesian optimization and multi - objective evolutionary algorithms;

[0024] Construct a parametric model for fault scenarios, including fault types, fault targets, fault intensities, and fault combination dimensions;

[0025] Construct an experiment execution engine responsible for converting parametric fault scenarios into specific executable fault injection instructions;

[0026] Construct an experiment result analysis module to extract valuable information from experimental data.

[0027] Furthermore, the experiment configuration engine combining Bayesian optimization and multi-objective evolutionary algorithm includes:

[0028] Establish a Gaussian process regression model to approximate the elastic scoring function;

[0029] Adopt the expected improvement sampling criterion to select the next test point;

[0030] Maintain a Pareto front composed of non-dominated solutions, and the evaluation objectives include information gain, operation risk, and experimental cost;

[0031] Use crossover and mutation operations to generate new candidate fault scenarios; Combine Bayesian optimization and multi-objective evolutionary algorithm through a serial combination strategy.

[0032] Furthermore, the steps to achieve progressive elastic boundary exploration include:

[0033] Construct a progressive elastic boundary exploration algorithm to accurately detect the system's elastic boundary by gradually increasing the fault intensity;

[0034] Construct a multi-dimensional system state monitoring module to monitor service-level metrics, resource-level metrics, business-level metrics, and channel-level metrics in real time;

[0035] Implement an intelligent security protection system, including an anomaly detection unit, a warning generation unit, an emergency intervention unit, and a protection boundary unit;

[0036] Construct an experimental parameter automatic adjustment algorithm to continuously optimize experimental parameters according to system responses.

[0037] Furthermore, the progressive elastic boundary exploration algorithm includes:

[0038] Use an adaptive step size to control the increase in fault intensity, and set the initial step size to a small value;

[0039] Adopt a composite performance index to evaluate the system state, and automatically reduce the step size when the performance index approaches the threshold to improve exploration accuracy;

[0040] Define intensity growth curves for different fault types respectively;

[0041] The multi-parameter exploration is carried out by using the alternating fixation method. Only one fault parameter is changed each time. The fault parameters include delay, packet loss rate, and error rate. After finding the elastic boundary point corresponding to the current fault parameter, the value of the current fault parameter is fixed, and the elastic boundaries corresponding to other fault parameters are continued to be explored.

[0042] Furthermore, the steps to build an elastic knowledge automated learning system include:

[0043] Implement an automatic analysis algorithm for experimental results, and analyze the experimental data based on machine learning methods to identify the correlation between fault patterns and system responses;

[0044] Build a routing policy optimization module based on reinforcement learning, and model the routing policy optimization as a Markov decision process;

[0045] Build an elastic knowledge graph generation module, and extract key information from the experimental results to build a knowledge graph including fault nodes, response nodes, effect nodes, and boundary nodes;

[0046] Build an elastic knowledge application and feedback module, apply the extracted elastic knowledge to the actual system operation, and collect feedback data for continuous improvement.

[0047] Furthermore, the automatic analysis algorithm for experimental results includes:

[0048] Perform time series alignment on the collected fault experimental data to ensure the time consistency between different metrics;

[0049] Apply sliding window aggregation to convert high-frequency sampling data into statistical features;

[0050] Use a Bayesian network to model the probability relationship between faults and system responses; implement an association rule mining algorithm to extract the association rules between faults and system responses;

[0051] Establish a scoring function for each elastic policy, taking into account the recovery time, service degradation level, and resource consumption.

[0052] Furthermore, a short message channel fault handling system based on dynamic routing adjustment, which is used to execute the steps in the above-mentioned short message channel fault handling method based on dynamic routing adjustment, includes:

[0053] A service mesh layer, including Sidecar agents associated with each microservice instance and a centralized control plane;

[0054] A multi-level circuit breaker module, which is used to automatically adopt different levels of response strategies according to the fault type and severity;

[0055] An intelligent chaos experiment module, which is used to automatically search for the most informative fault scenarios based on Bayesian optimization and multi-objective evolutionary algorithms;

[0056] A progressive elastic boundary exploration module for gradually enhancing the elastic boundary of the fault detection system;

[0057] An elastic knowledge automatic learning module for extracting knowledge from experimental results and optimizing routing rules and circuit breaker configurations.

[0058] The beneficial effects of the present invention are as follows:

[0059] It solves the problems of fault handling complexity and response delay in the microservice architecture SMS system, and improves the system's elastic ability. By constructing a service mesh layer and a Sidecar proxy system, precise control of traffic and dynamic routing adjustment are achieved, enabling the system to quickly respond and automatically switch channels when a fault occurs in the SMS channel, reducing the impact of the fault. The implementation of multi-level circuit breakers and real-time traffic control algorithms enables the system to adopt corresponding processing strategies according to the severity of the fault, from current limiting to partial traffic switching to complete fusing, providing a more refined fault handling mechanism.

[0060] In addition, through the intelligent chaos experiment system and progressive elastic boundary exploration technology, the present invention actively discovers the elastic boundary and potential fault points of the system, and continuously optimizes routing rules and circuit breaker configurations through the elastic knowledge automatic learning system. Description of the Drawings

[0061] Figure 1 is a flowchart of the SMS channel fault handling method based on dynamic routing adjustment of the present invention;

[0062] Figure 2 is a comparison of the fault response capabilities of the present invention (bar chart);

[0063] Figure 3 is the effect of the multi-level circuit breaker model of the present invention (radar chart);

[0064] Figure 4 is the result of the system elastic boundary exploration of the present invention (area chart). Detailed Embodiments

[0065] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the protection scope of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.

[0066] At least one embodiment of the present invention discloses a method for handling short message channel failures based on dynamic routing adjustment, as follows: Figure 1 It includes the following steps:

[0067] Step 1, deploy a service mesh layer in the short message system with a microservices architecture. The service mesh layer includes Sidecar proxies associated with each microservice instance and a centralized control plane;

[0068] This step constructs a service mesh layer and its Sidecar proxy system, which is responsible for communication control and fault handling among microservices in the short message system. Specifically, it includes:

[0069] Step 1.1, deploy lightweight Sidecar proxies beside each microservice instance;

[0070] These proxy programs take over all network traffic entering and leaving the microservices, forming a distributed data plane. The Sidecar proxy works with the microservices in the following ways:

[0071] Receive all outbound requests sent from the microservice and forward the requests to the target service according to the current routing policy;

[0072] Receive all inbound requests sent to the microservice and perform operations such as load balancing and circuit breaking protection before forwarding;

[0073] Collect performance metrics and health data of service calls and report them to the control plane in real time.

[0074] Step 1.2, construct a centralized control plane, which is responsible for global routing policy management and service discovery;

[0075] In addition, the control plane includes the following functional units:

[0076] Service registry: maintain real-time registration information of all microservice instances in the system;

[0077] Configuration storage: save global routing rules and fault handling policies;

[0078] Controller: process telemetry data collected from the data plane and dynamically update the routing configuration according to a preset algorithm;

[0079] API gateway: provide configuration query and management interfaces.

[0080] Step 1.3, implement a secure communication system between the control plane and the data plane to ensure that routing policies can be distributed to each Sidecar proxy quickly and accurately;

[0081] It should be noted that this communication system adopts an event-based asynchronous notification mode. When the routing policy changes, the control plane sends updates in real time to all relevant Sidecar proxies through push.

[0082] Step 2: Build a multi-level circuit breaker model based on the service mesh layer. The multi-level circuit breaker model automatically adopts different levels of response strategies according to the fault type and severity. The response strategies include current limiting, partial traffic switching, and full fuse and rerouting.

[0083] Based on the service mesh architecture built in Step 1, this step implements a multi-level circuit breaker model and a real-time traffic control algorithm, which can automatically adopt different levels of response strategies according to the type and severity of faults. Specifically, it includes:

[0084] Step 2.1: Build a multi-level circuit breaker model, which includes three levels of circuit breaker states:

[0085] First-level circuit breaker state: To cope with minor faults, mainly adopt current limiting measures.

[0086] Second-level circuit breaker state: To cope with moderate faults, execute partial traffic switching strategies.

[0087] Third-level circuit breaker state: To cope with severe faults, trigger full fuse and reroute all traffic.

[0088] It should be noted that the state transition of each circuit breaker is based on the following mathematical model:

[0089]

[0090] Among them, represents the current state, represents the next state, represents the set of monitoring metrics, represents the error rate metric, represents the request latency metric, represents the time window parameter, represents the state transition function;

[0091] The specific implementation method of this multi-level circuit breaker model is as follows:

[0092] State transition logic implementation:

[0093] Use a metric-based finite state machine model. Each circuit breaker can be in three basic states: closed, half-open, or open.

[0094] Error rate is calculated using a sliding window algorithm and is calculated by the following formula:

[0095]

[0096] where represents the number of failed requests, represents the total number of requests, and the error threshold is dynamically adjusted according to each level;

[0097] The exponential weighted moving average is used to calculate the response time :

[0098]

[0099] where is the response time smoothing coefficient, represents the current response time, represents the previous average response time;

[0100] The above index combination is evaluated to generate a state transition probability matrix, which determines the upgrade, downgrade or maintenance of the circuit breaker;

[0101] In some embodiments, the state transition logic may further include:

[0102] An anomaly detection model based on time series, using a long short-term memory network (LSTM) to predict the service performance trend and identify potential failures in advance. Optionally, a Bayesian change point detection algorithm is introduced to accurately identify the mutation points of performance indicators and avoid misjudgment caused by short-term fluctuations. The state transition can also consider the business priority factor and configure different trigger thresholds for critical business traffic and non-critical business traffic.

[0103] Implementation of the three levels of the multi-level circuit breaker model:

[0104] First-level circuit breaker: Triggered when the error rate exceeds 5% or the average response time exceeds 200 ms of the warning threshold, start the token bucket rate limiting algorithm to limit the traffic, but do not change the routing;

[0105] Second-level circuit breaker: Triggered when the error rate exceeds 15% or the average response time exceeds 500 ms of the emergency threshold, route 50% of the traffic to the backup channel, and at the same time implement stricter rate limiting on the original channel;

[0106] Third-level circuit breaker: Triggered when the error rate exceeds 30% or the service fails continuously for more than 5 times, completely fuse the original channel and route 100% of the traffic to the healthy channel, and at the same time issue a serious alarm;

[0107] Optionally, the number of levels of the multi - level circuit breaker can be adjusted according to the specific application scenario: in a financial transaction scenario with extremely high availability requirements, a five - level circuit breaker model can be implemented to provide more fine - grained fault response. For simple application scenarios, it can be simplified to a two - level circuit breaker model, which only includes two states: current limiting and fusing. In a resource - constrained edge computing environment, a lightweight implementation can be adopted to reduce the number of states and simplify the transition logic.

[0108] As Figure 2 shown, this is the comparison of the fault response capabilities of the present invention (bar chart), which shows the performance comparison between the traditional solution and the solution of the present invention in three key indicators: fault detection time, channel switching decision time, and complete fault recovery time. The values are normalized values, and the lower the value, the better the performance.

[0109] Step 2.2: Implement a traffic mirroring and real - time analysis module for fault detection and early warning;

[0110] Copy a part of the production traffic (by default 5%) to the mirror channel for analysis;

[0111] Perform multi - dimensional anomaly detection on the mirrored traffic, including response time, error rate, success rate, etc.;

[0112] When an anomaly is detected, trigger an early warning signal and provide a description of the anomaly characteristics.

[0113] Step 2.3: Build an adaptive traffic scheduling algorithm to dynamically adjust the routing rules according to the fault state:

[0114]

[0115] Among them, represents the request ratio assigned to the th channel, represents the health score of the th channel, represents the weight coefficient of the th channel, represents the total request volume, represents the number of available channels, represents the channel index.

[0116] Step 2.4: Implement a real - time update system for the routing rules of the control plane;

[0117] In addition, when a channel anomaly is detected, the control plane will perform the following operations:

[0118] Calculate the new routing rule allocation ratio;

[0119] Generate a routing configuration update instruction;

[0120] Push the update to the relevant Sidecar agent through the asynchronous event system;

[0121] Verify the deployment status of the new rule and collect feedback data.

[0122] Step 3: Based on the response data of the multi-level circuit breaker model, construct an intelligent chaotic experiment system using Bayesian optimization and multi-objective evolutionary algorithms to automatically search for the most informative fault scenarios;

[0123] In this step, an intelligent chaotic experiment system based on the combination of Bayesian optimization and multi-objective evolutionary algorithms is constructed to automatically search for and verify the most informative fault scenarios and improve the resilience of the system. Specifically, it includes:

[0124] Step 3.1: Implement an experimental configuration engine that combines Bayesian optimization and multi-objective evolutionary algorithms;

[0125] It should be noted that this engine automatically searches for the most informative fault scenarios based on the following core algorithms:

[0126]

[0127] Among them, represents the next fault scenario to be tested represents the set of all possible fault scenarios represents the expected improvement function used to evaluate the scenario the possible information value provided represents the scenario the difference from the historical test scenarios represents the set of historical tested scenarios represents the execution scenario the risk assessment value of represents the diversity weight parameter that controls the tendency to explore new scenarios represents the risk weight parameter that controls the degree of avoidance of high-risk scenarios;

[0128] The specific implementation method of this algorithm includes:

[0129] The Bayesian optimization part:

[0130] Establish a Gaussian process regression model to approximate the resilience scoring function, where the resilience score is:

[0131]

[0132] Among them, represents the predicted resilience score of scenario s, represents the time required for the system to recover from a fault to a normal state under scenario s, Indicates the service availability metric of the system under scenario s, usually expressed as a percentage. Indicates the resource utilization efficiency of the system under scenario s. , , Are the weight coefficients of recovery time, service availability, and resource utilization rate respectively;

[0133] The expected improvement sampling criterion is used to select the next test point, and the expected improvement function is defined as:

[0134]

[0135] Among them, Represents the expected improvement value of scenario s. Represents the mathematical expectation. Represents the predicted resilience score of scenario s. Is the current best resilience score. Is the exploration parameter, which controls the balance between exploration and exploitation. Indicates that there is a positive contribution only when the predicted score exceeds the current best score plus the exploration parameter;

[0136] The subsequent experimental scenarios are determined by maximizing the EI value, so as to achieve a balance between exploration (searching for untested areas) and exploitation (improving known good areas);

[0137] In some embodiments, the Bayesian optimization part can have the following variations:

[0138] Optionally, a radial basis function kernel (RBF) or a Markov kernel can be used to define the covariance function of the Gaussian process to adapt to different types of scenario spaces. In an environment with limited computing resources, a sparse Gaussian process approximation method, such as the fully independent training conditions (FITC) algorithm, can be adopted to improve the optimization efficiency. For high-dimensional scenario spaces, dimensionality reduction techniques, such as principal component analysis (PCA) or autoencoders, can be combined to reduce the computational complexity.

[0139] Multi-objective evolutionary algorithm part:

[0140] Maintains a Pareto front composed of non-dominated solutions, and the evaluation objectives include: information gain, operational risk, and experimental cost;

[0141] Uses crossover and mutation operations to generate new candidate failure scenarios;

[0142] Adopts non-dominated sorting and crowding distance calculation for population selection;

[0143] After each iteration, the best scenario is selected from the set of Pareto optimal solutions for testing;

[0144] For example, in the case of strict experimental time constraints, the multi-objective evolutionary algorithm can be adjusted as follows: Use the fast non-dominated sorting algorithm (NSGAII) to accelerate the population evaluation process. Adopt an adaptive mutation rate, using a higher mutation rate in the early stage of the search to promote global exploration and reducing the mutation rate in the later stage for local refinement. Introduce an environmental selection strategy, giving priority to scenarios that are expected to obtain high information benefits within a limited experimental window.

[0145] The fusion methods of the two algorithms:

[0146] Adopt a serial combination strategy, first use Bayesian optimization to generate a set of candidate solutions, and then use the multi-objective evolutionary algorithm to optimize on this basis;

[0147] Implement an alternating execution mechanism, dynamically switch the algorithms used according to the experimental progress and results;

[0148] Share the evaluation model so that the two algorithms can be based on a unified scenario evaluation standard;

[0149] Optionally, the algorithm fusion can also have the following implementation methods: Implement a parallel execution mode, run Bayesian optimization and the multi-objective evolutionary algorithm simultaneously, exchange information regularly and merge their respective optimal solutions. In an environment with sufficient computing resources, an integration method can be adopted, running multiple optimizers with different parameter configurations simultaneously and synthesizing the recommended results of each optimizer. For large-scale distributed systems, a hierarchical optimization strategy can be implemented, first performing local optimization at the subsystem level and then integrating the results at the global level.

[0150] Step 3.2, construct a parametric model of the fault scenario, including the following dimensions:

[0151] Fault type: such as network latency, packet loss, service crash, resource exhaustion, etc.;

[0152] Fault target: such as a specific microservice, network link, storage system, etc.;

[0153] Fault intensity: such as fault duration, impact degree, trigger frequency, etc.;

[0154] Fault combination: the simultaneous or sequential occurrence mode of multiple faults;

[0155] It should be understood that each fault scenario can be represented as a multi-dimensional vector :

[0156]

[0157] where represents the fault type, represents the fault target, represents the fault intensity, represents the fault combination mode.

[0158] Step 3.3, construct an experiment execution engine, which is responsible for converting the parameterized fault scenarios into specific executable fault injection instructions and controlling the execution process of the experiment:

[0159] Experiment preparation: Establish benchmark performance metrics and set monitoring parameters;

[0160] Fault injection: Execute precise fault injection according to the predefined parameters;

[0161] Data collection: Collect system response data, including service-level metrics and business-level metrics;

[0162] Safe termination: Continuously monitor key metrics during the experiment to ensure system safety.

[0163] 3.4 Construct an experiment result analysis module to extract valuable information from the experiment data:

[0164] Evaluation of the effectiveness of routing strategies: Analyze the response effects of routing strategies under various fault scenarios;

[0165] Identification of elastic boundaries: Determine the critical failure points of the system under various fault combinations;

[0166] Discovery of improvement opportunities: Identify the weak links in the current routing strategy;

[0167] In addition, through the constructed intelligent chaos experiment system, it is able to automatically search for and verify the most informative fault scenarios, ensuring the effectiveness of the SMS channel routing strategy in the face of complex faults, thus significantly enhancing the elasticity and reliability of the system.

[0168] As Figure 3 shown, it is the effect of the multi-level circuit breaker model (radar chart) of the present invention, showing the comparison between the traditional single fuse strategy and the multi-level circuit breaker model of the present invention in various key performance metrics. The values are normalized values (in the range of 0 - 1), and the higher the value, the better the performance.

[0169] Step 4, based on the fault scenarios generated by the intelligent chaos experiment system, implement progressive elastic boundary exploration, and detect the system elastic boundary by gradually increasing the fault intensity;

[0170] This step implements a progressive elastic boundary exploration strategy and a supporting safety protection system for exploring the system elastic limit while ensuring system safety and timely intervening in potential serious faults. Specifically, it includes:

[0171] 4.1 Construct a progressive elastic boundary exploration algorithm;

[0172] This algorithm precisely detects the system elastic boundary by gradually increasing the fault intensity and follows the following mathematical model:

[0173]

[0174] Wherein: represents the fault intensity at the current moment, represents the fault intensity at the next moment, represents the basic step size parameter, represents the sigmoid function, which is used to smoothly adjust the step size, represents the current system performance index, represents the performance index threshold;

[0175] The specific implementation method of this algorithm includes:

[0176] Fault intensity progressive growth strategy:

[0177] Use adaptive step size control, and the initial step size is set to a smaller value (for example, for network latency, the initial step size is 10ms);

[0178] Performance index Adopt a composite index, which is weighted and calculated by service availability, response time and resource utilization rate;

[0179] The Sigmoid function is implemented as , where is the slope parameter, which controls the sensitivity;

[0180] When is close to but not reaching the threshold, the step size automatically shrinks to improve the exploration accuracy; when is far from the threshold, the step size increases accordingly to accelerate the exploration;

[0181] In some embodiments, the fault intensity growth strategy can adopt the following variants: Optionally, a piecewise linear function can be used to replace the Sigmoid function, and different step size adjustment strategies are adopted in different performance intervals. For scenarios that require rapid exploration, a binary search strategy can be combined. First, make a large step size jump to find the approximate boundary, and then perform fine exploration. In a resource-constrained environment, a preset discrete fault intensity level can be adopted to avoid the additional computational overhead brought by continuous adjustment.

[0182] Multi-dimensional elastic boundary exploration implementation:

[0183] Define intensity growth curves for different fault types respectively, such as for network latency, CPU limit, memory limit, etc.;

[0184] Implement a hybrid strategy that combines binary search and gradient ascent to quickly locate the elastic critical point;

[0185] For multi-parameter exploration, the alternating fixation method is adopted, that is, only one parameter is changed each time. After finding the boundary points in this dimension, fix this parameter and explore other dimensions;

[0186] Support dynamic re-evaluation of elastic boundaries. When the system configuration or load condition changes, a new round of exploration is automatically triggered;

[0187] For example, when exploring multi-dimensional elastic boundaries, there can be the following implementation variants: Use the Response Surface Methodology to establish a relationship model between performance and multiple failure parameters. Adopt collaborative filtering technology to predict the possible boundaries of unexplored dimensions based on the results of explored dimensions. Prioritize exploring dimensions with high sensitivity shown in historical data to improve exploration efficiency.

[0188] 4.2 Build a multi-dimensional system status monitoring module;

[0189] It should be understood that this module monitors the following key metric categories in real time: Service-level metrics: such as response time, error rate, success rate, etc. Resource-level metrics: such as CPU usage, memory occupancy, network bandwidth, etc. Business-level metrics: such as SMS sending rate, arrival rate, delay distribution, etc. Channel-level metrics: the health status, load condition, error mode, etc. of each channel;

[0190] 4.3 Implement an intelligent security protection system;

[0191] This system includes the following functional units:

[0192] Anomaly detection unit: Analyze monitoring data in real time based on a statistical model to identify potential anomalies;

[0193] Early warning generation unit: Generate hierarchical warning signals when abnormal trends are detected;

[0194] Emergency intervention unit: Automatically terminate the experiment and trigger the recovery process before the system approaches the real failure point;

[0195] Protection boundary unit: Maintain known safe operation boundaries to prevent the experiment from exceeding the safe range;

[0196] Optionally, the intelligent security protection system can have the following implementation variants: In systems with extremely high requirements for high availability, a multi-level early warning strategy can be adopted, setting three levels of alarms: yellow, orange, and red. For anomaly detections of different severities, unsupervised learning methods such as IsolationForest or Local Outlier Factor (LOF) algorithms can be combined to improve the recognition ability of unknown anomaly patterns. When resources permit, predictive security protection can be implemented, predicting the future trend of the system state through a time series prediction model and taking intervention measures in advance. For distributed systems, a hierarchical security protection architecture can be implemented, monitoring and intervening at both the local component and global system levels.

[0197] 4.4 Construct an automatic experimental parameter adjustment algorithm;

[0198] Continuously optimize the experimental parameters according to the system response:

[0199]

[0200] Among them, represents the current experimental parameter set, represents the updated experimental parameter set, represents the learning rate, which controls the step size of parameter update, represents the parameter gradient, that is, the partial derivative of the loss function with respect to the parameter, represents the input features (such as system state, historical data, etc.), represents the target output (such as the expected system response pattern), represents the loss function, which measures the gap between the system response and the expected response under the current parameters;

[0201] In some embodiments, the automatic experimental parameter adjustment algorithm can have the following variants: Optionally, adaptive learning rate algorithms such as AdaGrad or Adam can be adopted to automatically adjust the learning rate according to the historical gradient of the parameters. For complex nonlinear systems, reinforcement learning methods based on policy gradients such as REINFORCE or PPO algorithms can be used to optimize the experimental parameters. An exploration-exploitation balance strategy based on Thompson sampling can be implemented to increase the exploration intensity in areas with high uncertainty.

[0202] Such as Figure 4 shown, is the exploration result of the system elastic boundary of the present invention (area chart), showing the change trends of the system in three key indicators: service availability, response time, and resource utilization during the progressive elastic boundary exploration process as the fault intensity increases. The response time is represented by a normalized value (normalized with respect to the maximum value of 1850ms).

[0203] Step 5: Based on the results of progressive elastic boundary exploration, construct an elastic knowledge automated learning system to extract knowledge from the experimental results and optimize routing rules and circuit breaker configurations;

[0204] In this step, an elastic knowledge automated learning system is constructed, which can automatically extract and apply knowledge from chaotic experimental results, continuously optimize routing rules and circuit breaker configurations, and form a visual elastic knowledge graph. Specifically, it includes:

[0205] 5.1 Implement an automatic analysis algorithm for experimental results;

[0206] This algorithm analyzes experimental data based on machine learning methods to identify the correlation between fault patterns and system responses, following the following model:

[0207]

[0208] Among them, represents the conditional probability that the system generates a response under the fault scenario ; represents the probability that the corresponding fault scenario is when the system response is ; represents the prior probability of the system response ; represents the marginal probability of the fault scenario ;

[0209] The specific implementation method of this algorithm is as follows:

[0210] Preprocessing of experimental data:

[0211] Perform time series alignment on the collected fault experimental data to ensure the time consistency between different metrics;

[0212] Apply sliding window aggregation to convert high-frequency sampling data into meaningful statistical features (mean, variance, peak, etc.);

[0213] Perform outlier detection and filtering to remove noise data and outliers caused by device fluctuations;

[0214] Standardize different types of metric data to make it comparable and suitable for subsequent modeling;

[0215] Mining of fault response patterns:

[0216] Use Bayesian networks to model the probability relationship between faults and system responses and learn the conditional probability table;

[0217] Implement association rule mining algorithms, such as an improved version of the Apriori algorithm, to extract rules in the form of "If fault type A + intensity B, then system response C";

[0218] Adopt time series pattern mining techniques to identify the time series pattern from the occurrence of a fault to the system response;

[0219] Apply clustering analysis to classify similar fault response patterns and form high-level pattern abstractions;

[0220] Evaluation of the effectiveness of resilience strategies:

[0221] Establish a scoring function for each resilience strategy, comprehensively considering the recovery time, degree of service degradation, and resource consumption;

[0222] Implement a comparative analysis function to compare the effectiveness differences of different resilience strategies in dealing with the same fault;

[0223] Build a policy sensitivity analysis module to evaluate the impact degree of policy parameter changes on the results;

[0224] Generate a resilience strategy ranking report to identify the most effective and weakest resilience measures;

[0225] 5.2 Build a routing policy optimization module based on reinforcement learning;

[0226] This module models the routing policy optimization as a Markov decision process and continuously optimizes the routing policy using deep reinforcement learning methods:

[0227]

[0228] Among them, represents the expected reward for executing action in state , represents the learning rate, which controls the speed of Q-value update, represents the immediate reward, which is the feedback obtained from the environment after executing action , represents the discount factor, which controls the importance weight of future rewards, represents the next state after executing action , that is, the new state that the system transfers to, represents all possible actions executable in the next state , represents the maximum Q-value of all possible actions in the next state ;

[0229] 5.3 Build a resilient knowledge graph generation module;

[0230] This module extracts key information from the experimental results and constructs an elastic knowledge graph containing the following elements:

[0231] Node types:

[0232] Fault nodes: Represent specific types of faults or fault combinations;

[0233] Response nodes: Represent the response measures of the system to faults;

[0234] Effect nodes: Represent the actual effects of the response measures;

[0235] Boundary nodes: Represent the critical points of system resilience;

[0236] Edge types:

[0237] Causal edges: Connect fault nodes and response nodes, representing the responses triggered by faults;

[0238] Evaluation edges: Connect response nodes and effect nodes, representing the effectiveness of the responses;

[0239] Association edges: Connect the implicit relationships between different nodes;

[0240] Temporal edges: Represent the temporal order relationship between nodes.

[0241] 5.4 Construct an elastic knowledge application and feedback module:

[0242] It should be noted that this module is responsible for applying the extracted elastic knowledge to the actual system operation and collecting feedback data for continuous improvement:

[0243] Routing rule generation: Generate optimized routing rules based on the elastic knowledge graph;

[0244] Circuit breaker configuration: Automatically adjust the thresholds, timeouts, and retry policies of the circuit breakers;

[0245] Predictive analysis: Predict potential fault scenarios based on historical data and adjust strategies in advance;

[0246] Effect evaluation: Monitor the system performance after application and form a closed-loop feedback.

[0247] Step 6, Apply the optimized routing rules and circuit breaker configurations of the elastic knowledge automated learning system to the service mesh layer to achieve dynamic routing adjustment;

[0248] This step applies the optimized routing rules and circuit breaker configurations generated by the elastic knowledge automated learning system to the service mesh layer to achieve dynamic routing adjustment of the SMS channel, thus forming a complete closed-loop system. Specifically, it includes:

[0249] 6.1 Implement a configuration conversion and verification module;

[0250] This module is responsible for converting abstract routing policies and circuit breaker parameters into configurations executable by the service mesh and validating them before application:

[0251]

[0252] Among them, represents the generated service mesh configuration, represents the configuration conversion function, represents the optimized routing policies and circuit breaker parameters, represents the current environmental status;

[0253] The specific implementation of the configuration conversion and validation module includes:

[0254] Configuration conversion engine:

[0255] Convert the abstract routing policy into service mesh-specific routing rules, including weight assignment, retry policy, timeout setting, etc.;

[0256] Convert the circuit breaker parameters into the fuse configuration of the service mesh, including consecutive error count, error percentage threshold, detection time window, etc.;

[0257] Generate a configuration difference report, clearly identifying the changed content and its expected impact;

[0258] Configuration validation mechanism:

[0259] Static validation: Check the syntax correctness, logical consistency, and resource reference validity of the configuration;

[0260] Simulation validation: Apply the configuration in a simulation environment and verify that its behavior meets expectations;

[0261] Security check: Ensure that configuration changes do not expose the system to security risks or violate security policies;

[0262] 6.2 Implement a progressive configuration deployment system;

[0263] This system adopts the canary release strategy to gradually apply the new configuration to the service mesh, minimizing potential risks:

[0264]

[0265] Among them, represents the deployment configuration at time , represents the initial configuration, represents the target configuration, represents the target time to complete the deployment, represents the time elapsed since the start of the deployment;

[0266] The specific implementation of the progressive configuration deployment system includes:

[0267] Multi-stage deployment strategy:

[0268] The first stage: Apply the new configuration to 5% of the traffic and monitor the changes in key metrics;

[0269] The second stage: If the metrics in the first stage are normal, expand to 25% of the traffic;

[0270] The third stage: If the metrics in the second stage are normal, expand to 75% of the traffic;

[0271] The final stage: Complete the configuration update for 100% of the traffic;

[0272] Automatic rollback mechanism:

[0273] Define the thresholds of key performance indicators (KPIs), such as the error rate increasing by more than 50% and the response time increasing by more than 100%;

[0274] Continuously monitor the changes in KPIs during the deployment process and automatically trigger a rollback when the threshold is exceeded;

[0275] The rollback process adopts a fast switching strategy to give priority to restoring system stability;

[0276] 6.3 Build a configuration effect evaluation system;

[0277] This system is responsible for evaluating the actual effect after the new configuration is applied and providing feedback for the next round of optimization:

[0278]

[0279] Among them, represents the effect score of the configuration change, represents the weight of the th indicator, represents the value of the th indicator under the new configuration, represents the value of the th indicator under the old configuration, represents the total number of evaluation indicators;

[0280] The specific implementation of the configuration effect evaluation system includes:

[0281] Multi-dimensional indicator collection:

[0282] Business indicators: SMS sending success rate, average delivery time, user experience score, etc.;

[0283] System indicators: CPU utilization rate, memory usage rate, network throughput, etc.;

[0284] Elasticity metrics: time to recover from failure, service availability, system stability, etc.;

[0285] A / B test analysis:

[0286] Run the old and new configurations simultaneously on a portion of the traffic and compare the performance differences;

[0287] Adopt the method of statistical significance testing to ensure that the observed differences are statistically significant;

[0288] Generate a detailed comparison report, including the percentage of improvement or degradation of each metric;

[0289] 6.4 Implement a configuration knowledge feedback loop:

[0290] This module feeds back the effect data of the configuration application to the elastic knowledge automated learning system to form a closed-loop optimization:

[0291] Convert the configuration effect evaluation result into a reward signal for the reinforcement learning algorithm;

[0292] Update the elastic knowledge graph to record the association between configuration changes and system behavior;

[0293] Adjust the prior distribution of the Bayesian optimization model to reflect the newly acquired knowledge;

[0294] Trigger a new round of chaos experiments to verify the effectiveness of the updated knowledge.

[0295] In an embodiment of the present invention, an application example of the short message channel fault handling method based on dynamic routing adjustment is provided:

[0296] This example focuses on the application in the transaction verification code short message system of a large financial institution. This financial institution processes approximately 20 million transaction verification code short messages per day, and processes more than 5,000 short message requests per second during peak periods. The system is connected to the channels of three major telecom operators and the channels of two third-party short message service providers, with a total of 8 independent short message channels. The system adopts a microservices architecture based on Kubernetes, including more than 20 microservice components such as short message receiving service, content filtering service, routing distribution service, sending queue service, and status receipt service.

[0297] The main challenges faced by this financial institution include: the short message channels are prone to overload during large promotional activities and require intelligent load balancing; occasional failures of telecom operator channels are frequent and require rapid detection and automatic switching; the system needs to ensure the high reliability and real-time nature of financial transaction verification codes, the sending success rate of verification code short messages is required to reach 99.99%, and the average delivery time does not exceed 3 seconds; traditional manual testing is difficult to cover complex fault combination scenarios, and the system elasticity assessment is insufficient.

[0298] Service Mesh Deployment Instance:

[0299] In the SMS system of this financial institution, an implementation solution based on the Istio service mesh framework is adopted. An Envoy proxy is deployed as a Sidecar container beside each microservice Pod to take over all traffic in and out of the microservices. The centralized control plane is provided by the Istiod component, which is responsible for configuration management, service discovery, and certificate management.

[0300] The control plane and the data plane communicate through the xDS protocol to update the routing rules in real time. To ensure high availability, the control plane is deployed with a 3-node cluster, and through the horizontal auto-scaling function of Kubernetes, the resource quota of the data plane proxy is automatically scaled during peak traffic periods.

[0301] The main configurations of the service mesh include: enabling mutual TLS authentication for all inter-service communications by default; configuring health check probes for each microservice to regularly detect the service health status; setting fine-grained access control policies to restrict the communication permissions between services; configuring the mesh telemetry system to collect key metrics such as service calls, latency, and error rates;

[0302] During the implementation process, through the Sidecar injection mechanism, existing microservices can be connected to the service mesh without modifying the code, achieving a complete decoupling of business logic and communication control.

[0303] Multi-level Circuit Breaker Implementation Instance:

[0304] A customized multi-level circuit breaker model is implemented to meet the high reliability requirements of financial institutions. This model is based on a sliding window to statistically calculate the error rate and response time of the SMS channels, and automatically activates corresponding fault handling strategies according to different levels of exceptions.

[0305] The specific parameter configurations are as follows: First-level circuit breaker: Triggered when the error rate of a certain channel exceeds 3% or the average response time exceeds 1.5 seconds. The system implements token bucket rate limiting for this channel, restricting the sending rate to 80% of the normal value, and sending a yellow warning to the monitoring system. Second-level circuit breaker: Triggered when the error rate exceeds 8% or the average response time exceeds 2.5 seconds for 30 seconds. The system dynamically distributes 40% of the traffic of this channel to other healthy channels and sends an orange warning to the monitoring system. Third-level circuit breaker: Triggered when the error rate exceeds 15% or there are more than 10 consecutive failures. The system completely fuses this channel, routes 100% of the traffic to the backup channel, and sends a red warning to the monitoring system;

[0306] The state recovery strategy adopts a half-open mode: After the circuit breaker channel has been isolated for a period of time (default is 60 seconds), the system will attempt to restore a small amount of traffic (initially 5%) to this channel for testing. If the test traffic performs normally, the traffic ratio will be gradually increased; if there are still abnormalities, the circuit breaker will be reset and the isolation time will be extended.

[0307] The state transition logic of the multi-level circuit breaker is implemented as a finite state machine. The health metrics of each channel are collected through the Prometheus metrics monitoring system, and the exponential weighted moving average algorithm is combined to smooth short-term fluctuations, ensuring that the circuit breaker will not be frequently triggered due to instantaneous jitters.

[0308] Implementation example of the intelligent chaos experiment system:

[0309] In the application of this financial institution, a dedicated chaos engineering test environment is built, isolated from the production environment but with the same architecture and configuration. The intelligent chaos experiment system automatically executes fault injection experiments in this environment to explore the elasticity boundaries of the system.

[0310] The experimental configuration engine is implemented by combining Bayesian optimization and the NSGAII multi-objective evolutionary algorithm. The elasticity scoring function is designed as:

[0311]

[0312] Among them, represents the predicted elasticity score of scenario s, represents the time required for the system to recover from a fault to the normal state, represents the maximum allowed recovery time threshold, represents the service availability metric of the system under a fault scenario, usually expressed as a percentage. The weights of each metric are determined through evaluation by business experts, reflecting the high-priority requirement of the financial transaction SMS system for recovery time.

[0313] The experimental system starts from simple faults and gradually explores complex scenarios. In the initial stage, the system generated and executed approximately 50 single-fault scenarios, including network latency, packet loss, service restart, etc. Based on these results, the system identified that the combination of inter-channel delay differences and bursty traffic is the fault dimension that has the greatest impact on the system.

[0314] In the further exploration stage, the system generated 120 complex fault combination scenarios, mainly concentrated in the following categories: multi-channel simultaneous delay + traffic surge, main channel complete interruption + partial packet loss in the standby channel, cluster node CPU limit + channel response fluctuation, network partition + service instance restart;

[0315] During the execution of each scenario, the security guard system continuously monitors key business metrics. Once approaching the preset thresholds (such as the SMS sending success rate being lower than 98% or the average latency exceeding 5 seconds), the experiment is immediately interrupted and the recovery process is triggered to ensure test security.

[0316] The experimental results show that the system can still operate stably when there is a 400ms delay and a 50% sudden increase in traffic simultaneously in the primary and backup channels, but the quality of service will decline when three main channels fail simultaneously. These findings provide a clear direction for enhancing the resilience of the system.

[0317] Implementation example of the elastic knowledge automated learning system:

[0318] For the chaotic experimental results, the function of automatically constructing an elastic knowledge graph is realized. The system uses Bayesian networks to analyze the causal relationship between faults and system responses, clusters the experimental data, and extracts key patterns.

[0319] The knowledge graph contains four types of core nodes: Fault nodes: Describe the types, locations, and intensities of the injected faults; Response nodes: Record the system's response measures, such as route adjustment, flow limiting, etc.; Effect nodes: Represent the actual effects and performance indicators of the response measures; Boundary nodes: Identify the critical state points of the system's resilience.

[0320] Through six months of continuous learning and optimization, the system automatically identifies multiple key opportunities for enhancing resilience. The most important ones include: For the sudden traffic scenario, preheating the backup channel in advance is more effective than temporary allocation; Channel switching should be progressive rather than instantaneous to reduce the risk of message loss; A channel combination with geographical dispersion provides better resilience than a channel combination concentrated in the same area.

[0321] Based on these findings, the system automatically generates and applies optimized routing rules and circuit breaker configurations. The optimized configurations include: More refined error rate thresholds, dynamically adjusted according to business priorities and time periods; An intelligent load balancing strategy based on the historical performance of the channels; Dedicated fast recovery paths for common fault patterns.

[0322] At the same time, the elastic knowledge graph is presented to the operation and maintenance team through a visualization interface, enabling them to intuitively understand the system's resilience characteristics and potential risk points, and guiding manual intervention decisions.

[0323] The above describes the embodiments of the present invention. However, these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.

Claims

1. A method for handling SMS channel failures based on dynamic routing adjustment, characterized in that: The following steps are involved: Deploy a service mesh layer in the SMS system of the microservice architecture. The service mesh layer includes the sidecar proxy and the centralized control plane deployed in association with each microservice instance. A multi-level circuit breaker model is built based on the service mesh layer. The multi-level circuit breaker model automatically adopts different levels of response strategies according to the fault type and severity. The response strategies include current limiting, partial traffic switching, and complete circuit breaker rerouting. Based on the response data of the multi-level circuit breaker model, an intelligent chaos experiment system is constructed using Bayesian optimization and multi-objective evolutionary algorithms to automatically search for the most informative fault scenarios; Based on the fault scenarios generated by the intelligent chaos experiment system, progressive elastic boundary exploration is achieved, and the elastic boundary of the system is detected by gradually increasing the fault intensity; Based on the results of progressive elastic boundary exploration, an elastic knowledge automatic learning system is built to extract knowledge from experimental results and optimize routing rules and circuit breaker configurations. Apply the routing rules and circuit breaker configuration optimized by the elastic knowledge automated learning system to the service mesh layer to achieve dynamic routing adjustment.

2. The method for handling SMS channel failure based on dynamic routing adjustment according to claim 1, characterized in that: The steps to deploy the service mesh layer include: Deploy a lightweight Sidecar proxy next to each microservice instance to form a data plane responsible for inter-service communication and traffic control; Build a centralized control plane responsible for global routing policy management and service discovery; Implement a secure communication system between the control plane and the data plane to ensure that routing policies can be distributed to each Sidecar proxy quickly and accurately.

3. The method for handling SMS channel failure based on dynamic routing adjustment according to claim 1, characterized in that: The steps to build a multi-level circuit breaker model include: Build state transition logic and use indicator-based finite state machine model to control circuit breaker state transition; The error rate is calculated using a sliding window algorithm, and the request latency index is calculated using an exponentially weighted moving average. Define the trigger conditions and response measures for multiple levels of circuit breaking states. The first-level circuit breaking state triggers current limiting when the error rate exceeds the preset threshold, the second-level circuit breaking state triggers partial traffic switching, and the third-level circuit breaking state triggers complete fusing and rerouting.

4. The method for handling SMS channel failure based on dynamic routing adjustment according to claim 1, characterized in that: The steps to build an intelligent chaos experiment system include: An experimental configuration engine that implements the combination of Bayesian optimization and multi-objective evolutionary algorithms; Construct a parameterized model of fault scenarios, including fault type, fault target, fault intensity, and fault combination dimensions; Build an experiment execution engine to convert parameterized fault scenarios into specific executable fault injection instructions; Build an experimental result analysis module to extract valuable information from experimental data.

5. The method for handling SMS channel failure based on dynamic routing adjustment according to claim 4 is characterized in that: The experimental configuration engine combining Bayesian optimization and multi-objective evolutionary algorithm includes: A Gaussian process regression model is established to approximate the elasticity scoring function; The expected improvement sampling criterion is used to select the next test point; Maintain a Pareto front consisting of non-dominated solutions, and the evaluation objectives include information gain, operation risk, and experimental cost; New candidate fault scenarios are generated using crossover and mutation operations; Bayesian optimization is combined with a multi-objective evolutionary algorithm through a serial combination strategy.

6. The method for handling SMS channel failure based on dynamic routing adjustment according to claim 1, characterized in that: The steps to achieve progressive elastic boundary exploration include: Construct a progressive elastic boundary exploration algorithm to accurately detect the elastic boundary of the system by gradually increasing the fault intensity; Build a multi-dimensional system status monitoring module to monitor service-level indicators, resource-level indicators, business-level indicators, and channel-level indicators in real time; Implement an intelligent security guard system, including anomaly detection unit, warning generation unit, emergency intervention unit and protection boundary unit; Build an automatic adjustment algorithm for experimental parameters to continuously optimize experimental parameters based on system response.

7. The method for handling SMS channel failure based on dynamic routing adjustment according to claim 6, characterized in that: The progressive elastic boundary exploration algorithm includes: Adaptive step size is used to control the increase of fault intensity, and the initial step size is set to a small value; A composite performance index is used to evaluate the system status. When the performance index approaches the threshold, the step size is automatically reduced to improve the exploration accuracy. Define intensity growth curves for different fault types; The alternating fixation method is used for multi-parameter exploration. Only one fault parameter is changed each time. The fault parameters include delay, packet loss rate, and error rate. After finding the elastic boundary point corresponding to the current fault parameter, the value of the current fault parameter is fixed, and the elastic boundaries corresponding to other fault parameters are continued to be explored.

8. The method for handling SMS channel failure based on dynamic routing adjustment according to claim 1, characterized in that: The steps to build an elastic knowledge automated learning system include: Implement an automatic analysis algorithm for experimental results, and analyze experimental data based on machine learning methods to identify the correlation between failure modes and system responses; Construct a routing strategy optimization module based on reinforcement learning and model routing strategy optimization as a Markov decision process; Construct a resilient knowledge graph generation module to extract key information from experimental results and construct a knowledge graph containing fault nodes, response nodes, effect nodes, and boundary nodes; Build an elasticity knowledge application and feedback module to apply the extracted elasticity knowledge to actual system operation and collect feedback data for continuous improvement.

9. The method for handling SMS channel failure based on dynamic routing adjustment according to claim 8, characterized in that: The automatic analysis algorithm of experimental results includes: Perform time series alignment on the collected fault experiment data to ensure the time consistency between different indicators; Apply sliding window aggregation to convert high-frequency sampled data into statistical features; Use Bayesian networks to model the probabilistic relationship between faults and system responses; implement association rule mining algorithms to extract association rules between faults and system responses; A scoring function is established for each resilience strategy, taking into account the recovery time, service degradation, and resource consumption.

10. The SMS channel fault handling system based on dynamic routing adjustment is characterized by: The method for executing the method for handling SMS channel failure based on dynamic routing adjustment according to any one of claims 1 to 9 comprises: The service mesh layer, which includes the sidecar proxies and centralized control plane deployed in association with each microservice instance; Multi-level circuit breaker module, used to automatically adopt different levels of response strategies according to the fault type and severity; Intelligent chaos experiment module, which is used to automatically search for the most informative fault scenarios based on Bayesian optimization and multi-objective evolutionary algorithms; Progressive elastic boundary exploration module, used to detect the elastic boundary of the system by gradually increasing the fault intensity; Elastic Knowledge Automated Learning Module, which is used to extract knowledge from experimental results and optimize routing rules and circuit breaker configurations.

Citation Information

Patent Citations

  • Financial micro-service fault injection detection system based on Istio

    CN116225510A

  • A fault isolation and rapid recovery system based on microservice architecture

    CN119782013A

  • System, method and device for realizing security business real-time self-healing service based on large model technology, processor and readable storage medium thereof

    CN119887394A

  • Microservice hardware and software deployment

    US20230059339A1

Cited By

  • Service orchestration system and optimization method based on rule engine

    CN120353453A

  • Short message marketing effect optimization method and system based on multi-channel dynamic weight

    CN120807020A

  • Land space planning intelligent optimization method and system

    CN121119789A