Automatic loss stopping and back switching method, device and equipment of active-active data center and medium

By constructing an intelligent evaluation model based on mean, standard deviation, and dynamic indicator weights, and a progressive rollback strategy, the inefficiency and high misjudgment problems caused by manual reliance in existing technologies are solved, and efficient and secure disaster recovery switching and rollback of active-active data centers are achieved.

CN121750451APending Publication Date: 2026-03-27FENGLING CHUANGJING (BEIJING) TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

The disaster recovery process of existing active-active data centers relies on manual operation and maintenance, which leads to low switching efficiency, high misjudgment rate, and easy to cause secondary failure risks. Furthermore, the stability of the data center has not been verified.

Method used

By integrating historical multi-dimensional indicator data with real-time multi-dimensional indicator data, an intelligent evaluation model based on mean, standard deviation, and dynamic indicator weights is constructed to accurately identify the data center to be stopped and the target migration data center. Based on the real-time load of the target migration data center, a traffic switching strategy is adaptively generated, and a dynamic observation duration mechanism and a gradual rollback strategy are introduced to ensure a safe rollback time.

Benefits of technology

It significantly improved the automation level and decision accuracy of disaster recovery switching and rollback, reduced the misjudgment rate at the time of rollback, avoided secondary failures and system instability, and ensured business continuity and system resilience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750451A_ABST
    Figure CN121750451A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of operation and maintenance, and discloses an automatic loss stopping and back switching method, device and equipment for an active-active data center and a medium, and the method comprises the steps: determining a mean value, a standard deviation and a dynamic index weight corresponding to each dimension index based on obtained historical multi-dimension index data; determining a stop loss data center and a target migration data center based on each dynamic index weight, the obtained real-time multi-dimensional index data, each mean value and each standard deviation; determining a flow switching strategy based on the real-time load index data of the target migration data center, and switching the flow of the stop loss data center to the target migration data center based on the flow switching strategy; and when the stop loss data center is in a healthy state for the first time, determining a back-switching moment based on the dynamic observation duration, and back-switching the traffic to the stop loss data center by adopting a progressive back-switching strategy at the back-switching moment. According to the scheme, the automation level of disaster recovery switching and back-switching is realized, the dependence of manual intervention is greatly reduced, and the switching and back-switching efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of operation and maintenance technology, and in particular to an automatic loss prevention and switchback method, device, equipment and medium for a dual-active data center. Background Technology

[0002] In existing technologies, disaster recovery for active-active data centers primarily relies on manual intervention in the field of operations and maintenance (O&M). When one data center fails, a full traffic failover strategy is employed, switching all traffic from the failed data center to the other. Furthermore, once the failure is resolved, the health status of the data centers is manually checked, and the timing of the switchback is determined to switch back the traffic from the failed data center to the original failed data center at that specific time. Therefore, current disaster recovery methods for active-active data centers heavily rely on manual intervention, leading not only to low switchback efficiency but also to a high rate of misjudgment regarding the switchback timing, which can easily trigger secondary failures. Summary of the Invention

[0003] The purpose of this invention is to provide at least one method, apparatus, device, and medium for automatic loss prevention and back-off in a dual-active data center. This invention can at least solve the technical problems of poor accuracy and low efficiency in identifying high-risk operations and high delay in blocking high-risk operations when performing operation control. It can at least improve the accuracy and efficiency of identifying high-risk operations and reduce the delay in blocking high-risk operations.

[0004] To address the aforementioned technical problems, at least one embodiment of this application provides an automatic loss mitigation and rollback method for active-active data centers, comprising: acquiring historical multi-dimensional indicator data for each data center over a historical period and real-time multi-dimensional indicator data at the current moment; determining the mean, standard deviation, and dynamic indicator weights corresponding to each dimension indicator based on the historical multi-dimensional indicator data; determining the loss mitigation data center and the target migration data center based on the dynamic indicator weights of each data center, the real-time multi-dimensional indicator data, the mean, and the standard deviation; determining a traffic switching strategy based on the real-time load indicator data of the target migration data center, and switching the traffic of the loss mitigation data center to the target migration data center based on the traffic switching strategy; when the loss mitigation data center is in a healthy state for the first time, determining the rollback time based on the dynamic observation period, and using a gradual rollback strategy at the rollback time to roll back the traffic switched out of the loss mitigation data center to the loss mitigation data center.

[0005] This solution integrates historical and real-time multi-dimensional indicator data from various data centers to construct an intelligent evaluation model based on mean, standard deviation, and dynamic indicator weights. This model accurately identifies the data centers requiring isolation for loss mitigation and the target migration data centers capable of handling the traffic. Based on the real-time load of the target migration data centers, it adaptively generates refined traffic switching strategies, abandoning the traditional "one-size-fits-all" approach to achieve a scientific and stable initial disaster recovery switchover. Once the data centers under loss mitigation recover to a healthy state for the first time, a dynamic observation period mechanism is introduced to fully verify their stability, thereby accurately determining the safe time for rollback. This significantly reduces the misjudgment rate of rollback timing. At this rollback timing, a gradual rollback strategy is adopted to migrate traffic in stages, effectively avoiding secondary failures and system instability caused by hasty or full rollback. Overall, this solution significantly improves the automation level, decision accuracy, and execution security of disaster recovery switchover and rollback, greatly reducing reliance on manual intervention while ensuring business continuity and improving switchover and rollback efficiency.

[0006] This solution integrates historical and real-time multi-dimensional indicator data to construct an intelligent evaluation model based on statistical characteristics (mean, standard deviation) and dynamic weights. This model can accurately identify the data center to be mitigated and the target data center to be migrated. Based on this, a matching traffic switching strategy is adaptively generated according to the real-time load of the target data center to be migrated, avoiding direct full-traffic switching and ensuring a scientific and stable initial disaster recovery switchover. Once the data center to be mitigated recovers, a dynamic observation period mechanism is introduced to fully verify its stability and determine an appropriate switchback time. This allows for a phased switchback strategy to migrate traffic back in stages during safe times, effectively avoiding secondary failures caused by hasty switchbacks and the uncontrollable risks associated with full-traffic switchbacks. This improves switchback efficiency and the accuracy of determining the switchback time.

[0007] In some examples, based on the weights of the dynamic indicators for each of the data centers, the real-time multi-dimensional indicator data, the means, and the standard deviations, the stop-loss data centers and target migration data centers are determined, including: for each data center, determining the deviation standard value corresponding to each dimension indicator based on the real-time multi-dimensional indicator data, the means, and the standard deviations of the data center; performing a weighted summation of the weights of the dynamic indicators for each data center and the deviation standard values ​​to obtain a total deviation value, wherein the dynamic indicator weights correspond one-to-one with the dimension indicators; using the difference between 100 and the total deviation value as the health score of the data center; if the health score is not less than a first health threshold, the data center is considered to be in a healthy state and is a candidate migration data center; if the health score is less than a second health threshold, the data center is considered to be in an unhealthy state and is a stop-loss data center, wherein the first health threshold is greater than or equal to the second health threshold; and selecting any data center from the candidate migration data centers as the target migration data center.

[0008] In some examples, determining the traffic switching strategy based on the real-time load metric data of the target migration data center includes: acquiring the real-time load metric data of the target migration data center; determining the capacity health of the target migration data center based on the load metric data; if the capacity health is not less than a first capacity health threshold, then selecting a full traffic switching strategy; if the capacity health is less than the first capacity health threshold but greater than a second capacity health threshold, then selecting a first tiered rate limiting switching strategy; if the capacity health is not greater than the second capacity health threshold, then selecting a second tiered rate limiting switching and cache release strategy; the rate limiting degree of the second tiered rate limiting switching and cache release strategy is higher than the rate limiting degree of the first tiered rate limiting switching strategy.

[0009] In some examples, the method further includes: real-time detection of the capacity decline status of the target migration data center; if the difference between the capacity health status and the capacity health status at the previous moment is greater than a preset health decline threshold, the target migration data center is considered to be in an abnormal capacity decline state; if the abnormal capacity decline state is the nth consecutive abnormal state, the target migration data center adopts a third-level rate limiting switching strategy; the rate limiting degree of the third-level rate limiting switching strategy is higher than the rate limiting degree of the switching strategy adopted by the target migration data center at the current moment.

[0010] In some examples, the load metric data includes multiple service load metric data that correspond one-to-one with each service. Determining the capacity health of the target migration data center based on the load metric data includes: for each service, obtaining the service priority weight corresponding to the service; determining the total service load corresponding to the service based on the service load metric data; calculating the product of the total service load and the service priority weight, and summing the products to obtain the load value of the target migration data center; and taking the absolute value of the difference between the ratio of the load value and the preset capacity limit and 1 as the capacity health.

[0011] In some examples, determining the cutback time based on the dynamic observation period includes: starting from the current time when the stop-loss data center first enters a healthy state, continuously monitoring the health status of the stop-loss data center for the dynamic observation period; determining the health status of the stop-loss data center at each time based on the multi-dimensional real-time load index data of the stop-loss data center at each time; if no real-time load index data in the multi-dimensional real-time load index data is higher than its corresponding preset threshold, then the stop-loss data center is considered to be in a healthy state at that time; otherwise, the stop-loss data center is considered to be in an unhealthy state at that time; counting the number of times the unhealthy state persists within the dynamic observation period, if the number does not exceed a preset number of abnormalities, then the time after the current time and the time distance from the current time is the dynamic observation period is taken as the cutback time.

[0012] In some examples, the progressive rollback strategy includes multiple rollback stages and sub-rollback strategies corresponding to each rollback stage. The method further includes: for each rollback stage, when executing the sub-rollback strategy of the rollback stage, real-time detection of the service indicator data of the service supported by the traffic rolled back by the sub-rollback strategy; if the service indicator data does not meet the preset service indicator standard, then the execution of the sub-rollback strategy is stopped.

[0013] At least one embodiment of this application also provides an automatic loss mitigation and rollback device for a dual-active data center, comprising: an acquisition unit, configured to acquire historical multi-dimensional indicator data for each data center over a historical period and real-time multi-dimensional indicator data at the current moment; a determination unit, configured to determine the mean, standard deviation, and dynamic indicator weights corresponding to each dimension indicator based on the historical multi-dimensional indicator data; the determination unit is further configured to determine the loss mitigation data center and the target migration data center based on the dynamic indicator weights of each data center, the real-time multi-dimensional indicator data, the mean, and the standard deviation; the determination unit is further configured to determine a traffic switching strategy based on the real-time load indicator data of the target migration data center, and switch the traffic of the loss mitigation data center to the target migration data center based on the traffic switching strategy; and a rollback unit, configured to determine the rollback time based on the dynamic observation duration when the loss mitigation data center is in a healthy state for the first time, and to use a gradual rollback strategy to roll back the traffic switched out of the loss mitigation data center to the loss mitigation data center at the rollback time.

[0014] At least one embodiment of this application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described automatic loss mitigation and rollback method for a dual-active data center.

[0015] At least one embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described automatic loss mitigation and rollback method for a dual-active data center. Attached Figure Description

[0016] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.

[0017] Figure 1 This is a flowchart of an automatic loss mitigation and rollback method for a dual-active data center provided in one embodiment of this application; Figure 2 This is a schematic diagram of a traffic back-cutting method provided in one embodiment of this application; Figure 3 This is a schematic diagram of an automatic loss prevention and back-switch device for a dual-active data center provided in another embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.

[0019] It should be noted that the acquisition or use of data in the embodiments of this application requires the user's consent. The relevant data can only be obtained after the user's authorization, and the acquisition or use of the data complies with the provisions of relevant laws and regulations.

[0020] To facilitate understanding of the embodiments of this application, the relevant content regarding the automatic loss mitigation and rollback method for dual-active data centers will be introduced first.

[0021] In existing technologies, disaster recovery for active-active data centers primarily relies on manual intervention in the field of operations and maintenance (O&M). When one data center fails, a full traffic failover strategy is employed, switching all traffic from the failed data center to the other. After the failure resolves, the health status of the data centers is manually checked, and the timing of the rollback is determined to switch back the traffic from the failed data center to the failed data center at that specific time. Therefore, existing disaster recovery methods for active-active data centers are heavily reliant on manual intervention, leading to low rollback efficiency and a high rate of misjudgment during rollback, potentially causing secondary failures. Furthermore, existing technologies do not perform data center stability verification during the rollback process, directly initiating the rollback, which easily triggers secondary failures.

[0022] To address the aforementioned technical problems, this invention proposes an automatic loss mitigation and rollback method for active-active data centers. The implementation details of this method are described below. The following content is for illustrative purposes only and is not essential for implementing this solution.

[0023] Example 1: The automatic loss prevention and switchback method for dual-active data centers in this embodiment can be applied to electronic devices with communication, computing, and data storage capabilities. The specific process can be as follows: Figure 1 As shown, it includes: Step 110: Obtain historical multi-dimensional indicator data for each data center during historical time periods, as well as real-time multi-dimensional indicator data for the current moment.

[0024] In this context, a data center refers to one of the active-active data centers that provides services, processes business traffic, and has real-time data synchronization capabilities. An active-active data center consists of two or more data centers that simultaneously provide services and process business traffic in parallel. In other words, there is no primary or secondary distinction among the data centers in an active-active data center; all are operational and share the business traffic. If one data center fails, the remaining operational data center can seamlessly take over all business traffic from the failed data center. A historical time period refers to a time interval with the current time as the end time and a preset duration. No point in this time period is later than the current time. The preset duration can be set according to actual needs, such as 5 minutes, 10 minutes, or other durations; this is just an example and not a specific limitation. Historical multi-dimensional indicator data refers to the multi-dimensional indicator data of the data center during a historical time period. Real-time multi-dimensional indicator data refers to the multi-dimensional indicator data of the data center at the current moment. Multi-dimensional metrics include network layer metrics, service layer metrics, and business layer metrics. Network layer metrics include at least latency, packet loss rate, bandwidth, jitter, network device CPU utilization, bandwidth utilization, TCP retransmission rate, and availability. Service layer metrics include at least API success rate, response time, service throughput, application instance CPU utilization, database connection pool utilization, garbage collection frequency, garbage collection duration, and error rate. Business layer metrics include at least transaction volume, business error rate, conversion rate, average order value, and business activity. Metric data refers to the actual data collected for each dimension of the metrics.

[0025] Step 120: Determine the mean, standard deviation, and dynamic indicator weights for each dimension indicator based on historical multi-dimensional indicator data.

[0026] In this context, the mean refers to the average level of a set of data; the standard deviation refers to the dispersion or volatility of a set of data relative to its mean, which can also be understood as the data volatility of a dimensional indicator over a historical period. Dynamic indicator weights represent the importance of a dimensional indicator in assessing the health of a data center. The larger the standard deviation, the larger the dynamic indicator weight, and the more important the corresponding dimensional indicator is in assessing the health of the data center. Dynamic indicator weights are dynamically determined based on historical multi-dimensional indicator data and are not fixed values. A larger dynamic indicator weight corresponds to a more important dimensional indicator.

[0027] Specifically, the mean, standard deviation, and dynamic indicator weights for each dimension indicator are determined based on historical multi-dimensional indicator data. This includes: for each dimension indicator, obtaining multiple indicator data corresponding to the dimension indicator from historical multi-dimensional indicator data, determining the mean and standard deviation of the dimension indicator based on the multiple indicator data, and using the Sigmoid function to calculate and process the standard deviation to obtain the dynamic indicator weights corresponding to the dimension indicator.

[0028] For example, the Sigmoid function is used to calculate the standard deviation to obtain the dynamic indicator weights corresponding to the dimensional indicators, as shown in the following formula: w =

[0029] Where w is the dynamic indicator weight corresponding to the dimension indicator, and x is the standard deviation corresponding to the dimension indicator.

[0030] In addition, dynamic indicator weights are determined based on historical multi-dimensional indicator data, including: for each dimension indicator, obtaining multiple indicator data corresponding to the dimension indicator from historical multi-dimensional indicator data; and using the entropy weight method to calculate and process the multiple indicator data to obtain the dynamic indicator weights corresponding to the dimension indicators.

[0031] Among them, the entropy weight method refers to a method for determining the dynamic weight of a dimensional indicator based on the degree of data fluctuation of that dimensional indicator. The greater the data fluctuation of a dimensional indicator, the greater its dynamic weight, and the more important that dimensional indicator is in assessing the health of a data center. For details on how to use the entropy weight method to calculate and process multiple indicator data to obtain the dynamic weights of dimensional indicators, please refer to existing technologies; further details will not be elaborated here.

[0032] Therefore, dynamically adjusting the weights of dimensional indicators based on their data volatility allows indicators with significant data fluctuations to receive greater attention and influence when assessing data center health. This effectively enhances the sensitivity and responsiveness of the data center health assessment system to abnormal data changes, thereby improving the accuracy of the assessment results.

[0033] Step 130: Based on the weights of each dynamic indicator for each data center, real-time multi-dimensional indicator data, each mean and each standard deviation, determine the data center to be stopped and the target data center to be migrated.

[0034] Among them, the stop-loss data center refers to the data center with a fault in the dual-active data center, and the target migration data center refers to the data center in the dual-active data center that has no fault and is used to take over the traffic of the stop-loss data center.

[0035] Specifically, in step 130 above, the stop-loss data center and the target migration data center are determined based on the weights of each dynamic indicator of each data center, real-time multi-dimensional indicator data, each mean, and each standard deviation. This includes: for each data center, determining the deviation standard value corresponding to each dimension indicator based on the real-time multi-dimensional indicator data, each mean, and each standard deviation of the data center; performing weighted summation on the weights of each dynamic indicator of the data center and each deviation standard value to obtain the total deviation value, wherein the dynamic indicator weights correspond one-to-one with the dimension indicators; using the difference between 100 and the total deviation value as the health score of the data center; if the health score is not less than the first health threshold, the data center is considered to be in a healthy state and is selected as a candidate migration data center; if the health score is less than the second health threshold, the data center is considered to be in an unhealthy state and is selected as the stop-loss data center, wherein the first health threshold is greater than or equal to the second health threshold; and selecting any data center from the candidate migration data centers as the target migration data center.

[0036] The deviation from standard values ​​measures the degree of anomaly in real-time metric data relative to its historical statistical distribution (mean and standard deviation). Total deviation refers to the overall deviation of multiple dimensions of the data center from its historical normal state. A larger total deviation indicates greater fluctuations across multiple dimensions, suggesting a more unstable and unhealthy data center. 100 indicates the data center is in an ideal stable state, or ideally healthy state. The candidate data center for migration refers to the healthy data center within a dual-active data center setup. The first and second health thresholds can be set according to actual needs; for example, the first health threshold could be 80, and the second health threshold could be 60.

[0037] Furthermore, based on the real-time multi-dimensional indicator data, mean, and standard deviation of the data center, the deviation standard value corresponding to each dimension indicator is determined, including: for each dimension indicator, obtaining the real-time indicator data corresponding to the dimension indicator from the real-time multi-dimensional indicator data of the data center; obtaining the mean and standard deviation corresponding to the dimension indicator; and taking the absolute value of the ratio of the difference between the real-time indicator data and the mean to the standard deviation as the deviation standard value.

[0038] For example, the absolute value of the ratio of the difference between the real-time indicator data and the mean to the standard deviation can be used as the deviation from the standard value, as shown in the following formula: Z = | | Where Z represents the deviation from the standard value; B represents the real-time indicator data. The mean; The standard deviation is denoted as .

[0039] Furthermore, the weights of each dynamic indicator of the data center and each deviation standard value are weighted and summed to obtain the total deviation value, including: for each dimension indicator of the data center, obtaining the deviation standard value and dynamic indicator weight corresponding to the dimension indicator; taking the product of the deviation standard value and the dynamic indicator weight as the sub-deviation value corresponding to the dimension indicator; and taking the sum of the sub-deviation values ​​corresponding to each dimension indicator as the total deviation value of the data center.

[0040] Furthermore, if the health score is less than the first health threshold but not less than the second health threshold, the data center is considered to be in a sub-healthy state and is designated as a data center to be observed. Real-time health scores of the data centers to be observed are acquired at each moment within the preset observation period. If the real-time health score is less than the second health threshold, the data center to be observed is designated as a stop-loss data center. If, at no moment within the preset observation period, the real-time health score of the data center to be observed is less than the second health threshold, the data center to be observed is designated as a candidate for migration.

[0041] Among them, the data center under observation refers to the data center in the dual-active data center that may have a fault, and it is necessary to determine whether the data center has a fault based on a period of observation.

[0042] In addition, when the health score is less than the second health threshold, a data center switching event is triggered to execute a traffic switching strategy based on real-time load metric data of the target migration data center, and to switch the traffic of the stop-loss data center to the target migration data center based on the traffic switching strategy.

[0043] In some cases, when the health score is less than the second health threshold, the method also includes performing data analysis on real-time multi-dimensional indicator data based on graph neural networks to locate the root cause of the failure, and triggering predefined operation and maintenance actions based on the root cause of the failure to resolve the failure of the data center and restore the data center to a healthy state.

[0044] For details on how to perform data analysis on real-time multi-dimensional indicator data based on graph neural networks to pinpoint the root cause of the fault, please refer to existing technologies; further details will not be elaborated here.

[0045] Step 140: Determine the traffic switching strategy based on the real-time load metric data of the target migration data center, and switch the traffic of the stop-loss data center to the target migration data center based on the traffic switching strategy.

[0046] Real-time load metrics data refers to the load metrics data collected at the target migration data center at the current moment. Traffic switching strategy refers to the switching strategy that guides the transfer of traffic from the stop-loss data center to the target migration data center. Load metrics include 12 core load metrics from the target migration data center, such as CPU utilization, load queue, memory usage, swap frequency, disk utilization, disk space, remaining disk space, context switch count, active connections, TCP retransmission rate, network throughput, and packet loss rate. Specifically, in step 140 above, determining the traffic switching strategy based on the real-time load metric data of the target migration data center includes: obtaining the real-time load metric data of the target migration data center; determining the capacity health of the target migration data center based on the load metric data; if the capacity health is not less than the first capacity health threshold, then selecting the full traffic switching strategy; if the capacity health is less than the first capacity health threshold but greater than the second capacity health threshold, then selecting the first tiered rate limiting switching strategy; if the capacity health is not greater than the second capacity health threshold, then selecting the second tiered rate limiting switching and cache release strategy; the rate limiting degree of the second tiered rate limiting switching and cache release strategy is higher than the rate limiting degree of the first tiered rate limiting switching strategy.

[0047] Capacity health measures the current resource capacity of the target migration data center. A higher capacity health indicates more abundant resources in the target migration data center, making it more suitable for handling increased traffic. The full traffic switching strategy refers to the scheme where the target migration data center takes over all traffic from the stop-loss data center. The first-tier rate-limiting switching strategy divides the traffic in the stop-loss data center into multiple tiers. For each tier, a portion of the traffic is switched to the target migration data center according to its corresponding traffic proportion, allowing the target migration data center to handle different proportions of traffic at different tiers. The second-tier rate-limiting switching and cache release strategy divides the traffic in the stop-loss data center into multiple tiers. For each tier, a portion of the traffic is switched to the target migration data center according to its corresponding traffic proportion, allowing the target migration data center to handle different proportions of traffic at different tiers, and simultaneously triggers cache release. However, for the same tier of traffic, the proportion of traffic for that tier under the second-tier rate-limiting switching and cache release strategy is less than or equal to the proportion of traffic for that tier under the first-tier rate-limiting strategy. Traffic tiers include core business traffic, important business traffic, and ordinary business traffic. Core business traffic consists of requests for core services, important business traffic consists of requests for important services, and ordinary business traffic consists of requests for ordinary services. The importance and priority of core services are higher than those of important services; conversely, the importance and priority of important services are higher than those of ordinary services. Rate limiting refers to the proportion of traffic requests rejected by the target migration data center from the data center that is being migrated to.

[0048] For example, with a first capacity health threshold of 30% and a second capacity health threshold of 10%, if the capacity health is greater than or equal to 30%, a full traffic switching strategy is adopted, switching all traffic from the stop-loss data center to the target migration data center. If the capacity health is less than 30% but greater than 10%, a first-tier rate limiting strategy is adopted, switching 100% of the core business traffic in the stop-loss data center to the target migration data center, switching 80% of the important business traffic in the stop-loss data center to the target migration data center (that is, the target migration data center takes over 80% of the important business traffic in the stop-loss data center and rejects the 20% of important business traffic requested by the stop-loss data center), and switching 50% of the ordinary business traffic in the stop-loss data center to the target migration data center. If the capacity health is less than or equal to 10%, a two-tier rate limiting switching and cache release strategy is adopted, switching 100% of the core business traffic in the stop-loss data center to the target migration data center; switching 60% of the important business traffic in the stop-loss data center to the target migration data center; switching 30% of the ordinary business traffic in the stop-loss data center to the target migration data center, and simultaneously releasing the cache in the target migration data center.

[0049] Furthermore, the load metric data includes multiple service load metric data that correspond one-to-one with each service. Based on the load metric data, the capacity health of the target migration data center is determined, including: obtaining the service priority weight corresponding to each service; determining the total service load corresponding to each service based on the service load metric data; calculating the product of the total service load and the service priority weight, and summing the products to obtain the load value of the target migration data center; and using the absolute value of the difference between the load value and the preset capacity limit and 1 as the capacity health.

[0050] The business is divided into core transactions, important queries, and ordinary tasks. Core transactions correspond to the first business priority weight, important queries to the second business priority weight, and ordinary tasks to the third business priority weight. The first business priority weight is greater than the second business priority weight, and the second business priority weight is greater than the third business priority weight. The total business load is the sum of all resources currently used by the business, and the total business load corresponds one-to-one with the business. The preset capacity limit refers to the maximum business load that the target migration data center can safely and stably bear under the current resource configuration and operating conditions.

[0051] For example, calculate the product of the total business load and the business priority weight, and sum the products to obtain the load value of the target migration data center, as shown in the following formula: Load value = ∑(Total business load × Business priority weight) The total workload includes the total workload corresponding to core transactions, the total workload corresponding to important queries, and the total workload corresponding to ordinary tasks. The business priority weight includes a first business priority weight, a second business priority weight, and a third business priority weight. If the total workload equals the total workload corresponding to core transactions, the business priority weight is the first business priority weight; if the total workload equals the total workload corresponding to important queries, the business priority weight is the second business priority weight; and if the total workload equals the total workload corresponding to ordinary tasks, the business priority weight is the third business priority weight.

[0052] Secondly, the absolute value of the difference between the load value and the preset capacity limit and 1 is used as the capacity health status, as shown in the following formula: Capacity health = 1 - load value / preset capacity limit Furthermore, the method also includes: real-time detection of the capacity decline status of the target migration data center; if the difference between the capacity health status and the capacity health status at the previous moment is greater than the preset health decline threshold, the target migration data center is considered to be in an abnormal capacity decline state; if the abnormal capacity decline state is the nth consecutive abnormal state, the target migration data center adopts a third-level rate limiting switching strategy; the rate limiting degree of the third-level rate limiting switching strategy is higher than the rate limiting degree of the switching strategy adopted by the target migration data center at the current moment.

[0053] The switching strategy can be a full-traffic switching strategy, a first-level rate limiting switching strategy, or a second-level rate limiting switching and cache release strategy.

[0054] For example, if the preset health decline threshold is 5% and n is 3, the switching strategy adopted by the target migration data center at the current time is the first-level rate limiting switching strategy. If the difference between the capacity health at the current time and the capacity health at the previous time is greater than 5% for three consecutive times, the third-level rate limiting switching strategy is immediately adopted. At this time, the third-level rate limiting switching strategy can be the second-level rate limiting switching and cache release strategy.

[0055] In addition, the method includes: real-time monitoring of the capacity status of the target migration data center; if the capacity health is greater than the capacity safety threshold, the target migration data center adopts a phased traffic rollback strategy. The phased traffic rollback strategy refers to periodically increasing the proportion of traffic taken over by the target migration data center according to a preset cycle.

[0056] For example, if the target migration data center is currently handling 80% of the business traffic, and the capacity health of the target migration data center is greater than the capacity safety threshold in the next moment, then starting from the next moment, the proportion of business traffic handled by the target migration data center will be increased by 10% every 10 seconds. Here, business traffic refers to the sum of core business traffic, ordinary business traffic, and important business traffic.

[0057] Therefore, by monitoring the capacity health trend of the target migration data center in real time, enhanced rate limiting strategies can be triggered in advance when the capacity significantly declines. Compared with traditional solutions that rely solely on static thresholds, this proactively avoids overload risks at least 30% earlier. Simultaneously, after capacity recovers to a safe level, a phased and gradual increase in traffic takeover ratio is adopted to achieve a smooth transition, effectively avoiding secondary system instability caused by sudden traffic surges. Overall, this method significantly improves the stability, security, and automation of disaster recovery switching and rollback processes, ensuring business continuity while enhancing the system's resilience and adaptability under high load scenarios.

[0058] Step 150: When the stop-loss data center is in a healthy state for the first time, the switchback time is determined based on the dynamic observation period, and a gradual switchback strategy is adopted at the switchback time to switch back the traffic switched out from the stop-loss data center to the stop-loss data center.

[0059] Specifically, in step 150 above, determining the cutback time based on the dynamic observation period includes: starting from the current time when the stop-loss data center is first in a healthy state, continuously monitoring the health status of the stop-loss data center for a dynamic observation period; determining the health status of the stop-loss data center at each time based on the multi-dimensional real-time load index data of the stop-loss data center at each time; if there is no real-time load index data in the multi-dimensional real-time load index data that is higher than its corresponding preset threshold, then the stop-loss data center is considered to be in a healthy state at that time; otherwise, the stop-loss data center is considered to be in an unhealthy state at that time; counting the number of times the unhealthy state persists within the dynamic observation period, if the number does not exceed the preset number of abnormalities, then the time after the current time and the time from the current time is the dynamic observation period is taken as the cutback time.

[0060] Furthermore, if the number of anomalies exceeds the preset number within the dynamic observation period, the dynamic observation period is extended. This extended observation period continues to monitor the health status of the data center and counts the number of times unhealthy states persist within the extended observation period. If this number does not exceed the preset number of anomalies, the time after the current moment, where the time interval from the current moment equals the total observation period, is used as the cutoff point. The total observation period is the sum of the dynamic observation period and the extended observation period.

[0061] Furthermore, in step 150 above, the progressive rollback strategy includes multiple rollback stages and sub-rollback strategies corresponding to each rollback stage. The method also includes: for each rollback stage, when executing the sub-rollback strategy of the rollback stage, real-time detection of the service indicator data of the service supported by the traffic rolled back by the sub-rollback strategy; if the service indicator data does not meet the preset service indicator standard, then the execution of the sub-rollback strategy is stopped.

[0062] The gradual rollback strategy includes at least three rollback phases: the first phase rolls back 10% of read-only traffic; the second phase rolls back 30% of read / write traffic; and the third phase rolls back 100% of all traffic. Rollback refers to switching back traffic migrated from the stop-loss data center to the target migration data center back to the stop-loss data center. Business metrics can include success rate, latency, etc.

[0063] Furthermore, if the business metric data is the business success rate, then if the business success rate is lower than the preset success rate threshold, the business metric data is considered to fail to meet the preset business metric standards. If the business metric data is the delay duration, then if the delay duration is higher than the preset delay duration, the business metric data is considered to fail to meet the preset business metric standards.

[0064] Specifically, while stopping the execution of the sub-switchback strategy, the switchback traffic is rolled back to the target migration data center. For example, when the securities system switches back, it automatically stops and switches back to the target migration data center because the order success rate drops to 99.2%.

[0065] In addition, when a new fault is detected in the data center under damage control, the sub-switchback strategy should be stopped immediately and the operations and maintenance personnel should be notified to handle the situation.

[0066] For example, such as Figure 2 As shown, once the stop-loss data center recovers to a healthy state, a dynamic observation period is initiated to continuously monitor its health status and collect business indicator data in real time. If the business indicator data does not meet the preset business indicator standards, the sub-switchback strategy is stopped. If the business indicator data meets the preset business indicator standards, the sub-switchback strategy continues to be executed.

[0067] Therefore, by introducing a dynamic observation period mechanism, the health and stability of the data center under control are fully and continuously verified, effectively avoiding the risk of repeated business interruptions and secondary failures caused by hasty rollback. Simultaneously, combined with a three-stage gradual rollback strategy (e.g., 10% read-only → 30% read-write → full rollback), system performance is monitored in real time during the gradual recovery of traffic. Once an anomaly is detected, the rollback can be immediately stopped and rolled back, minimizing potential risks within a controllable range. This solution significantly improves the security and reliability of the rollback operation, fundamentally solving the hidden danger of secondary failures easily caused by traditional "one-time full rollback," ensuring a smooth business transition and high system availability.

[0068] In summary, this solution acquires historical multi-dimensional indicator data for each data center over historical periods, as well as real-time multi-dimensional indicator data at the current moment. Based on the historical multi-dimensional indicator data, it determines the mean, standard deviation, and dynamic indicator weights for each dimension indicator. Based on the dynamic indicator weights for each data center, the real-time multi-dimensional indicator data, the mean, and the standard deviation, it identifies the stop-loss data center and the target migration data center. Based on the real-time load indicator data of the target migration data center, it determines the traffic switching strategy and switches the traffic from the stop-loss data center to the target migration data center based on the traffic switching strategy. When the stop-loss data center is in a healthy state for the first time, it determines the back-switch time based on the dynamic observation period and uses a gradual back-switch strategy to switch the traffic switched out from the stop-loss data center back to the stop-loss data center. By integrating the historical multi-dimensional indicator data and the current real-time multi-dimensional indicator data of each data center, an intelligent evaluation model based on the mean, standard deviation, and dynamic indicator weights is constructed to accurately identify the stop-loss data center that needs to be isolated and the target migration data center that can handle the traffic. Based on the real-time load of the target migration data center, it adaptively generates a refined traffic switching strategy, abandoning the traditional "one-size-fits-all" approach and achieving a scientific and stable initial disaster recovery switchover. Once the data center recovers to a healthy state for the first time, a dynamic observation period mechanism is introduced to fully verify its stability, thereby accurately determining the safe time for switchback. This significantly reduces the misjudgment rate of switchback timing. At this switchback timing, a gradual switchback strategy is adopted to migrate traffic in stages, effectively avoiding secondary failures and system instability caused by hasty or full switchbacks. Overall, this solution significantly improves the automation level, decision accuracy, and execution security of disaster recovery switching and switchback, greatly reducing reliance on manual intervention while ensuring business continuity and improving switching and switchback efficiency.

[0069] Example 2: Another embodiment of this application relates to an automatic loss prevention and rollback device for a dual-active data center. The implementation details of this embodiment are described below. The following content is only for ease of understanding and is not essential for implementing this solution. A schematic diagram of the automatic loss prevention and rollback device 30 for the dual-active data center in this embodiment can be seen as follows: Figure 3 As shown, it includes an acquisition unit 301, a determination unit 302, and a back-cut unit 303.

[0070] The acquisition unit 301 is used to acquire historical multi-dimensional indicator data for historical periods of each data center and real-time multi-dimensional indicator data for the current moment.

[0071] The determining unit 302 is used to determine the mean, standard deviation and dynamic indicator weights of each dimension indicator based on the historical multi-dimensional indicator data.

[0072] The determining unit 302 is further configured to determine the stop-loss data center and the target migration data center based on the dynamic indicator weights of each of the data centers, the real-time multi-dimensional indicator data, the mean and the standard deviation of each of the data centers.

[0073] The determining unit 302 is further configured to determine a traffic switching strategy based on the real-time load index data of the target migration data center, and switch the traffic of the stop-loss data center to the target migration data center based on the traffic switching strategy.

[0074] The back-off unit 303 is used to determine the back-off time based on the dynamic observation time when the stop-loss data center is in a healthy state for the first time, and to use a gradual back-off strategy to back-off the traffic cut out by the stop-loss data center to the stop-loss data center at the back-off time.

[0075] In some examples, when determining the stop-loss data center and the target migration data center based on the dynamic indicator weights of each of the data centers, the real-time multi-dimensional indicator data, the mean and the standard deviation of each data center, the determining unit 302 is specifically used for: determining the deviation standard value corresponding to each dimension indicator for each data center based on the real-time multi-dimensional indicator data, the mean and the standard deviation of each data center; performing weighted summation on the dynamic indicator weights and the deviation standard values ​​of each data center to obtain the total deviation value, wherein the dynamic indicator weights correspond one-to-one with the dimension indicators; using the difference between 100 and the total deviation value as the health score of the data center; if the health score is not less than a first health threshold, the data center is considered to be in a healthy state and is a candidate migration data center; if the health score is less than a second health threshold, the data center is considered to be in an unhealthy state and is a stop-loss data center, wherein the first health threshold is greater than or equal to the second health threshold; and selecting any data center among the candidate migration data centers as the target migration data center.

[0076] In some examples, when determining the traffic switching strategy based on the real-time load index data of the target migration data center, the determining unit 302 is specifically used to: acquire the real-time load index data of the target migration data center; determine the capacity health of the target migration data center based on the load index data; if the capacity health is not less than a first capacity health threshold, then select a full traffic switching strategy; if the capacity health is less than the first capacity health threshold but greater than a second capacity health threshold, then select a first tiered rate limiting switching strategy; if the capacity health is not greater than the second capacity health threshold, then select a second tiered rate limiting switching and cache release strategy; the rate limiting degree of the second tiered rate limiting switching and cache release strategy is higher than the rate limiting degree of the first tiered rate limiting switching strategy.

[0077] In some examples, the switchback unit 303 is also used to: detect the capacity decline status of the target migration data center in real time; if the difference between the capacity health status and the capacity health status at the previous moment is greater than a preset health decline threshold, then the target migration data center is considered to be in an abnormal capacity decline state; if the abnormal capacity decline state is the nth consecutive abnormal state, then the target migration data center adopts a third-level rate limiting switching strategy; the rate limiting degree of the third-level rate limiting switching strategy is higher than the rate limiting degree of the switching strategy adopted by the target migration data center at the current moment.

[0078] In some examples, when determining the capacity health of the target migration data center based on the load index data, the determining unit 302 is used to: obtain the service priority weight corresponding to each service for each service; determine the total service load corresponding to each service based on the service load index data; calculate the product of the total service load and the service priority weight, and sum the products to obtain the load value of the target migration data center; and take the absolute value of the difference between the ratio of the load value and the preset capacity limit and 1 as the capacity health.

[0079] In some examples, the back-cut unit 303, when used to determine the back-cut time based on the dynamic observation duration, is specifically used to: start from the current time when the stop-loss data center is first in a healthy state, continuously monitor the health status of the stop-loss data center for the dynamic observation duration; determine the health status of the stop-loss data center at the current time based on the multi-dimensional real-time load index data of the stop-loss data center at each time; if there is no real-time load index data higher than its corresponding preset threshold in the multi-dimensional real-time load index data, then the stop-loss data center is considered to be in a healthy state at that time; otherwise, the stop-loss data center is considered to be in an unhealthy state at that time; count the number of times the unhealthy state persists within the dynamic observation duration; if the number does not exceed the preset number of abnormalities, then the time after the current time and the time distance from the current time is the dynamic observation duration is taken as the back-cut time.

[0080] In some examples, the progressive rollback strategy includes multiple rollback stages and sub-rollback strategies corresponding to each rollback stage. The rollback unit 303 is further configured to: for each rollback stage, when executing the sub-rollback strategy of the rollback stage, detect in real time the service indicator data of the service supported by the traffic rolled back by the sub-rollback strategy; if the service indicator data does not meet the preset service indicator standard, stop executing the sub-rollback strategy.

[0081] It is worth mentioning that all units involved in this embodiment are logical units. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.

[0082] Example 3: Another embodiment of this application relates to an electronic device, such as... Figure 4 As shown, it includes: at least one processor 901; and a memory 902 communicatively connected to the at least one processor 901; wherein the memory 902 stores instructions executable by the at least one processor 901, the instructions being executed by the at least one processor 901 to enable the at least one processor 901 to execute the automatic loss mitigation and rollback method for dual-active data centers in the above embodiments.

[0083] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.

[0084] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.

[0085] Example 4: Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.

[0086] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0087] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.

Claims

1. An automatic loss prevention and rollback method for a dual-active data center, characterized in that, include: Obtain historical multi-dimensional indicator data for each data center during historical time periods, as well as real-time multi-dimensional indicator data for the current moment; Based on the historical multi-dimensional indicator data, determine the mean, standard deviation, and dynamic indicator weights for each dimension indicator. Based on the weights of the dynamic indicators for each of the data centers, the real-time multi-dimensional indicator data, the mean and the standard deviation are used to determine the stop-loss data center and the target migration data center. Based on the real-time load metric data of the target migration data center, a traffic switching strategy is determined, and based on the traffic switching strategy, the traffic of the stop-loss data center is switched to the target migration data center. When the stop-loss data center is in a healthy state for the first time, the switchback time is determined based on the dynamic observation period, and a gradual switchback strategy is adopted at the switchback time to switch back the traffic cut out from the stop-loss data center to the stop-loss data center.

2. The automatic loss prevention and rollback method for a dual-active data center according to claim 1, characterized in that, Based on the weights of the dynamic indicators for each of the data centers, the real-time multi-dimensional indicator data, the means, and the standard deviations of each indicator determine the stop-loss data center and the target migration data center, including: For each of the aforementioned data centers, the deviation from the standard value corresponding to each dimension indicator is determined based on the real-time multi-dimensional indicator data of the data center, each mean, and each standard deviation. The weights of each dynamic indicator in the data center and each deviation standard value are weighted and summed to obtain the total deviation value, wherein the weights of the dynamic indicators correspond one-to-one with the dimensional indicators; The difference between 100 and the total deviation value is used as the health score of the data center; If the health score is not less than the first health threshold, the data center is considered to be in a healthy state and is a candidate data center for migration. If the health score is less than the second health threshold, the data center is considered to be in an unhealthy state, and the data center is a stop-loss data center, wherein the first health threshold is greater than or equal to the second health threshold; Choose any one of the candidate migration data centers as the target migration data center.

3. The automatic loss prevention and rollback method for a dual-active data center according to claim 1, characterized in that, The process of determining the traffic switching strategy based on the real-time load metric data of the target migration data center includes: Obtain real-time load metric data for the target migration data center; The capacity health of the target migration data center is determined based on the load index data. If the capacity health status is not less than the first capacity health threshold, then the full traffic switching strategy is selected; If the capacity health status is less than the first capacity health threshold and greater than the second capacity health threshold, then the first tiered rate limiting switching strategy is selected. If the capacity health status is not greater than the second capacity health threshold, then the second-level rate limiting switching and cache release strategy is selected; The rate limiting level of the second-level rate limiting switching and cache release strategy is higher than that of the first-level rate limiting switching strategy.

4. The automatic loss prevention and rollback method for a dual-active data center according to claim 3, characterized in that, The method further includes: Real-time monitoring of capacity degradation in the target migration data center; If the difference between the capacity health status and the capacity health status at the previous moment is greater than the preset health status decline threshold, then the target migration data center is considered to be in an abnormal capacity decline state. If the capacity reduction anomaly is the nth consecutive anomaly, the target migration data center will adopt the third-level rate limiting and switching strategy. The rate limiting level of the third-tier rate limiting switching strategy is higher than the rate limiting level of the switching strategy adopted by the target migration data center at the current moment.

5. The automatic loss prevention and rollback method for a dual-active data center according to claim 3, characterized in that, The load metric data includes multiple service load metric data that correspond one-to-one with each service. Determining the capacity health of the target migration data center based on the load metric data includes: For each service, obtain the service priority weight corresponding to that service; The total business load corresponding to the business is determined based on the business load index data; Calculate the product of the total service load and the service priority weight, and sum the products to obtain the load value of the target migration data center; The absolute value of the difference between the ratio of the load value and the preset capacity limit and 1 is taken as the capacity health.

6. The automatic loss prevention and rollback method for a dual-active data center according to claim 1, characterized in that, The determination of the cut-off time based on dynamic observation duration includes: Starting from the current moment when the stop-loss data center first enters a healthy state, the health status of the stop-loss data center is continuously monitored for the specified dynamic observation period; The health status of the stop-loss data center at each time point is determined based on the multi-dimensional real-time load index data of the stop-loss data center at that time point. If no real-time load indicator data in the multi-dimensional real-time load indicator data is higher than its corresponding preset threshold, then the stop-loss data center is considered to be in a healthy state at that time; otherwise, the stop-loss data center is considered to be in an unhealthy state at that time. The number of times the unhealthy state persists within the dynamic observation period is counted. If the number of times does not exceed the preset number of abnormalities, the time after the current time and the time distance from the current time is the dynamic observation period is taken as the cut-off time.

7. The automatic loss prevention and rollback method for a dual-active data center according to claim 1, characterized in that, The progressive back-cut strategy includes multiple back-cut stages and sub-back-cut strategies corresponding to each back-cut stage. The method further includes: For each back-switching stage, when executing the sub-back-switching strategy of the back-switching stage, the service indicator data of the service supported by the traffic backed by the sub-back-switching strategy is detected in real time; If the business indicator data does not meet the preset business indicator standards, the sub-switch strategy will be stopped.

8. An automatic loss prevention and switchback device for a dual-active data center, characterized in that, include: The acquisition unit is used to acquire historical multi-dimensional indicator data for historical periods of each data center and real-time multi-dimensional indicator data for the current moment. The determining unit is used to determine the mean, standard deviation, and dynamic indicator weights of each dimension indicator based on the historical multi-dimensional indicator data. The determining unit is further configured to determine the stop-loss data center and the target migration data center based on the dynamic indicator weights of each of the data centers, the real-time multi-dimensional indicator data, the mean and the standard deviation of each of the data centers; The determining unit is further configured to determine a traffic switching strategy based on the real-time load index data of the target migration data center, and switch the traffic of the stop-loss data center to the target migration data center based on the traffic switching strategy. The back-off unit is used to determine the back-off time based on the dynamic observation period when the stop-loss data center is in a healthy state for the first time, and to use a gradual back-off strategy to switch the traffic cut out by the stop-loss data center back to the stop-loss data center at the back-off time.

9. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the automatic stop-loss and rollback method for a dual-active data center as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the automatic loss mitigation and rollback method for a dual-active data center as described in any one of claims 1 to 7.