Disaster recovery scheduling method and system for financial-grade PAAS platform based on unitization architecture

By receiving fault alarm signals in a financial-grade PaaS platform, collecting and analyzing unit data, quantifying and calculating center-of-gravity parameters, dynamically filtering target units, and formulating traffic switching strategies, the problem of disaster recovery scheduling relying on manual intervention and insufficient correlation assessment in existing technologies is solved, thereby achieving automation of disaster recovery scheduling and improvement of business continuity.

CN121567703BActive Publication Date: 2026-04-24SHANGHAI NEWTOUCH SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI NEWTOUCH SOFTWARE CO LTD
Filing Date
2026-01-26
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing financial-grade PaaS platforms based on modular architecture suffer from several drawbacks in disaster recovery scheduling, including reliance on manual intervention for disaster recovery scheduling decisions, insufficient representation and evaluation of disaster recovery relationships between units, neglect of the overall balance of the disaster recovery system, and lack of dynamic adaptability in traffic switching strategies. These issues lead to business interruptions and a decline in user experience.

Method used

By receiving fault alarm signals, collecting and aggregating real-time resource load status data and business traffic data of all business units within the platform, analyzing the disaster recovery correlation between the main unit in the same city, the backup unit in the same city, and the disaster recovery unit in a different location, quantifying and calculating the center of gravity parameters, dynamically selecting target takeover units, and formulating and executing business traffic switching strategies.

Benefits of technology

It achieves full-process automation and intelligence in disaster recovery scheduling, accurately represents the disaster recovery association status between core reference entities, selects target takeover units more scientifically and rationally, and ensures smooth and stable business traffic switching, thereby ensuring rapid and uninterrupted recovery of core financial businesses and improving the platform's disaster recovery reliability and business continuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121567703B_ABST
    Figure CN121567703B_ABST
Patent Text Reader

Abstract

The application provides a financial-grade PAAS platform disaster recovery scheduling method and system based on a unit architecture, relates to the technical field of computer software, and comprises the following steps: step 1, receiving a failure alarm signal of a main unit; collecting and converging real-time resource load state data and real-time service traffic data of all service units in the platform according to the failure alarm signal; and step 2, performing load adaptability analysis on all healthy service units at present according to the real-time resource load state data and the real-time service traffic data, selecting a city main unit, a city backup unit and a remote disaster recovery unit in the platform as three core reference entities, analyzing a disaster recovery correlation between the city main unit, the city backup unit and the remote disaster recovery unit and performing parameterization characterization, so as to obtain a parameterized characterization result and a corresponding operation state level. The application realizes fast, smooth and uninterrupted disaster recovery takeover of financial core services under the condition of unit failure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer software technology, and in particular to a disaster recovery scheduling method and system for a financial-grade PaaS platform based on a modular architecture. Background Technology

[0002] In the process of digital transformation in finance, financial-grade PaaS platforms based on modular architecture are widely used to support core financial businesses such as payment settlement and credit approval. Their disaster recovery and scheduling capabilities are directly related to the continuity and security of financial businesses.

[0003] In a PaaS platform deployed with a unitized architecture, when the main business unit responsible for processing personal online banking transactions experiences a sudden failure, the existing disaster recovery scheduling solution first manually assesses the scope of the failure, then collects resource load data from some healthy business units, and selects a local backup unit for business takeover. However, during the switchover process, due to the lack of systematic analysis of the disaster recovery correlation between the local main backup unit and the remote disaster recovery unit, and the failure to consider the overall balance of the disaster recovery system, the selected takeover unit, although having a low load, experiences implicit delays in data synchronization with other disaster recovery units. After the business traffic switchover, some transactions exhibit abnormal responses. Furthermore, due to the lack of a global dynamic adjustment mechanism, it cannot adapt to real-time changes in business traffic in a timely manner, ultimately causing short-term business interruption and a decline in user experience. This case exposes technical flaws in the existing technology, such as reliance on manual intervention in disaster recovery scheduling decisions, insufficient representation and evaluation of disaster recovery correlations between units, neglect of the global balance of the disaster recovery system, and a lack of dynamic adaptability in traffic switching strategies. These flaws make it difficult to meet the high reliability and timeliness requirements of financial-grade business for disaster recovery scheduling. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a disaster recovery scheduling method and system for a financial-grade PaaS platform based on a modular architecture, so as to realize rapid, stable and uninterrupted disaster recovery takeover of core financial businesses in the event of unit failure.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] Firstly, a disaster recovery scheduling method for a financial-grade PaaS platform based on a modular architecture, the method comprising:

[0007] Step 1: Receive the fault alarm signal from the main unit; based on the fault alarm signal, collect and aggregate real-time resource load status data and real-time business traffic data of all business units within the platform;

[0008] Step 2: Based on real-time resource load status data and real-time business traffic data, perform load adaptability analysis on all current healthy business units, select the same-city main unit, same-city backup unit and off-site disaster recovery unit in the platform as three core reference entities, analyze the disaster recovery correlation between the same-city main unit, same-city backup unit and off-site disaster recovery unit and perform parameterized characterization to obtain the parameterized characterization results and the corresponding operating status level.

[0009] Step 3: The results of parameterized characterization and the corresponding operating status level are converted into a feature point set. By calculating the feature point set, the center of gravity parameter reflecting the disaster recovery correlation balance status is obtained. The center of gravity parameter is used as a dynamic adjustment value. The load adaptability analysis results of each health business unit are corrected by the dynamic adjustment value to obtain a comprehensive evaluation result.

[0010] Step 4: Based on the comprehensive evaluation results, dynamically select the target takeover unit from all healthy business units;

[0011] Step 5: Based on the target takeover unit, formulate and execute scheduling instructions to switch service traffic to the target takeover unit according to the service type corresponding to the fault alarm signal and real-time service traffic data, so as to complete the disaster recovery scheduling.

[0012] Furthermore, it receives fault alarm signals from the main unit; based on the fault alarm signals, it collects and aggregates real-time resource load status data and real-time service traffic data from all business units within the platform, including:

[0013] Receive fault alarm signals sent by the main unit and parse the fault alarm signals to determine the unique identifier of the service unit that has failed and the corresponding fault level;

[0014] Based on the fault unit identifier and fault level, all healthy business units within the platform that need to be evaluated are dynamically analyzed and screened to form a target unit set for data collection. The target unit set is determined based on healthy units that have direct business call relationships, data synchronization relationships, and resource pool sharing relationships with the fault master unit.

[0015] Based on the target unit set, a data collection request is synchronously initiated to each healthy business unit within the set, and real-time resource load status data and real-time business traffic data returned by each unit are received.

[0016] Furthermore, step 2 above includes:

[0017] Based on real-time resource load status data and real-time business traffic data, the initial load adaptability score of each healthy business unit in the platform is calculated.

[0018] Based on the initial load adaptability score and combined with the predefined geographical location attributes of each business unit, three core reference entities are selected from all healthy business units: the unit with the highest score in the same city area is determined as the main unit in the same city, the unit with the second highest score in the same city area is determined as the backup unit in the same city, and the unit with the highest score in a different region is determined as the disaster recovery unit in a different region.

[0019] Based on the same-city main unit, same-city backup unit and off-site disaster recovery unit, real-time operation and maintenance indicators among the same-city main unit, same-city backup unit and off-site disaster recovery unit are extracted and quantitatively analyzed to form a disaster recovery correlation indicator set including data synchronization delay, network transmission bandwidth utilization and service switching time cost.

[0020] By standardizing the disaster recovery correlation index set and assigning weight coefficients to each index according to the pre-configured strategy for weighted calculation, a quantitative parameterized representation result is obtained.

[0021] By combining the obtained parameterized characterization results with the real-time resource load status data of the core reference entities, the current operating status of each core reference entity in the disaster recovery system is evaluated and determined.

[0022] Furthermore, the results of the parameterized representation and the corresponding operational status levels are transformed into a feature point set. By calculating this feature point set, a centroid parameter reflecting the disaster recovery equilibrium status is obtained, and this centroid parameter is used as a dynamic adjustment value. This dynamic adjustment value is then used to correct the load adaptability analysis results of each healthy business unit, resulting in a comprehensive evaluation result, including:

[0023] Using the parameterized representation results as the coordinate reference, and mapping the operational status level of each core reference entity to a weight factor, the three core reference entities—the same-city main unit, the same-city backup unit, and the off-site disaster recovery unit—are each represented as a multi-dimensional feature point in the disaster recovery association space.

[0024] Based on multidimensional feature points, according to their respective weighting factors, the weighted geometric centroid of the multidimensional feature points in the disaster recovery association space is calculated, and the coordinate value of the weighted geometric centroid is defined as the centroid parameter reflecting the current disaster recovery association equilibrium state among core reference entities.

[0025] The center of gravity parameter is used as a global dynamic adjustment value to correct the initial load adaptability score of each health business unit.

[0026] Based on the correction logic, the dynamic adjustment value is integrated with the initial load adaptability score of each healthy business unit to generate a comprehensive evaluation result that simultaneously reflects the individual load capacity of the unit and the overall disaster recovery system balance.

[0027] Furthermore, based on multi-dimensional feature points and their respective weighting factors, the weighted geometric centroid of each multi-dimensional feature point in the disaster recovery association space is calculated. The coordinate values ​​of the weighted geometric centroid are defined as centroid parameters reflecting the current disaster recovery association equilibrium state among core reference entities, including:

[0028] By obtaining the multidimensional feature points and corresponding running state weight factors of each generated core reference entity;

[0029] Based on the coordinate values ​​of each multidimensional feature point and the associated weighting factors, the weighted coordinate components of each feature point in each coordinate dimension of the disaster recovery association space are calculated respectively.

[0030] The weighted coordinate components of all core reference entities in each coordinate dimension are summed and divided by the sum of all weight factors to obtain the final coordinate value of the weighted geometric centroid in the disaster recovery association space.

[0031] The calculated final coordinate value is defined as the centroid parameter, which characterizes the real-time disaster recovery association balance among the three units: the primary unit in the same city, the backup unit in the same city, and the disaster recovery unit in another location.

[0032] Furthermore, based on the comprehensive evaluation results, target takeover units are dynamically selected from all health business units, including:

[0033] Based on the comprehensive evaluation results and a preset takeover capability threshold, the obtained comprehensive evaluation results are screened, and healthy business units with evaluation values ​​below the threshold are removed to obtain a preliminary candidate unit set.

[0034] Based on the preliminary candidate unit set, and considering the business type and resource requirements of the faulty main unit, a business and resource matching degree analysis is performed on each unit in the preliminary candidate unit set. The units are then sorted according to their matching degree to generate a sorted candidate unit list.

[0035] Based on the sorted candidate unit list, the candidate unit with the highest matching degree and whose comprehensive evaluation results meet the final conditions is selected as the target takeover unit.

[0036] Furthermore, based on the target takeover unit, and according to the service type corresponding to the fault alarm signal and real-time service traffic data, a scheduling instruction for switching service traffic to the target takeover unit is formulated and executed to complete disaster recovery scheduling, including:

[0037] Based on the target takeover unit identifier, and combined with the service type parsed from the fault alarm signal and real-time service traffic data, a detailed service traffic switching strategy is formulated, including the switching time window, traffic ratio and routing strategy.

[0038] Based on the detailed business traffic switching strategy, generate a dispatchable instruction containing specific execution parameters and execution sequence;

[0039] The scheduling command is sent to the platform's traffic scheduling component for execution, so as to switch the business traffic originally directed to the faulty main unit to the target takeover unit and complete the disaster recovery scheduling.

[0040] Secondly, a disaster recovery and scheduling system for a financial-grade PaaS platform based on a modular architecture includes:

[0041] The acquisition module is used to receive fault alarm signals from the main unit; based on the fault alarm signals, it collects and aggregates real-time resource load status data and real-time business traffic data of all business units within the platform.

[0042] The analysis module is used to perform load adaptability analysis on all healthy business units based on real-time resource load status data and real-time business traffic data. It selects the same-city main unit, same-city backup unit and off-site disaster recovery unit in the platform as three core reference entities, analyzes the disaster recovery correlation between the same-city main unit, same-city backup unit and off-site disaster recovery unit and performs parameterized characterization to obtain the parameterized characterization results and the corresponding operating status level.

[0043] The evaluation module is used to convert the results of parameterized characterization and the corresponding operating status level into a feature point set. By calculating the feature point set, the center of gravity parameter reflecting the disaster recovery correlation balance status is obtained, and the center of gravity parameter is used as a dynamic adjustment value. The load adaptability analysis results of each health business unit are corrected by the dynamic adjustment value to obtain a comprehensive evaluation result.

[0044] The screening module is used to dynamically select target takeover units from all healthy business units based on comprehensive evaluation results;

[0045] The processing module is used to formulate and execute scheduling instructions for switching service traffic to the target takeover unit based on the service type corresponding to the fault alarm signal and real-time service traffic data, so as to complete disaster recovery scheduling.

[0046] Thirdly, a computing device, comprising:

[0047] One or more processors;

[0048] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0049] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0050] The above-described solution of the present invention has at least the following beneficial effects:

[0051] This system employs a technical approach that involves accurately filtering target unit sets after fault alarm analysis, synchronously collecting and aggregating real-time resource load and business traffic data, selecting local primary units, local backup units, and off-site disaster recovery units as core reference entities, and quantitatively analyzing their disaster recovery correlations to achieve parameterized representation. The parameterized results and operational status levels are then transformed into feature point sets to calculate center-of-gravity parameters, thereby correcting load adaptability scores and obtaining comprehensive evaluation results. Through hierarchical screening, matching degree analysis, and ranking, target takeover units are determined, and a refined traffic switching strategy including switching time windows, traffic ratios, and routing policies is formulated and executed. This approach overcomes the data acquisition blindness in traditional financial-grade PaaS platform disaster recovery scheduling. The previous approach addressed technical issues such as inefficiency, lack of quantitative disaster recovery correlation assessment, neglect of global disaster recovery balance in load adaptability assessment, lack of standardized processes for selecting target takeover units, and crude business traffic switching strategies. These issues led to reliance on manual intervention in disaster recovery decisions, delayed responses, low reliability of business switching, and susceptibility to business interruptions. The solution aims to improve the automation and intelligence of the entire disaster recovery scheduling process, accurately represent the disaster recovery correlation status between core reference entities, select more scientific and reasonable target takeover units, and ensure smooth and stable business traffic switching. This effectively guarantees the rapid recovery and uninterrupted operation of core financial businesses after unit failures, thereby enhancing the platform's disaster recovery reliability and business continuity. Attached Figure Description

[0052] Figure 1 This is a flowchart illustrating the disaster recovery scheduling method for a financial-grade PaaS platform based on a modular architecture, provided by an embodiment of the present invention.

[0053] Figure 2 This is a schematic diagram of a financial-grade PaaS platform disaster recovery and scheduling system based on a modular architecture, provided by an embodiment of the present invention. Detailed Implementation

[0054] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0055] like Figure 1 As shown, embodiments of the present invention propose a disaster recovery scheduling method for a financial-grade PaaS platform based on a modular architecture. The method includes the following steps:

[0056] Step 1: Receive the fault alarm signal from the main unit; based on the fault alarm signal, collect and aggregate real-time resource load status data and real-time business traffic data of all business units within the platform;

[0057] Step 2: Based on real-time resource load status data and real-time business traffic data, perform load adaptability analysis on all current healthy business units, select the same-city main unit, same-city backup unit and off-site disaster recovery unit in the platform as three core reference entities, analyze the disaster recovery correlation between the same-city main unit, same-city backup unit and off-site disaster recovery unit and perform parameterized characterization to obtain the parameterized characterization results and the corresponding operating status level.

[0058] Step 3: The results of parameterized characterization and the corresponding operating status level are converted into a feature point set. By calculating the feature point set, the center of gravity parameter reflecting the disaster recovery correlation balance status is obtained. The center of gravity parameter is used as a dynamic adjustment value. The load adaptability analysis results of each health business unit are corrected by the dynamic adjustment value to obtain a comprehensive evaluation result.

[0059] Step 4: Based on the comprehensive evaluation results, dynamically select the target takeover unit from all healthy business units;

[0060] Step 5: Based on the target takeover unit, formulate and execute scheduling instructions to switch service traffic to the target takeover unit according to the service type corresponding to the fault alarm signal and real-time service traffic data, so as to complete the disaster recovery scheduling.

[0061] In this embodiment of the invention, the technical means of receiving fault alarm signals, collecting and aggregating real-time resource load and business traffic data of all platform business units, selecting three core reference entities through load adaptability analysis and parameterizing their disaster recovery correlation, converting the representation results and operating status levels into feature point set calculation center parameters as dynamic adjustment values ​​to correct the load adaptability analysis results to obtain a comprehensive evaluation result, dynamically selecting target takeover units based on the comprehensive evaluation results, and formulating and executing business traffic switching scheduling instructions in combination with fault business types and real-time traffic data, overcome the technical problems of existing technologies such as disaster recovery scheduling decisions relying on manual intervention, insufficient representation and evaluation of disaster recovery correlation between units, neglect of the global balance of the disaster recovery system, and lack of dynamic adaptability of traffic switching strategies. This achieves the technical effects of realizing fully automated decision-making in disaster recovery scheduling, accurately quantifying and representing the disaster recovery correlation between units, ensuring the global balance of the disaster recovery system, improving the dynamic adaptability and stability of traffic switching, avoiding business interruption and user experience degradation, ensuring the continuity and security of the core business of the financial-grade PaaS platform, and meeting the high reliability and timeliness requirements of financial-grade business for disaster recovery scheduling.

[0062] In a preferred embodiment of the present invention, step 1 above may include:

[0063] Step 1.1: Receive fault alarm signals sent by the main unit and parse the fault alarm signals to determine the unique identifier of the faulty business unit and its corresponding fault level. Specifically, this includes: continuously monitoring the running status message queues of all business units within the platform; when the main unit responsible for core business processing experiences hardware, software, or network failures, the fault monitoring sensors built into the main unit will automatically generate fault alarm signals; receiving fault alarm signals in real time; and parsing the data packets of the fault alarm signals layer by layer. First, extract the unit identity field from the data packet. The unit identity field contains the rack number, cabinet number, and unit number of the business unit. The unique identifier of the faulty business unit can be determined through the unit identity field. Then, parse the fault severity field in the data packet and determine the fault level according to the preset fault level classification standard. The fault levels are divided into four levels: Level 1 fault corresponds to a state where the unit is completely paralyzed and cannot process any business requests; Level 2 fault corresponds to a state where the unit's performance drops sharply and its processing capacity decreases by more than 50%; Level 3 fault corresponds to a state where the unit has partial functional abnormalities that only affect some business processing; and Level 4 fault corresponds to a state where the unit has minor alarms that do not affect core business processing. The unique identifier of the faulty unit and its corresponding fault level are obtained through the parsing results.

[0064] Step 1.2: Based on the fault unit identifier and fault level, dynamically analyze and filter all healthy business units within the platform that need to be evaluated, forming a target unit set for data collection. The target unit set is determined based on healthy units that have direct business call relationships, data synchronization relationships, and resource pool sharing relationships with the faulty main unit. Specifically, this includes retrieving the platform's pre-set unit association mapping table, which pre-stores the business call relationships, data synchronization relationships, and resource pool sharing relationships between all business units within the platform. The unique identifier of the determined fault unit is input into the unit association mapping table, and dynamic analysis is performed in conjunction with the obtained fault level. When the fault level is level one or two... The analysis scope is expanded to include upstream sending units and downstream receiving units that have direct business call relationships with the faulty main unit, backup data units and transaction record units that have real-time data synchronization relationships, and computing resource units, storage resource units, and network resource units that have resource pool sharing relationships. When the fault level is level three or four, the analysis scope is narrowed to healthy units that only have core business call relationships and core data synchronization relationships with the faulty main unit. All healthy business units within the platform are matched and screened one by one, and healthy units that have no relationship with the faulty main unit are eliminated. The healthy business units that meet the screening conditions are summarized to form a target unit set for data collection.

[0065] Step 1.3: Based on the target unit set, synchronously initiate data collection requests to each healthy business unit within the set, and receive real-time resource load status data and real-time business traffic data returned by each unit. Specifically, this includes: generating a data collection task list based on the determined target unit set, the data collection task list containing unique identifiers for all healthy business units within the target unit set; synchronously initiating data collection requests to each healthy business unit in the list based on the platform's internal dedicated monitoring data transmission protocol, the collection requests specifying the types and collection frequencies of real-time resource load status data and real-time business traffic data to be collected; real-time resource load status data includes CPU utilization, memory utilization, disk I / O utilization, and network bandwidth utilization; real-time business traffic data includes the number of transactions processed per second, the number of concurrent user connections, the business request response time, and the business request success rate; setting a data collection timeout of 3 seconds, waiting for each healthy business unit to return data within the timeout period, marking healthy business units that do not return data within the timeout as abnormal data collection units and removing them from the target unit set; receiving and aggregating the data returned by each healthy business unit, and storing the aggregated real-time resource load status data and real-time business traffic data in the platform's temporary data cache.

[0066] In this embodiment of the invention, a technical approach is adopted to first analyze the fault alarm signal to accurately determine the unique identifier and fault level of the fault unit, then dynamically filter healthy business units to form a target unit set based on the fault identifier and level combined with business call relationship, data synchronization relationship and resource pool sharing relationship, and finally synchronously initiate data collection requests to the target unit set and aggregate real-time resource load and business traffic data. This approach overcomes the technical problems of unclear fault location, blind data collection scope leading to low collection efficiency and invalid data interfering with decision-making, and lack of targeting in healthy unit selection in traditional disaster recovery scheduling. As a result, it achieves accurate and efficient fault location, strong targeting of data collection with improved efficiency and data quality, and a target unit set that meets the disaster recovery assessment requirements.

[0067] In a preferred embodiment of the present invention, step 2 above may include:

[0068] Step 2.1: Based on real-time resource load status data and real-time business traffic data, calculate the initial load adaptability score for each healthy business unit within the platform. This specifically includes: retrieving real-time resource load status data and real-time business traffic data for each healthy business unit stored in the temporary data cache, and determining a two-dimensional scoring rule that includes both resource load and business traffic dimensions. The resource load dimension score accounts for 60%, covering four sub-indicators: CPU utilization, memory utilization, disk I / O utilization, and network bandwidth utilization. Each sub-indicator accounts for 15%. A maximum score of 15 points is awarded for each sub-indicator when CPU utilization is below 70%, memory utilization is below 65%, disk I / O utilization is below 60%, and network bandwidth utilization is below 50%. For each sub-indicator value exceeding the threshold by 5 points... For each percentage point deducted, 3 points are deducted, until all points are deducted. The business traffic dimension accounts for 40% of the score, covering four sub-indicators: transaction processing per second, concurrent user connections, business request response time, and business request success rate. Each sub-indicator accounts for 10%. A sub-indicator can receive a full score of 10 points when the transaction processing per second is 30% higher than the current business demand, the concurrent user connections are 50% lower than the unit's maximum capacity, the business request response time is less than 200 milliseconds, and the business request success rate is higher than 99.9%. For each sub-indicator value deviating from the threshold by 5%, 2 points are deducted, until all points are deducted. By scoring each sub-indicator of each healthy business unit one by one, and then weighting and summing them according to the dimension's proportion, the initial load adaptability score of each healthy business unit is finally obtained. The total score is set at 100 points.

[0069] Step 2.2: Based on the initial load adaptability score and combined with the predefined geographical location attributes of each business unit, select three core reference entities from all healthy business units: the unit with the highest score within the same city is designated as the primary unit within the same city; the unit with the second highest score within the same city is designated as the backup unit within the same city; and the unit with the highest score in a different location is designated as the off-site disaster recovery unit. Specifically, this includes: retrieving the predefined geographical location attribute information of all healthy business units within the platform. This geographical location attribute information is pre-stored in the unit's basic information database. The criteria for dividing the same-city and different-city areas are then determined. The same-city area is defined as those located within the same city... Business units within the same data center park in the city are defined as those located in different data center parks in different provinces. All healthy business units are categorized into local and remote regions. Within each region, healthy business units are ranked from highest to lowest based on their initial load adaptability score. For healthy business units within the same city, the unit ranked first is designated as the primary local unit, and the unit ranked second is designated as the backup local unit. For healthy business units within a remote region, the unit ranked first is designated as the remote disaster recovery unit. These three units are marked as the core reference entities of the disaster recovery system.

[0070] Step 2.3: Based on the same-city primary unit, same-city backup unit, and off-site disaster recovery unit, extract and quantify the real-time operation and maintenance indicators among the same-city primary unit, same-city backup unit, and off-site disaster recovery unit to form a disaster recovery correlation indicator set including data synchronization latency, network transmission bandwidth utilization, and service switching time cost. Specifically, for the selected three core reference entities—the same-city primary unit, same-city backup unit, and off-site disaster recovery unit—three types of real-time operation and maintenance indicators are continuously extracted at a collection frequency of once every 5 seconds. The first type of indicator is data synchronization latency, which is calculated by comparing the business data consistency timestamps among the three core reference entities in real time to determine the time it takes for data to be transmitted from the source unit to the target unit. The first category of indicators is the network transmission bandwidth utilization rate. This is calculated by monitoring the real-time transmission bandwidth data of the dedicated communication links between the three core reference entities, calculating the ratio of real-time transmission bandwidth to the total link bandwidth, and taking the average value after collecting data 10 times consecutively. The second category of indicators is the service switching time cost. This is calculated by simulating the entire process of switching service traffic from the same-city primary unit to the same-city backup unit and the off-site disaster recovery unit, recording the total time from initiating the switching command to the complete restoration of normal service, and taking the average value after simulating the switching 5 times consecutively. These three categories of indicators are then summarized to form a disaster recovery correlation indicator set.

[0071] Step 2.4 involves standardizing the disaster recovery correlation index set and assigning weight coefficients to each index according to a pre-configured strategy for weighted calculation, thereby obtaining a quantified parameterized representation result. Specifically, this includes: using extreme value standardization to process the values ​​of each index in the disaster recovery correlation index set, mapping all index values ​​to a range of 0 to 100 points. Lower data synchronization latency values ​​correspond to higher standardized scores, network transmission bandwidth utilization values ​​are highest when they are between 40% and 60%, and lower service switching time cost values ​​correspond to higher standardized scores. After standardization, weight coefficients are assigned to each index according to the pre-configured index weight allocation strategy: 40% for data synchronization latency, 30% for network transmission bandwidth utilization, and 30% for service switching time cost. The standardized scores of each index are then weighted and summed according to these weight coefficients. The resulting weighted sum is the quantified parameterized representation result, with a total score set to 100 points.

[0072] Step 2.5: Integrate the obtained parameterized characterization results with the real-time resource load status data of the core reference entities to evaluate and determine the current operational status of each core reference entity within the disaster recovery system. Specifically, this includes: retrieving the obtained parameterized characterization results and simultaneously retrieving the real-time resource load status data of the three core reference entities; determining the evaluation rules for the operational status of the core reference entities; and setting the operational status evaluation criteria into three levels: good, average, and poor. The evaluation criteria for a good level are: a parameterized characterization score higher than 80 points and the core reference entity's CPU utilization, memory utilization, disk I / O utilization, and network bandwidth utilization... The utilization rate of all four resource load indicators is below 70%; the evaluation criteria for the general level are that the parameterized characterization result is between 60 and 80 points and the values ​​of all four resource load indicators of the core reference entity are below 80%; the evaluation criteria for the poor level are that the parameterized characterization result is below 60 points or any one of the resource load indicators of the core reference entity is above 80%; the parameterized characterization result and real-time resource load status data of each core reference entity are substituted into the evaluation rules to complete the operation status level determination of the same-city main unit, same-city backup unit, and off-site disaster recovery unit one by one, and finally output the current operation status of each core reference entity in the disaster recovery system.

[0073] In this embodiment of the invention, the technical means of quantitatively calculating the initial load adaptability score of each healthy business unit based on real-time resource load and business traffic data, accurately selecting the same-city main unit, same-city backup unit, and off-site disaster recovery unit as core reference entities by combining the score with geographical location attributes, extracting and quantitatively analyzing real-time operation and maintenance indicators among core reference entities to form a disaster recovery correlation indicator set, and obtaining a quantitative parameterized representation result by weighted calculation after standardization of the indicator set, and comprehensively evaluating the operating status of the core reference entities by combining the result with the real-time resource load data of the core reference entities, overcomes the technical problems of traditional disaster recovery scheduling, such as lack of quantitative basis for load adaptability assessment, lack of standard for core disaster recovery unit selection, lack of systematic analysis and quantitative representation of disaster recovery correlation between units, and one-sided evaluation of the operating status of core disaster recovery units. Thus, it achieves scientific and reasonable load adaptability assessment, accurate selection of core reference entities to meet disaster recovery needs, clear and quantifiable disaster recovery correlation, and comprehensive and accurate evaluation of the operating status of core reference entities.

[0074] In a preferred embodiment of the present invention, step 3 above may include:

[0075] Step 3.1: Using the parameterized representation results as the coordinate reference, and mapping the operational status level corresponding to each core reference entity to a weight factor, the three core reference entities—the same-city main unit, the same-city backup unit, and the off-site disaster recovery unit—are each represented as a multi-dimensional feature point in the disaster recovery association space. Specifically, this includes: first, retrieving the obtained quantitative parameterized representation results, which contain the quantitative values ​​of three core indicators: data synchronization latency, network transmission bandwidth utilization, and service switching time cost. These three indicators are then mapped to the three coordinate dimensions of the disaster recovery association space: data synchronization latency corresponds to the X-axis, network transmission bandwidth utilization corresponds to the Y-axis, and service switching time cost corresponds to the Z-axis. The specific values ​​of each indicator in the parameterized representation results are then used as... The baseline coordinate values ​​are used for the corresponding coordinate dimensions. Subsequently, the operating status levels of each core reference entity are retrieved, and a mapping table between operating status levels and weight factors is established. The weight factor corresponding to the good level is set to 1.2, the weight factor corresponding to the average level is set to 1.0, and the weight factor corresponding to the poor level is set to 0.8. The parameterized representation results of each core reference entity are decomposed into dimensions, and the index values ​​corresponding to each coordinate dimension are extracted as the coordinate components of that dimension. Combined with the weight factors obtained from the mapping, the same-city main unit, same-city backup unit, and off-site disaster recovery unit are represented as a three-dimensional feature point in the disaster recovery association space. The coordinate values ​​of each feature point are composed of the three index values ​​of the corresponding core reference entity, and the corresponding weight factors are bound to them.

[0076] Step 3.2: Based on the multi-dimensional feature points and their respective weighting factors, calculate the weighted geometric centroid of the multi-dimensional feature points in the disaster recovery association space. Define the weighted geometric centroid coordinates as centroid parameters reflecting the current disaster recovery association equilibrium state among the core reference entities. Specifically, this includes: first, synchronously acquiring the coordinates of the three-dimensional feature points corresponding to the three generated core reference entities and their bound running status weighting factors, establishing a list of feature points and weighting factors to ensure that the coordinate dimensions of each feature point accurately match the weighting factors; then, based on this list, calculate the weighted coordinate components for each feature point's three coordinate dimensions. The calculation method is to multiply the coordinate value of each dimension by the corresponding weighting factor. For example, if the X-axis coordinate value of the main unit in the same city is 30 and the weighting factor is 1.2, then the weighted X-axis coordinate component of this unit is 30 × 1.2 = 36. This process is repeated to calculate the weighted components of all coordinate dimensions for the three feature points. Then, the weighted coordinate components of the three core reference entities in the same coordinate dimension are summarized, i.e., all weighted components on the X-axis, Y-axis, and Z-axis are summed, while simultaneously calculating the sum of the weight factors for the three core reference entities. Next, the summarized weighted components for each coordinate dimension are divided by the sum of the weight factors to obtain the final X-axis, Y-axis, and Z-axis coordinate values ​​of the weighted geometric centroid within the disaster recovery association space. For example, if the summarized X-axis component is 100 and the total weight is 3.2, then the final X-axis coordinate value is 100 ÷ 3.2 = 31.25. Finally, the coordinate data formed by combining the three sets of final coordinate values ​​is defined as the centroid parameter representing the real-time disaster recovery association equilibrium state among the same-city primary unit, same-city backup unit, and off-site disaster recovery unit.

[0077] Step 3.3 uses the centroid parameter as a global dynamic adjustment value to correct the initial load adaptability score of each healthy business unit. Specifically, this includes: pre-setting the balanced coordinate range of the disaster recovery association space, where the X-axis balanced range is 40 to 60, the Y-axis balanced range is 40 to 60, and the Z-axis balanced range is 35 to 55. When the centroid parameter coordinate value obtained in Step 3.2 is within this range, the current disaster recovery association system is determined to be in a balanced state; if it exceeds this range, it is determined to be in an unbalanced state. The centroid parameter is used as a global dynamic adjustment value. First, calculate the deviations of the center of gravity parameters from the center values ​​of the balanced coordinate range (X-axis 50, Y-axis 50, Z-axis 45). The deviation is calculated by subtracting the center value of the corresponding dimension from the center of gravity parameter coordinate value. Then, determine the adjustment range based on the magnitude and direction of the deviation. Set the absolute value of the dynamic adjustment value to increase by 0.5 for every unit the deviation exceeds the boundary of the balanced range. If the deviation is positive, the adjustment value is negative; if the deviation is negative, the adjustment value is positive. This achieves a reverse correction of the initial load adaptability score, ensuring that the corrected score reflects the balanced requirements of the overall disaster recovery system.

[0078] Step 3.4: Based on the correction logic, the dynamic adjustment value is integrated with the initial load adaptability score of each healthy business unit to generate a comprehensive evaluation result that simultaneously reflects the individual load capacity of the unit and the overall disaster recovery system balance. Specifically, this includes: first, obtaining the correction logic, i.e., adding the initial load adaptability score of each healthy business unit to the corresponding dynamic adjustment value to obtain a preliminary correction score; simultaneously, setting scoring constraints: if the preliminary correction score is higher than 100, it is taken as 100; if it is lower than 0, it is taken as 0, ensuring that the final comprehensive evaluation result is within a reasonable range; subsequently, retrieving the obtained... The initial load adaptability score of each health business unit, and the dynamic adjustment value of the associated storage, are used to calculate the score fusion for each health business unit one by one according to the correction logic. For example, if the initial score of a health business unit is 85 points and the corresponding dynamic adjustment value is -3 points, then the preliminary corrected score is 82 points, which meets the constraints and is directly used as the comprehensive evaluation result. If another unit has an initial score of 98 points and a dynamic adjustment value of 4 points, then the preliminary corrected score is 102 points, and 100 points is taken as the comprehensive evaluation result according to the constraints. Finally, a list containing all health business unit identifiers and corresponding comprehensive evaluation results is generated.

[0079] In this embodiment of the invention, the technical means of using parameterized characterization results as coordinate references, mapping the operating status level of core reference entities to weight factors and constructing multi-dimensional feature points in the disaster recovery association space, calculating the weighted geometric centroid based on the multi-dimensional feature points and weight factors to obtain the centroid parameter reflecting the disaster recovery association balance state, and using the centroid parameter as the global dynamic adjustment value to correct the initial load adaptability score of each healthy business unit and integrate the calculation to generate a comprehensive evaluation result overcomes the technical problems in traditional disaster recovery scheduling where load adaptability evaluation only focuses on the individual load capacity of the unit, ignores the overall disaster recovery system balance, and the evaluation result is one-sided, leading to a lack of global rationality in the selection of the target takeover unit. Thus, the comprehensive evaluation result can simultaneously and accurately reflect the individual load capacity of the unit and the overall disaster recovery system balance, improving the comprehensiveness and reliability of the evaluation result.

[0080] In a preferred embodiment of the present invention, step 3.2 above may include:

[0081] Step 3.21 involves acquiring the multi-dimensional feature points and corresponding operational status weight factors of each generated core reference entity. Specifically, this includes: retrieving relevant data for each generated core reference entity, including multi-dimensional feature point data and bound operational status weight factors for the same-city main unit, same-city backup unit, and off-site disaster recovery unit. The multi-dimensional feature points are three-dimensional, containing specific values ​​in three coordinate dimensions: data synchronization delay (X-axis), network transmission bandwidth utilization (Y-axis), and service switching time cost (Z-axis). The operational status weight factors are mapped based on the operational status levels of each core reference entity: 1.2 for good, 1.0 for average, and 0.8 for poor. The acquired data is verified one by one to ensure the completeness of the three coordinate dimensions of the multi-dimensional feature points for each core reference entity and that the operational status weight factors match the corresponding operational status levels correctly, avoiding data loss or mismatches. After successful verification, a one-to-one correspondence list of core reference entity identifiers, multi-dimensional feature points, and weight factors is established.

[0082] Step 3.22: Based on the coordinate values ​​and associated weight factors of each acquired multi-dimensional feature point, calculate the weighted coordinate components of each feature point in each coordinate dimension of the disaster recovery association space. Specifically, this includes: first, extracting the coordinate values ​​of the multi-dimensional feature points of a single core reference entity and the associated running status weight factors from the corresponding list, obtaining the data indicators corresponding to the three coordinate dimensions of the multi-dimensional feature point, namely, the X-axis for data synchronization latency, the Y-axis for network transmission bandwidth utilization, and the Z-axis for service switching time cost; then, calculating the weighted coordinate components one by one according to the coordinate dimensions, with the calculation logic being for each coordinate dimension... The specific coordinate values ​​are multiplied by the weight factor of the operating status corresponding to the core reference entity. For example, if a core reference entity is at a good level with a weight factor of 1.2, its X-axis coordinate value is 30, its Y-axis coordinate value is 50, and its Z-axis coordinate value is 40. Then, the weighted coordinate component of the X-axis is calculated as 30 × 1.2 = 36, the weighted coordinate component of the Y-axis is 50 × 1.2 = 60, and the weighted coordinate component of the Z-axis is 40 × 1.2 = 48. The weighted coordinate components of all coordinate dimensions of the three core reference entities—the same-city main unit, the same-city backup unit, and the off-site disaster recovery unit—are calculated in sequence.

[0083] Step 3.23: Summarize the weighted coordinate components of all core reference entities in each coordinate dimension, and divide by the sum of all weight factors to obtain the final coordinate value of the weighted geometric centroid within the disaster recovery association space. Specifically, this includes: first, classifying and summarizing the weighted coordinate components of the three core reference entities according to their coordinate dimensions; then, summing the total weighted coordinate components for the X-axis data synchronization delay dimension, the Y-axis network transmission bandwidth utilization dimension, and the Z-axis service switching time cost dimension. For example, if the X-axis weighted components of the three core reference entities are 36, 32, and 30 respectively, then the summed X-axis value is 36 + 32 + 30 = 98; and the Y-axis weighted components are 60, 55, and 50 respectively. The sum of values ​​for the Y-axis is 60 + 55 + 50 = 165; the weighted components for the Z-axis are 48, 44, and 42, so the sum of values ​​for the Z-axis is 48 + 44 = 134; simultaneously, the sum of the weight factors for the running status of the three core reference entities is calculated. If the three weight factors are 1.2, 1.0, and 1.0, the sum of weights is 1.2 + 1.0 + 1.0 = 3.2; then, the sum of weighted components for each coordinate dimension is divided by the sum of weight factors to obtain the final coordinate values ​​for each dimension, keeping two decimal places in the calculation. For example, the final coordinate value for the X-axis is 98 ÷ 3.2 = 30.62, the final coordinate value for the Y-axis is 165 ÷ 3.2 = 51.56, and the final coordinate value for the Z-axis is 134 ÷ 3.2 = 41.88.

[0084] Step 3.24 defines the calculated final coordinate values ​​as the centroid parameter, which characterizes the real-time disaster recovery association balance among the three units: the primary unit, the backup unit, and the off-site disaster recovery unit. Specifically, this includes combining the calculated final coordinate values ​​of the X, Y, and Z axes to form a complete set of three-dimensional coordinate data. This three-dimensional coordinate data is the final coordinate value of the weighted geometric centroid within the disaster recovery association space. The meaning of the final coordinate values ​​is then defined as the centroid parameter, which is specifically used to accurately reflect the real-time disaster recovery association balance among the three units: the primary unit, the backup unit, and the off-site disaster recovery unit.

[0085] In this embodiment of the invention, the technical means of first obtaining the multi-dimensional feature points and corresponding operating status weight factors of each core reference entity, then calculating the weighted coordinate components of each coordinate dimension based on the feature point coordinate values ​​and weight factors, then summing the weighted coordinate components of each dimension and dividing them by the sum of weight factors to obtain the final coordinate value of the weighted geometric centroid, and finally defining the coordinate value as the centroid parameter, overcomes the technical problems of the inability to accurately quantify and represent the disaster recovery association balance state among core reference entities in traditional disaster recovery scheduling, and the lack of intuitive and scientific balance measurement indicators. Thus, it achieves the ability to accurately and objectively quantify and reflect the real-time disaster recovery association balance state among the three entities: the same-city main unit, the same-city backup unit, and the off-site disaster recovery unit.

[0086] In a preferred embodiment of the present invention, step 4 above may include:

[0087] Step 4.1: Based on the comprehensive assessment results and a preset takeover capability threshold, the obtained comprehensive assessment results are screened, and healthy business units with assessment values ​​below the threshold are removed, resulting in a preliminary candidate unit set. This set includes a targeted, fully generated list containing all healthy business unit identifiers and their corresponding comprehensive assessment results. The comprehensive assessment results in this list simultaneously reflect the individual load capacity of each unit and the overall disaster recovery system's balance. Subsequently, data preprocessing is performed, verifying the comprehensive assessment results of each healthy business unit in the list to confirm whether the results are complete and without omissions, and whether they fall within a reasonable range of 0 to 100 points. If a unit is found to have missing assessment results or values ​​exceeding the reasonable range, that unit is immediately marked as invalid and removed from the list to avoid invalid data interfering with the screening process. Results: Based on the high reliability requirements of a financial-grade PaaS platform for core financial business operations, a takeover capability threshold of 70 points was preset. This threshold was determined by comprehensively referencing the lowest comprehensive evaluation scores of takeover units in over 200 successful disaster recovery scheduling cases, as well as the security redundancy standards required for core business operations such as payment settlement and credit approval. After data preprocessing, the list of valid healthy business units was traversed, and the comprehensive evaluation result of each unit was compared with the 70-point takeover capability threshold. If the comprehensive evaluation result of a unit was greater than or equal to 70 points, it was determined to have basic takeover capability and was retained. If the evaluation result was lower than 70 points, it was determined to not meet the core business takeover capability requirements and was removed. After the traversal, all retained healthy business units were summarized and organized to form a preliminary candidate unit set.

[0088] Step 4.2: Based on the preliminary candidate unit set, and considering the business type and resource requirements of the faulty main unit, perform a business and resource matching degree analysis on each unit in the preliminary candidate unit set, and sort them according to the matching degree to generate a sorted candidate unit list. Specifically, this includes: retrieving the core information of the faulty main unit obtained from the parsing, determining that the business type of the faulty main unit is personal online banking transaction processing, and extracting the resource requirement parameters of the faulty main unit, specifically including CPU utilization below 70%, memory utilization below 65%, disk I / O utilization below 60%, network bandwidth utilization below 50%, and a transaction processing capacity of no less than 1000 transactions per second; then, starting the matching degree analysis program, conducting quantitative analysis on each unit in the preliminary candidate unit set from two core dimensions: business type matching degree and resource requirement satisfaction degree; in the business type matching degree analysis stage, the scoring standard is set as 100 points for a candidate unit whose supported business type is completely consistent with that of the faulty main unit, and support department The core business functions are scored out of 60 points, while those that do not support core business types receive 0 points. Dimensional scoring is completed by comparing the business type support range of each candidate unit with that of the faulty main unit. In the resource requirement satisfaction analysis, five resource requirement parameters of the faulty main unit are used as evaluation criteria. Each parameter meeting the requirement receives 20 points, and not meeting it receives 0 points, for a total of 100 points. The real-time resource configuration parameters and operational data of each candidate unit are checked one by one to determine whether each resource requirement is met and a score is completed. After completing the two-dimensional scoring, a weighted sum is calculated based on a 60% weighting for business type matching and a 40% weighting for resource requirement satisfaction, yielding the total matching score for each candidate unit. Finally, all candidate units are sorted from highest to lowest total matching score. If multiple candidate units have the same total matching score, the unit with the higher overall evaluation result is ranked first. If the overall evaluation results are still the same, the unit with greater resource redundancy is ranked first, resulting in a final ranked list of candidate units.

[0089] Step 4.3: Based on the sorted candidate unit list, select the candidate unit with the highest matching degree and whose comprehensive evaluation result meets the final condition, and determine it as the target takeover unit. Specifically, the final condition for obtaining the comprehensive evaluation result is that the comprehensive evaluation result is not lower than 80 points. The condition is set to ensure that the target takeover unit not only has a high matching degree with the faulty main unit, but also has excellent individual load-bearing capacity and the ability to adapt to the overall disaster recovery system balance, meeting the high security requirements of financial-grade core business for disaster recovery takeover. Then, start the target unit screening procedure, starting from the first candidate unit in the sorted candidate unit list, and check in turn whether the comprehensive evaluation result of each candidate unit meets the final condition of not lower than 80 points. Conditions: If the highest-ranked candidate unit in the list has a comprehensive evaluation score of 80 or above, the unit is directly designated as the target takeover unit. If the comprehensive evaluation score of a unit is below 80, it is determined that the unit has the highest matching degree but insufficient overall capacity and balance adaptation capability. The unit is skipped and the comprehensive evaluation result of the next candidate unit is checked. The check logic is followed sequentially until the candidate unit with the highest matching degree and the comprehensive evaluation result meets the final conditions is found and designated as the target takeover unit. If, after traversing the entire candidate unit list, no candidate unit with a comprehensive evaluation score of not less than 80 is found, in order to ensure the continuity of financial business, the candidate unit with the highest comprehensive evaluation result in the list is designated as the target takeover unit.

[0090] In this embodiment of the invention, a preliminary candidate unit set is obtained based on comprehensive evaluation results and preset takeover capability thresholds. The business and resource matching degree of the preliminary candidate units is analyzed and ranked in combination with the business type and resource requirements of the faulty main unit. The unit with the highest matching degree and the comprehensive evaluation result that meets the final conditions is selected as the target takeover unit. This technical approach overcomes the technical problems in traditional disaster recovery scheduling, such as the lack of standardized process for target takeover unit selection, single selection dimension, and low business resource matching degree between the selected unit and the faulty main unit, which leads to unstable operation after business switchover. This achieves the technical effects of standardized target takeover unit selection process, selection results that better meet disaster recovery takeover requirements, improved operational stability after business switchover, and ensuring rapid recovery and uninterrupted operation of financial services.

[0091] In a preferred embodiment of the present invention, step 5 above may include:

[0092] Step 5.1: Based on the target takeover unit identifier and combined with the business type parsed from the fault alarm signal and real-time business traffic data, formulate a detailed business traffic switching strategy, including the switching time window, traffic ratio, and routing strategy. Specifically, this includes: first, retrieving the determined target takeover unit identifier; simultaneously, extracting the faulty main unit's business type (personal online banking transaction processing) obtained from parsing the fault alarm signal, and collecting the real-time business traffic data of the faulty main unit before the fault, including core data such as peak transaction processing count of 800 transactions per second, valley count of 100 transactions per second, and average of 500 transactions per second; peak concurrent user connections of 5000 connections per second, valley count of 500 connections per second, and average of 2500 connections per second; and average business request response time of 150 milliseconds. The key takeaways are the three core elements: switching time window traffic ratio and routing strategy. The switching time window is determined based on the peak business patterns of personal online banking transactions, avoiding the high-concurrency period from 9:00 AM to 6:00 PM daily, selecting 0:00 AM to 2:00 AM as the switching time window, with a window duration of 120 minutes, while reserving a 30-minute emergency buffer time. In case of anomalies during the switchover process, a rollback can be completed within the buffer period. The traffic ratio is configured in a phased, progressive manner to avoid a sudden increase in load or traffic congestion on the target takeover unit due to a one-time switch. Specifically, it is divided into four phases: Phase 1 switches 10% of the service traffic for 15 minutes; Phase 2 switches 30% of the service traffic for 30 minutes; Phase 3 switches 50% of the service traffic for 45 minutes; and Phase 4 switches the remaining 10% of the service traffic for 15 minutes. Regarding the routing strategy, the dedicated communication link identifier corresponding to the designated target takeover unit is LT-2024001, and two redundant routing links, LT-2024002 and LT-2024003, are configured. When the main link network transmission bandwidth utilization exceeds 80% or an interruption occurs, automatic switching to the redundant link ensures the stability of traffic transmission. In addition, the routing strategy includes traffic filtering rules, allowing only personal online banking transaction-related traffic corresponding to the faulty main unit to pass through, filtering out irrelevant traffic interference. Finally, all the above content is integrated to form a complete and detailed service traffic switchover strategy.

[0093] Step 5.2: Based on the detailed service traffic switching strategy, generate a dispatchable scheduling instruction containing specific execution parameters and execution sequence. Specifically, this includes: based on the established detailed service traffic switching strategy, starting the scheduling instruction generation program to obtain the core components of the instruction, including specific execution parameters and execution sequence; the generation of specific execution parameters requires converting the content of the switching strategy into quantifiable parameters recognizable by the traffic scheduling component, where the switching time window parameter is obtained as a start time of 00:00:00 and an end time of 02:00. 0:00 buffer time 02:00:00 to 02:30:00; the phased traffic proportion parameters correspond to: Phase 1 00:00:00 to 00:15:00 traffic share 10%; Phase 2 00:15:01 to 00:45:00 traffic share 30%; Phase 3 00:45:01 to 01:30:00 traffic share 50%; Phase 4 01:30:01 to 02:00:00 traffic share 10%; the routing parameters yield the primary link identifier LT-2024001 and the redundant link identifier LT- The switching threshold for links 2024002 and LT-2024003 is 80% bandwidth utilization. Additionally, auxiliary execution parameters are configured, including a 30-second traffic transmission timeout, 3 retries, and a 5-second data verification interval. The execution sequence must follow a logical progression principle to ensure an orderly and controllable switching process. The specific execution sequence is as follows: Step 1: 00:00: Trigger route initialization to complete connectivity checks on the main and redundant links; Step 2: 00:00: Load traffic filtering rules to filter traffic related to personal online banking transactions; Step 3: Execute routing and forwarding configurations for each proportion of traffic according to the phased time nodes; Step 4: After each phase of traffic switching is completed, perform data verification to confirm normal traffic transmission; Step 5: After the full traffic switching is completed at 02:00:00, perform link status stability monitoring; Step 6: From 02:00:00 to 02:30:00, enter the buffer monitoring phase to continuously monitor the business operation status. Execution parameters and sequences are integrated and encapsulated according to the platform's unified instruction format specifications to generate standardized scheduling instructions that can be directly issued.

[0094] Step 5.3 involves sending the scheduling command to the platform's traffic scheduling component for execution, switching the business traffic originally directed to the faulty primary unit to the target takeover unit, thus completing disaster recovery scheduling. Specifically, this includes: first, verifying the completeness and validity of the generated standardized scheduling command, checking whether the execution parameters in the command are complete, whether the execution sequence is logically coherent, and whether the command format meets the interface requirements of the traffic scheduling component. After verification, the scheduling command is sent to the traffic scheduling component through the platform's internal encrypted transmission channel; after the command is sent, continuous communication is established with the traffic scheduling component to obtain real-time command execution status feedback data, including the completion time of each execution step, the actual proportion of traffic switching at each stage, the operating status of the primary and backup links (such as bandwidth utilization, transmission delay, packet loss rate, etc.), and real-time operating data of business traffic (such as transaction response time, success rate, etc.). If monitoring detects that after a certain stage of traffic switching, the business request response time of the target takeover unit exceeds 2... If the success rate drops below 99.9% or within 0.00 milliseconds, a dynamic adjustment mechanism is immediately triggered. The traffic switching ratio for this stage is adjusted through the traffic scheduling component, with each adjustment not exceeding 5%, and the duration of this stage is extended by 10 minutes. If the main link bandwidth utilization reaches the 80% switching threshold, redundant link switching is automatically triggered to ensure uninterrupted traffic transmission. After the scheduling command is executed to the last stage and the full traffic switching is completed, a 30-minute buffer period is continuously monitored. If no abnormal data is found during this period, it is confirmed that the personal online banking transaction traffic originally pointing to the faulty main unit has been smoothly switched to the target takeover unit and that the business is operating normally. Finally, the complete process data of this disaster recovery scheduling is recorded, including command execution logs, traffic switching data, and business operation monitoring data, and archived in the disaster recovery scheduling historical database. At the same time, a disaster recovery scheduling completion notification is sent to the platform operation and maintenance management module to complete the entire disaster recovery scheduling process.

[0095] In this embodiment of the invention, a detailed service traffic switching strategy, including a switching time window, traffic ratio, and routing strategy, is formulated based on the target takeover unit identifier, fault service type, and real-time service traffic data. This strategy generates a dispatchable scheduling instruction containing specific execution parameters and execution sequence, which is then sent to the traffic scheduling component for execution. This overcomes the technical problems of traditional disaster recovery scheduling, such as the coarse service traffic switching strategy, the lack of standardized execution parameters for scheduling instructions, and the easy occurrence of traffic loss or congestion during the switching process, which can lead to service interruption. As a result, the service traffic switching process is smooth and controllable, the accuracy and timeliness of traffic switching are significantly improved, service interruption is avoided, and the technical effect of ensuring the continuity and stability of the core business of the financial-grade PaaS platform is achieved.

[0096] like Figure 2 As shown, embodiments of the present invention also provide a disaster recovery scheduling system for a financial-grade PaaS platform based on a modular architecture, comprising:

[0097] The acquisition module is used to receive fault alarm signals from the main unit; based on the fault alarm signals, it collects and aggregates real-time resource load status data and real-time business traffic data of all business units within the platform.

[0098] The analysis module is used to perform load adaptability analysis on all healthy business units based on real-time resource load status data and real-time business traffic data. It selects the same-city main unit, same-city backup unit and off-site disaster recovery unit in the platform as three core reference entities, analyzes the disaster recovery correlation between the same-city main unit, same-city backup unit and off-site disaster recovery unit and performs parameterized characterization to obtain the parameterized characterization results and the corresponding operating status level.

[0099] The evaluation module is used to convert the results of parameterized characterization and the corresponding operating status level into a feature point set. By calculating the feature point set, the center of gravity parameter reflecting the disaster recovery correlation balance status is obtained, and the center of gravity parameter is used as a dynamic adjustment value. The load adaptability analysis results of each health business unit are corrected by the dynamic adjustment value to obtain a comprehensive evaluation result.

[0100] The screening module is used to dynamically select target takeover units from all healthy business units based on comprehensive evaluation results;

[0101] The processing module is used to formulate and execute scheduling instructions for switching service traffic to the target takeover unit based on the service type corresponding to the fault alarm signal and real-time service traffic data, so as to complete disaster recovery scheduling.

[0102] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A disaster recovery scheduling method for a financial-grade PaaS platform based on a modular architecture, characterized in that, The method includes: Step 1: Receive the fault alarm signal from the main unit; based on the fault alarm signal, collect and aggregate real-time resource load status data and real-time business traffic data of all business units within the platform; Step 2: Based on real-time resource load status data and real-time business traffic data, perform load adaptability analysis on all current healthy business units, select the same-city main unit, same-city backup unit and off-site disaster recovery unit in the platform as three core reference entities, analyze the disaster recovery correlation between the same-city main unit, same-city backup unit and off-site disaster recovery unit and perform parameterized characterization to obtain the parameterized characterization results and the corresponding operating status level. Step 3: The results of parameterized characterization and the corresponding operating status level are converted into a feature point set. By calculating the feature point set, the center of gravity parameter reflecting the disaster recovery correlation balance status is obtained. The center of gravity parameter is used as a dynamic adjustment value. The load adaptability analysis results of each health business unit are corrected by the dynamic adjustment value to obtain a comprehensive evaluation result. Step 4: Based on the comprehensive evaluation results, dynamically select the target takeover unit from all healthy business units; Step 5: Based on the target takeover unit, formulate and execute scheduling instructions to switch service traffic to the target takeover unit according to the service type corresponding to the fault alarm signal and real-time service traffic data, so as to complete the disaster recovery scheduling.

2. The disaster recovery scheduling method for a financial-grade PaaS platform based on a modular architecture as described in claim 1, characterized in that, Receive fault alarm signals from the main unit; Based on fault alarm signals, real-time resource load status data and real-time business traffic data of all business units within the platform are collected and aggregated, including: Receive fault alarm signals sent by the main unit and parse the fault alarm signals to determine the unique identifier of the service unit that has failed and the corresponding fault level; Based on the fault unit identifier and fault level, all healthy business units within the platform that need to be evaluated are dynamically analyzed and screened to form a target unit set for data collection. The target unit set is determined based on healthy units that have direct business call relationships, data synchronization relationships, and resource pool sharing relationships with the fault master unit. Based on the target unit set, a data collection request is synchronously initiated to each healthy business unit within the set, and real-time resource load status data and real-time business traffic data returned by each unit are received.

3. The disaster recovery scheduling method for a financial-grade PaaS platform based on a modular architecture as described in claim 2, characterized in that, Step 2 above includes: Based on real-time resource load status data and real-time business traffic data, the initial load adaptability score of each healthy business unit in the platform is calculated. Based on the initial load adaptability score and combined with the predefined geographical location attributes of each business unit, three core reference entities are selected from all healthy business units: the unit with the highest score in the same city area is determined as the main unit in the same city, the unit with the second highest score in the same city area is determined as the backup unit in the same city, and the unit with the highest score in a different region is determined as the disaster recovery unit in a different region. Based on the same-city main unit, same-city backup unit and off-site disaster recovery unit, real-time operation and maintenance indicators among the same-city main unit, same-city backup unit and off-site disaster recovery unit are extracted and quantitatively analyzed to form a disaster recovery correlation indicator set including data synchronization delay, network transmission bandwidth utilization and service switching time cost. By standardizing the disaster recovery correlation index set and assigning weight coefficients to each index according to the pre-configured strategy for weighted calculation, a quantitative parameterized representation result is obtained. By combining the obtained parameterized characterization results with the real-time resource load status data of the core reference entities, the current operating status of each core reference entity in the disaster recovery system is evaluated and determined.

4. The disaster recovery scheduling method for a financial-grade PaaS platform based on a modular architecture as described in claim 3, characterized in that, The results of parameterized characterization and the corresponding operational status levels are transformed into a set of feature points. By calculating the set of feature points, the centroid parameter reflecting the disaster recovery correlation equilibrium state is obtained, and the centroid parameter is used as a dynamic adjustment value. By dynamically adjusting values, the load adaptability analysis results of each health service unit are corrected to obtain a comprehensive evaluation result, including: Using the parameterized representation results as the coordinate reference, and mapping the operational status level of each core reference entity to a weight factor, the three core reference entities—the same-city main unit, the same-city backup unit, and the off-site disaster recovery unit—are each represented as a multi-dimensional feature point in the disaster recovery association space. Based on multidimensional feature points, according to their respective weighting factors, the weighted geometric centroid of the multidimensional feature points in the disaster recovery association space is calculated, and the coordinate value of the weighted geometric centroid is defined as the centroid parameter reflecting the current disaster recovery association equilibrium state among core reference entities. The center of gravity parameter is used as a global dynamic adjustment value to correct the initial load adaptability score of each health business unit. Based on the correction logic, the dynamic adjustment value is integrated with the initial load adaptability score of each healthy business unit to generate a comprehensive evaluation result that reflects both the individual load capacity of the unit and the overall disaster recovery system balance.

5. The disaster recovery scheduling method for a financial-grade PaaS platform based on a modular architecture as described in claim 4, characterized in that, Based on multi-dimensional feature points, and according to their respective weighting factors, the weighted geometric centroid of the multi-dimensional feature points in the disaster recovery association space is calculated. The coordinate values ​​of the weighted geometric centroid are defined as centroid parameters reflecting the current disaster recovery association equilibrium state among core reference entities, including: By obtaining the multidimensional feature points and corresponding running state weight factors of each generated core reference entity; Based on the coordinate values ​​of each multidimensional feature point and the associated weighting factors, the weighted coordinate components of each feature point in each coordinate dimension of the disaster recovery association space are calculated respectively. The weighted coordinate components of all core reference entities in each coordinate dimension are summed and divided by the sum of all weight factors to obtain the final coordinate value of the weighted geometric centroid in the disaster recovery association space. The calculated final coordinate value is defined as the centroid parameter, which characterizes the real-time disaster recovery association balance among the three units: the primary unit in the same city, the backup unit in the same city, and the disaster recovery unit in another location.

6. The disaster recovery scheduling method for a financial-grade PaaS platform based on a modular architecture as described in claim 5, characterized in that, Based on the comprehensive evaluation results, target takeover units are dynamically selected from all healthy business units, including: Based on the comprehensive evaluation results and a preset takeover capability threshold, the obtained comprehensive evaluation results are screened, and healthy business units with evaluation values ​​below the threshold are removed to obtain a preliminary candidate unit set. Based on the preliminary candidate unit set, and considering the business type and resource requirements of the faulty main unit, the business and resource matching degree of each unit in the preliminary candidate unit set is analyzed, and the units are sorted according to the matching degree to generate a sorted candidate unit list. Based on the sorted candidate unit list, the candidate unit with the highest matching degree and whose comprehensive evaluation results meet the final conditions is selected as the target takeover unit.

7. The disaster recovery scheduling method for a financial-grade PaaS platform based on a modular architecture as described in claim 6, characterized in that, Based on the target takeover unit, and according to the service type corresponding to the fault alarm signal and real-time service traffic data, a scheduling instruction for switching service traffic to the target takeover unit is formulated and executed to complete disaster recovery scheduling, including: Based on the target takeover unit identifier, and combined with the service type and real-time service traffic data parsed from the fault alarm signal, a detailed service traffic switching strategy is formulated, including the switching time window, traffic ratio and routing strategy. Based on the detailed business traffic switching strategy, generate a dispatchable instruction containing specific execution parameters and execution sequence; The scheduling command is sent to the platform's traffic scheduling component for execution, so as to switch the business traffic originally directed to the faulty main unit to the target takeover unit and complete the disaster recovery scheduling.

8. A disaster recovery and scheduling system for a financial-grade PaaS platform based on a modular architecture, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The acquisition module is used to receive fault alarm signals from the main unit; Based on fault alarm signals, collect and aggregate real-time resource load status data and real-time business traffic data of all business units within the platform; The analysis module is used to perform load adaptability analysis on all healthy business units based on real-time resource load status data and real-time business traffic data. It selects the same-city main unit, same-city backup unit and off-site disaster recovery unit in the platform as three core reference entities, analyzes the disaster recovery correlation between the same-city main unit, same-city backup unit and off-site disaster recovery unit and performs parameterized characterization to obtain the parameterized characterization results and the corresponding operating status level. The evaluation module is used to transform the results of parameterized representation and the corresponding operating status level into a feature point set. By calculating the feature point set, the centroid parameter reflecting the disaster recovery correlation equilibrium state is obtained, and the centroid parameter is used as a dynamic adjustment value. By dynamically adjusting the values, the load adaptability analysis results of each health business unit are corrected to obtain a comprehensive evaluation result; The screening module is used to dynamically select target takeover units from all healthy business units based on comprehensive evaluation results; The processing module is used to formulate and execute scheduling instructions for switching service traffic to the target takeover unit based on the service type corresponding to the fault alarm signal and real-time service traffic data, so as to complete disaster recovery scheduling.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • A statistical method and device of access events

    CN101087465A

  • Multi-data center disaster recovery backup and rapid recovery method and system based on private cloud

    CN120934981A