Intelligent integrated disaster recovery method and disaster recovery system

By building a unified resource pool for x86 and ARM architectures, and combining software-defined storage and network technologies, the problems of resource dispersion and slow switching in traditional disaster recovery technologies in heterogeneous environments are solved, achieving efficient data synchronization and rapid fault recovery, and meeting the requirements of financial-grade business continuity.

CN120803820AInactive Publication Date: 2025-10-17ANHUI FENGRUI COMM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510903912.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional disaster recovery technologies suffer from problems such as resource dispersion, slow system switching, high data synchronization failure rate, lack of data consistency guarantee, and insufficient scientific basis for fault switching when facing heterogeneous hardware environments, cloud-native applications, and domestic substitution.

Method used

An intelligent integrated disaster recovery approach is adopted. By connecting x86 and ARM architecture devices, and utilizing heterogeneous device driver libraries and software-defined storage and networking technologies, a unified resource pool is built to realize resource virtualization and load computing. Combined with the sliding window weighted average method and priority scoring function, disaster recovery switching is performed. Sandbox exercises and resource performance monitoring are introduced to support unified management and dynamic scheduling of heterogeneous resources.

Benefits of technology

It improves equipment compatibility and the accuracy and flexibility of resource scheduling, enhances the accuracy and stability of fault recovery, and achieves efficient data synchronization and rapid fault switching across platforms, meeting financial-grade business continuity requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803820A_ABST
    Figure CN120803820A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent integrated disaster recovery method and system, and the method comprises the steps: accessing a device of an X86 architecture and a device of an ARM architecture, and calling a control interface corresponding to each device in a preset heterogeneous device driver library; accessing a computing resource Computeri, a storage resource Storagei and a network resource Networki of each piece of equipment i through the control interface, and constructing a heterogeneous resource pool (ResourcePool); the method comprises the following steps: virtualizing each resource in a resource pool into a standardized resource unit based on a software defined storage and software defined network technology; extracting attribute parameters for each resource unit j to form a resource identification tuple Rj corresponding to the resource unit j, and constructing a uniform resource abstract view table VRDL = [R1, R2, R3,..., Rj,...]; service data synchronization is carried out based on the constructed resources, the service load is monitored for each resource unit j, and the resource load index Wt at the time t is calculated; and when the Wt exceeds a preset threshold value Threshold, triggering a disaster recovery prompt and executing a disaster recovery switching operation by using the VRDL.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of information technology, in particular to an intelligent integrated disaster recovery method and a disaster recovery system. BACKGROUND

[0002] In the current information system highly dependent, disaster recovery system as the core infrastructure to ensure business continuity, its technical evolution has always been accompanied by the diversification of computing architecture, exponential growth of data size and increasing policy compliance requirements. Traditional disaster recovery solutions rely on storage replication, virtualization snapshot and other methods, but when facing heterogeneous hardware environment, cloud native applications and the wave of domestic substitution, they gradually expose systematic problems such as resource dispersion, system switching delay, data synchronization failure rate, etc.

[0003] Chinese invention patent CN107066319A discloses a multi-dimensional scheduling system for heterogeneous resources, which proposes to realize the allocation and scheduling of computing resources through a two-level scheduling architecture, and has certain resource integration capability. However, this solution is mainly used for general computing resource management, lacks the full-stack virtualization integration capability of computing, storage and network in disaster recovery scenarios, and does not consider the key requirements of data consistency, fault switching and sandbox drilling in disaster recovery. SUMMARY

[0004] The present application provides an intelligent integrated disaster recovery method and a disaster recovery system to solve the technical problem that the prior art does not consider the key requirements of data consistency, fault switching and sandbox drilling in disaster recovery.

[0005] To solve the above technical problems, the present application provides the following technical solutions:

[0006] In a first aspect, the present application provides an intelligent integrated disaster recovery method, comprising:

[0007] Accessing devices of X86 architecture and devices of ARM architecture, and calling control interfaces corresponding to each of the devices in a preset heterogeneous device driver library;

[0008] Accessing computing resources Computer i , storage resources Storage i and network resources Network i of each of the devices i through the control interfaces to build a heterogeneous resource pool ResourcePool:

[0009]

[0010] Based on software-defined storage and software-defined network technology, each resource in the resource pool is virtualized into a standardized resource unit;

[0011] Extracting attribute parameters for each resource unit j to form a resource identification tuple R corresponding to the resource unit j j , and constructing a unified resource abstraction view table VRDL = [R1, R2, R3,..., R j ]; wherein the attribute parameters at least include central processing unit (CPU) resource, input / output per second (IOPS) and latency (Latency);

[0012] Based on the constructed resource, business data synchronization is carried out, and for each resource unit j, business load is monitored and resource load indicator W t at time t is calculated:

[0013] W t = α · CPU util + β · IOPS + γ · Latency

[0014] Wherein, CPU util represents CPU usage, and α, β and γ are weight coefficients;

[0015] In the case where W t exceeds a preset threshold Threshold, a disaster recovery prompt is triggered and a disaster recovery switching operation is performed using the VRDL.

[0016] In an optional embodiment, the control interface corresponding to each device in the preset heterogeneous device driver library is called, including:

[0017] For the device of the X86 architecture or the device of the ARM architecture, BIOS information and network card MAC address of the device are parsed;

[0018] Using the BIOS information and the network card MAC address, a device compatibility entry is matched, and according to the matching result, the control interface corresponding to the device in the heterogeneous device driver library is called.

[0019] In an optional embodiment, based on software-defined storage and software-defined network technology, each resource in the resource pool is virtualized into a standardized resource unit, including:

[0020] The computing resource is mapped to a virtual CPU unit and registered as a computing resource unit through a standard interface;

[0021] The storage resource is divided into a logical volume block with a label and registered as a storage resource unit through a standard interface;

[0022] A virtual network interface is created for the network resource, and bandwidth limit, forwarding rule and QoS parameter are configured, and the network resource unit is encapsulated through a standard interface.

[0023] In an alternative embodiment, the method further comprises:

[0024] The weight coefficients a, b, g are dynamically updated according to historical monitoring data of the service load by using a sliding window weighted average method.

[0025] In an alternative embodiment, the performing disaster recovery switchover operation by using the VRDL comprises:

[0026] The priority score of the resource to be scheduled in the VRDL is calculated by using the following formula:

[0027] P = d * QoS score + m * SLA weight + h * Latency

[0028] Wherein, d, m, h are weight coefficients, QoS score represents the service quality of the resource, SLA weight represents the importance of the resource.

[0029] In an alternative embodiment, the performing disaster recovery switchover operation by using the VRDL comprises:

[0030] Determining at least one candidate resource in the VRDL for performing disaster recovery switchover operation;

[0031] Isolating the resource replica to build a sandbox environment, performing a fault script, and sandbox pre-rehearsing the candidate resource; the fault script comprises: network interruption simulation, storage input and output congestion simulation, and / or system call blocking simulation;

[0032] Determining the resource for performing disaster recovery switchover operation based on the results of the sandbox pre-rehearsing.

[0033] In an alternative embodiment, the method further comprises:

[0034] Benchmarking the performance of each resource unit j and recording the initial performance parameters of the resource unit j;

[0035] Performing performance check on the resource unit j according to a predetermined period, the performance check comprising: whether the hardware is faulty, whether the performance is degraded, whether the network is delayed, and / or whether the resource consumption is abnormal;

[0036] In the case where the performance check result is performance abnormality, triggering a disaster recovery prompt and performing disaster recovery switchover operation by using the VRDL.

[0037] In an alternative embodiment, the method further comprises:

[0038] For the resource unit j whose performance check result is performance anomaly, determining an anomaly cause of the resource unit j;

[0039] For the determined anomaly cause, taking a preset corresponding repair measure to perform a fault recovery operation; the repair measure includes restarting a device, replacing a network channel, scheduling a backup resource, and / or starting manual intervention.

[0040] In an optional embodiment, the method further includes:

[0041] Using a time series analysis algorithm to construct a resource demand prediction model, a sample of the model being load data of each resource unit j in different time periods;

[0042] Using the resource demand prediction model to predict resource demand, and dynamically adjusting allocation of each resource unit in the resource pool according to a prediction result.

[0043] In a second aspect, the application provides an intelligent integrated disaster recovery system, characterized in that it includes:

[0044] A device access module, a resource pool construction module, a resource virtualization module, a view construction module, a load calculation module, and a resource scheduling module;

[0045] The device access module is configured to access devices of X86 architecture and devices of ARM architecture, and to call control interfaces corresponding to each device in a preset heterogeneous device driver library;

[0046] The resource pool construction module is configured to access computing resources Computer i , storage resources Storage i , and network resources Network i of each device i through the control interfaces, and to construct a heterogeneous resource pool ResourcePool:

[0047]

[0048] The resource virtualization module is configured to virtualize each resource in the resource pool into a standardized resource unit based on software-defined storage and software-defined network technology;

[0049] The view construction module is configured to extract attribute parameters for each resource unit j to form a resource identification tuple R j corresponding to the resource unit j, and to construct a unified resource abstraction view table VRDL=[R1,R2,R3,...,R j ,...]; wherein the attribute parameters include at least central processing unit CPU resources, read-write operations per second IOPS, and latency Latency.

[0050] The load calculation module is used for carrying out business data synchronization based on the constructed resources, and monitoring business load and calculating resource load index W of each resource unit j at time t t :

[0051] W t = alpha CPU util + beta IOPS + gamma Latency, wherein CPU util represents CPU usage, alpha, beta and gamma are weight coefficients;

[0052] The resource scheduling module is used for triggering disaster recovery prompt and performing disaster recovery switching operation by using the VRDL when W t exceeds a preset threshold Threshold.

[0053] The technical scheme provided by the application has at least the following beneficial effects:

[0054] 1. The application supports unified access and driving adaptation of heterogeneous devices, realizes unified control interface loading of X86 and ARM architecture devices by calling preset heterogeneous device driving library and combining BIOS and MAC information for compatibility matching, has stronger device compatibility and localization adaptation capability compared with the resource scheduling mode supporting only general computing architecture in the prior art, and meets the disaster recovery demand in the signal creation environment.

[0055] 2. The application realizes deep virtualization of computing, storage and network resources by using software-defined storage and software-defined network technology, encapsulates resources into independent computing resource units, storage resource units and network resource units through a standard interface, is beneficial to unified management and scheduling, and does not specifically involve the bottom virtualization implementation details in the prior art, so it is difficult to directly support horizontal expansion of multi-dimensional resource pool.

[0056] 3. The application proposes a structured resource abstract view table VRDL modeling mechanism, forms a tuple structure by extracting multi-dimensional attributes (such as CPU resource, IOPS, Latency, etc.) of resource units, and uses the tuple structure as the basis for subsequent scheduling and evaluation, which improves the decision accuracy and flexibility of resource scheduling compared with the prior art in which resource state information is scattered and the scheduling logic does not have the advantage of modeling.

[0057] 4. The application introduces a sliding window weight adjustment mechanism and a priority scoring function, so that the resource scheduling not only depends on the real-time load index W t , but also combines QoS score, SLA weight and response delay for multi-factor priority evaluation, significantly enhances the controllability and service quality guarantee capability of resource scheduling, and the prior art does not consider factors such as QoS and service level.

[0058] 5、The application innovatively adds a disaster recovery drill mechanism in a sandbox environment, executes a preset fault script (such as network interruption, IO congestion, etc.) through an isolated copy, and performs elastic verification on candidate resources outside the real environment, compared with a scheduling model in the prior art that lacks verification ability before recovery, the application effectively improves the accuracy and stability of fault recovery operations;

[0059] 6、The application integrates resource performance baseline monitoring and fault recovery processes, can intervene in advance and take a preset repair strategy (such as restart, backup switching, etc.) when resource performance degrades, forms a complete early warning, judgment, and recovery closed loop, and has stronger practicability and engineering deployment feasibility;

[0060] 7、The application integrates a resource demand prediction mechanism based on time series, supports dynamic modeling and pre-expansion of resource load trends, enhances the forward-looking and elastic capacity management capabilities of the system, and the scheduling logic in the prior art does not have a forward prediction capability and has poor adaptability. BRIEF DESCRIPTION OF DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0062] Figure 1 is a flowchart of an intelligent integrated disaster recovery method provided by an embodiment of the application.

[0063] Figure 2 is a module schematic diagram of an intelligent integrated disaster recovery system provided by an embodiment of the application. DETAILED DESCRIPTION

[0064] The application will be described in detail below with reference to the drawings and specific embodiments. The following embodiments are implemented on the premise of the technical solutions of the application, and detailed implementation modes and specific operation processes are given, but the protection scope of the application is not limited to the following embodiments.

[0065] In the application, the words such as “in a possible embodiment”, “exemplary” or “for example” are used to represent an example, illustration or description. Any embodiment or design scheme described as “in a possible embodiment”, “exemplary” or “for example” in the application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words such as “in a possible embodiment”, “exemplary” or “for example” are intended to present the relevant concept in a specific way.

[0066] Disaster recovery technology, as the core means to ensure the continuity of information system business, has always been closely related to the development of computing architecture, data size and policy requirements. Financial, power and other industries have been refining the recovery time objective (RTO) and recovery point objective (RPO) quantitative indicators year by year, marking the entry into the deep regulatory stage of parallel promotion of regulations combined with cases. At the same time, the rapid development of digital economy makes business continuity necessary - the downtime of financial systems will cause huge losses, the interruption of medical image archiving and communication systems (PACS) will affect medical procedures, and systems with a hybrid cloud architecture penetration rate of 68% may still face the challenge of cross-cloud disaster recovery coverage of less than 30%.

[0067] At the technical hierarchical architecture level, the traditional disaster recovery technology system can be divided into four types of technology paths: storage layer, virtualization layer, database layer, and application layer.

[0068] Traditional storage layer disaster recovery technology is based on block-level data synchronization mechanisms of storage devices (such as EMC SRDF, NetApp SnapMirror), which implement cross-data center data replication through storage area networks (SAN) and network-attached storage (NAS) protocols. In synchronous mode, RPO can be achieved at the second level, and in asynchronous mode, RPO can be extended to the hour level. However, research shows that the compatibility of this technology for heterogeneous storage devices is less than 30%, and it cannot perceive the upper-layer business logic state. Test data from a commercial bank's disaster recovery system shows that storage layer replication caused up to 27% of database transactions to be lost, exposing the lack of coordination between storage and database layers;

[0069] Traditional virtualization layer disaster recovery technology is represented by VMware vSphere Replication and Hyper-V Replica, which is a virtual machine incremental snapshot technology. It implements disaster recovery through Hypervisor-level data capture, with a typical RPO of 5 minutes. However, cross-virtualization platform migration tests show that the hot migration failure rate from VMware to KVM environment is as high as 65.3%, and the coverage rate for containerized applications is less than 15%. In a thousand-node virtual machine cluster scenario, the full replication time and bandwidth occupancy grow exponentially. Actual measurements show that it takes more than 72 hours to fully replicate 100 virtual machines, with a link bandwidth occupancy rate of 95%;

[0070] Traditional database layer disaster recovery technology relies on log parsing mechanisms such as Oracle DataGuard and MySQL master-slave replication to achieve transaction-level data synchronization through Redo Log / Binlog replay. This technology can theoretically achieve sub-second RPO, but is limited by database type heterogeneity. Data from a provincial government cloud project shows that the disaster recovery transformation from Oracle to domestic Gauss database requires the development of additional data conversion middleware, with a manual input of 1800 man-days. In addition, in large transaction processing scenarios (single transaction over 1GB), synchronization delay increases by more than 300%, indicating that traditional log parsing algorithms are sensitive to transaction size.

[0071] Traditional application layer disaster recovery technology is based on the global load balancing scheme of Domain Name System (DNS) and Global Server Load Balancing (GSLB), which achieves business traffic redirection through health checks and IP switching. The typical RTO is on the order of 15 minutes, but lacks a mechanism to ensure data consistency. A disaster recovery drill case of an e-commerce platform shows that 35% of user session states are lost after switching, causing abnormal order payment links.

[0072] However, under the multiple drivers of cloud computing, hybrid architecture, and alternative to Xintian, traditional disaster recovery technology system faces the following common challenges in data synchronization, fault switching, and drill verification, and gradually exposes systematic defects such as architectural rigidity, low coordination efficiency, and insufficient compatibility.

[0073] Data synchronization performance bottleneck: There is a compatibility gap between traditional protocol stacks (such as iSCSI / FC / NFS) and new storage protocols (such as NVMe over Fabrics). Experimental data shows that in a 100Gb network environment, the throughput of NVMe-oF protocol is 83% higher than that of iSCSI, but existing disaster recovery tools are generally not compatible with new protocols. At the same time, the incremental capture mechanism based on file system monitoring produces more than 90% CPU occupancy in the scenario of millions of small files, forcing some cloud service providers to downgrade to full backup strategy.

[0074] Fault switching reliability defects: Manual intensive switching process has significant operational risks. In a disaster recovery drill of a securities company, 2.7TB of data was damaged due to incorrect storage disconnection sequence. The deeper problem is the lack of cross-system state consistency guarantee mechanism. A case of a financial institution shows that the switching time difference between Oracle database and Redis cache is 47 minutes, causing end-of-day reconciliation failure.

[0075] Insufficient scientificity of drill verification: About 90% of traditional solutions rely on direct testing in production environment, and a city commercial bank mistakenly triggered a real switch during a drill, causing a core system outage for 11 hours. The roughness of the evaluation method further exacerbates the risk, and in a case where the disaster recovery drill was marked as successful, subsequent audits found that 27% of the order tables had primary key conflicts.

[0076] Technology evolution trend and contradiction: From the perspective of technology stratigraphy, disaster recovery technology has gone through four stages: tape backup (1990s), storage replication (2000s), virtualization disaster recovery (2010s), and cloud native solution (2020s). Although the technical indicators are continuously optimized, there is always the following structural contradiction: the contradiction between vertical closure and horizontal expansion: storage, virtualization, and database technologies are self-contained, and cross-layer collaboration requires customized development (the cost of certain project integration accounts for 42% of the total investment); the contradiction between stable architecture and cloud native: dynamic load such as containers and serverless has an essential conflict with static disaster recovery strategy, and the time consumption of K8s fault recovery is 3 times that of virtual machines; the contradiction between cost control and performance target: to meet the financial-level requirement of RPO<60 seconds, the storage procurement cost of a certain bank's disaster recovery system has increased from 15% to 62%.

[0077] In summary, the traditional disaster recovery technology system has a significant capability fault in the dimensions of hybrid architecture adaptation, intelligent operation and maintenance, and cross-cloud collaboration. Under the dual pressure of policy supervision and business sensitivity, the industry urgently needs to break through technical bottlenecks such as multi-source heterogeneous integration, intelligent switching, and cross-cloud compatibility, and build a new generation of disaster recovery system and product technology that takes into account compliance, economy, and agility, with features such as heterogeneous resource management, semantic-level data synchronization, and automated policy orchestration.

[0078] Figure 1 is a flowchart of an intelligent integrated disaster recovery method provided by an embodiment of the present application. As shown in Figure 1 , an intelligent integrated disaster recovery method can include:

[0079] S101, access X86 architecture devices and ARM architecture devices, and call the control interface corresponding to each device in the preset heterogeneous device driver library;

[0080] S102, access the computing resource Computer i , storage resource Storage i , and network resource Network i of each device i through the control interface, and build a heterogeneous resource pool ResourcePool:

[0081]

[0082] S103, based on software-defined storage and software-defined network technology, virtualize each resource in the resource pool into a standardized resource unit;

[0083] S104, extract attribute parameters for each resource unit j to form a resource identification tuple R jand construct a uniform resource abstract view table VRDL=[R1, R2, R3,..., R j ,...]; wherein the attribute parameters at least include central processing unit (CPU) resources, read and write operations per second (IOPS) and latency (Latency);

[0084] S105, based on the constructed resource, carrying out business data synchronization, and monitoring business load and calculating resource load index W t :

[0085] W t =α·CPU util +β·IOPS+γ·Latency

[0086] wherein CPU util represents CPU usage, and alpha, beta and gamma are weight coefficients;

[0087] S106, in the case where W t exceeds a preset threshold (Threshold), triggering a disaster recovery prompt and performing disaster recovery switching operation by using the VRDL.

[0088] The heterogeneous resource pool in the application can refer to unified management and pool management of computing resources (CPU), storage devices (SSD, SAS, SAN, etc.) and network devices (10G / 25G switch, virtual switch, etc.) of different types, different manufacturers and different architectures (such as X86, ARM, domestic Feiteng, etc.), that is, abstracting them into the same kind of logical resources, which provides a technical basis for cross-platform disaster recovery (such as cloud+local, Oracle+Dream).

[0089] In a possible embodiment, the calling of the control interface corresponding to each device in the preset heterogeneous device driver library can include: for the device of the X86 architecture or the device of the ARM architecture, parsing BIOS information and network card MAC address of the device; matching a device compatibility table item by using the BIOS information and the network card MAC address, and calling the control interface corresponding to the device in the heterogeneous device driver library according to a matching result.

[0090] In the application, virtualization technology (such as KVM, VMware, Hyper-V, Docker, etc.) can be used to abstract and uniformly manage physical hardware resources, and originally dispersed devices (such as multiple servers) are integrated into one or more virtual resource pools. After integration, the system can automatically allocate computing / storage / network resources on demand, without physical limitations, and this process is an important process of establishing a heterogeneous resource pool.

[0091] In a possible embodiment, the virtualization of the resources in the resource pool into standardized resource units based on the software-defined storage and software-defined network technologies can include: mapping the computing resources into virtual CPU units and registering as computing resource units through a standard interface; dividing the storage resources into logical volume blocks with labels and registering as storage resource units through a standard interface; creating virtual network interfaces for the network resources and configuring bandwidth limits, forwarding rules and QoS parameters, and encapsulating as network resource units through a standard interface.

[0092] In a possible embodiment, the method can further include: dynamically updating the weight coefficients a, b and g according to historical monitoring data of the service load by using a sliding window weighted average method.

[0093] In the present application, the heterogeneous resources (such as different CPU architectures, storage device types, etc.) after virtualization can be uniformly modeled, labeled and structured to generate a logical view that can be recognized and invoked by a scheduling system, i.e., a unified resource abstract view table VRDL. For example, the VRDL can be generated by a resource orchestration system (such as OpenStack, K8s or a private platform) and can include the following information: node ID, CPU core number, memory, network bandwidth, architecture type (X86 / ARM), scheduling state, resource label (such as “high IO” and “low latency”), etc.

[0094] For example, the virtualization of the storage resources by SDS can include: scanning physical storage devices (such as SSD, HDD, NVMe); dividing the physical disks into logical volume blocks (LV) or object storage blocks (Object) through LVM or Ceph backend systems; labeling each volume block (such as “high IOPS” and “low latency”); encapsulating as a standard interface service and registering to the resource scheduling layer through a unified API (such as CSI).

[0095] For example, the virtualization of the network resources by SDN can include: discovering physical network interfaces and VLAN configurations; creating virtual network bridge interfaces (such as veth and tap) through Open vSwitch or similar controllers; setting bandwidth and QoS rules (such as HTB) and binding with resource units; registering the configuration results as virtual network resource units for use by the scheduling engine.

[0096] The disaster recovery switching operation in the present application can refer to dynamically adjusting the allocation number and range of resources according to the changes of the service load. For example, it can include: scaling up (adding more computing nodes, storage space or network bandwidth for the load); scaling down (releasing idle resources after the load decreases to improve efficiency and reduce energy consumption).

[0097] In the present application, the capacity expansion and contraction action can be an automatic triggering behavior, and does not need to be manually configured by an administrator.

[0098] In a possible embodiment, the performing of the disaster recovery switching operation using the VRDL can include: calculating a priority score of a resource to be scheduled in the VRDL using the following formula:

[0099] P = δ * QoS + μ * SLA + η * Latency score weight

[0100] wherein δ, μ, η are weight coefficients, QoS score represents the service quality of the resource, SLA weight represents the importance of the resource.

[0101] For example, after performing the disaster recovery switching operation, a data synchronization step can be further started, for example: through a protocol adaptive conversion module, syntax tree parsing is performed on a plurality of database Redo Logs, semantic level data conversion is completed; the converted structured data is synchronized to a heterogeneous database environment, including but not limited to Oracle, Dream DM, and Renmin University Gold Warehouse.

[0102] For example, the syntax tree parsing can use a recursive descent parsing algorithm oriented to operator priority and statement block structure, to ensure the conversion integrity of compatible DDL and DML statements.

[0103] In a possible embodiment, the performing of the disaster recovery switching operation using the VRDL includes: determining at least one candidate resource in the VRDL for performing the disaster recovery switching operation; isolating resource replicas to build a sandbox environment, performing a fault script, and sandbox pre-rehearsing the candidate resource; the fault script includes: network interruption simulation, storage input / output congestion simulation, and / or system call blocking simulation; based on a result of the sandbox pre-rehearsing, determining the resource for performing the disaster recovery switching operation.

[0104] For example, an evaluation report can be generated during the sandbox rehearsing process, and the report can include the following dimension indicators: a) data consistency verification result (based on improved Merkle tree for hash comparison); b) service recovery integrity evaluation (verify business continuity through a self-defined process engine); c) fault recovery RTO, RPO indicators.

[0105] ​​In a possible embodiment, the method can further include: benchmarking the performance of each resource unit j, and recording the initial performance parameters of the resource unit j; performing performance inspection on the resource unit j according to a predetermined period, the performance inspection including: whether the hardware is malfunctioning, whether the performance is degraded, whether the network is delayed, and / or whether the resource consumption is abnormal; in the case where the result of the performance inspection is performance abnormality, triggering a disaster recovery prompt and performing a disaster recovery switching operation by using the VRDL.

[0106] In a possible embodiment, the method can further include: determining the abnormal cause of the resource unit j in the case where the result of the performance inspection is performance abnormality; taking a preset corresponding repair measure according to the determined abnormal cause, and performing a fault recovery operation; the repair measure including: restarting the device, replacing the network channel, scheduling the backup resource, and / or starting manual intervention.

[0107] In a possible embodiment, the method can further include: constructing a resource demand prediction model by using a time series analysis algorithm, the sample of the model being the load data of each resource unit j in different time periods; predicting the resource demand by using the resource demand prediction model, and dynamically adjusting the allocation of each resource unit in the resource pool according to the prediction result.

[0108] Figure 2 FIG. 1 is a schematic diagram of a module of an intelligent integrated disaster recovery system provided by an embodiment of the present application. As shown in FIG. 1, an intelligent integrated disaster recovery system 10 can include: Figure 2

[0109] a device access module 101, a resource pool construction module 102, a resource virtualization module 103, a view construction module 104, a load calculation module 105, and a resource scheduling module 106;

[0110] The device access module 101 is configured to access devices of X86 architecture and devices of ARM architecture, and to call control interfaces corresponding to each device in a preset heterogeneous device driver library;

[0111] The resource pool construction module 102 is configured to access the computing resource Computer i , the storage resource Storage i , and the network resource Network i of each device i by using the control interface, and to construct a heterogeneous resource pool ResourcePool

[0112]

[0113] ​The resource virtualization module 103 is configured to virtualize each resource in the resource pool into a standardized resource unit based on software-defined storage and software-defined network technology.

[0114] The view construction module 104 extracts attribute parameters for each resource unit j to form a resource identification tuple R corresponding to the resource unit j j , and constructs a unified resource abstract view table VRDL = [R1, R2, R3,..., R j ,...]; wherein the attribute parameters at least include central processing unit (CPU) resource, input / output per second (IOPS), and latency.

[0115] The load calculation module 105 is configured to perform business data synchronization based on the constructed resource, monitor the business load and calculate a resource load indicator W t at time t for each resource unit j.

[0116] W t = a CPU util + b IOPS + g Latency

[0117] wherein CPU util represents CPU usage, and a, b, and g are weight coefficients.

[0118] The resource scheduling module 106 is configured to trigger a disaster recovery prompt and perform a disaster recovery switching operation using the VRDL when the W t exceeds a preset threshold Threshold.

[0119] It should be noted that the above is only an exemplary description, and in actual application, different needs of users can be set, and the present embodiment does not limit this.

[0120] The intelligent integrated disaster recovery method and system of the present application will be described below.

[0121] 1. Overall architecture

[0122] The core goal of the present application is to build a cross-platform, adaptive, and high-intelligent disaster recovery management system, focusing on solving the systematic problems such as resource fragmentation, inefficient data synchronization, and complex switching process under the hybrid cloud architecture.

[0123] The technical solution of the present application adopts a three-layer two-wing architecture: the basic layer adopts a hyper-converged resource pool to realize integration of heterogeneous hardware; the capability layer adopts a data synchronization bus and an automatic switching engine; the management layer adopts a unified monitoring and policy control platform; the security wing adopts a ransomware isolation area and a drill isolation area; and the national security wing adopts a full-stack adaptation module.

[0124] The architecture realizes the synergy of capabilities of each layer through a dynamic orchestration engine, meeting the high reliability and high compliance requirements of energy, finance, government affairs, medical treatment and other scenarios for disaster recovery systems.

[0125] The following is a specific description:

[0126] (1) Super-converged heterogeneous resource pooling scheduling model

[0127] ① Technical principle: Adopt software-defined technology to build a logical resource pool, breaking through the limitation of physical device heterogeneity. Propose a virtual resource description language VRDL to realize automatic parsing of hardware characteristics and generate a unified resource view.

[0128] ② Function modules: Can include intelligent load balancer, heterogeneous driver library and elastic scaling engine. Among them, the intelligent load balancer dynamically adjusts resource allocation based on Q-Learning algorithm; the heterogeneous driver library pre-integrates more than 100 mainstream device drivers (including domestic Feiteng / Kunpeng chipsets); the elastic scaling engine automatically expands and shrinks capacity according to business pressure (minute-level response).

[0129] ③ Technical features: Based on the super-converged architecture, realize resource virtualization and reorganization of cross-X86 / ARM heterogeneous devices, and propose a dynamic reorganization operator of software-defined resources:

[0130]

[0131] Among them, represents the dynamic weighted aggregation operation of computing storage network resources, and α, β, γ are load balancing coefficients, which can be dynamically calculated by the following formula:

[0132]

[0133] ④ Technical effect: Financial cloud test data shows that this model improves the utilization rate of mixed resources of DELL PowerStore and Huawei OceanStor from 38% to 82%, and shortens the joint debugging period by 68% (from 42 days to 13.5 days).

[0134] (2) Data synchronization bus

[0135] ① Technical principle: Design a protocol adaptive conversion module PACM to support bidirectional synchronization of object storage between AWS S3 / Alibaba Cloud OSS / local Ceph, realize semantic-level synchronization of multi-source data:

[0136]

[0137] Among them, λ is the protocol conversion efficiency factor.

[0138] ② Functional modules: can include a database syntax tree converter (or syntax tree reconstruction module), a file system incremental capturer and a third-party tool scheduler. Among them, the database syntax tree converter supports mutual conversion between Dameng DM_SQL and Oracle PL / SQL; the file system incremental capturer realizes second-level file change tracking; the third-party tool scheduler seamlessly integrates commercial synchronization tools such as ADG and OGG.

[0139] ③ Technical effect: According to actual measurement, the value from Dameng database to Oracle database is 0.89, and the value from Huawei cloud to Tencent cloud is 0.93.

[0140] (3) Automatic disaster recovery switching engine

[0141] ① Technical principle: Build a topology-aware process orchestration model, and abstract the switching operation as an atomic action.

[0142] ② Functional modules: can include a health assessment model, an intelligent routing controller and a rollback guarantee mechanism. Among them, the health assessment model calculates the system health index HI based on more than 200 indicators; the intelligent routing controller supports BGP / OSPF / VxLAN multi-protocol linkage; the rollback guarantee mechanism stores multiple versions of transaction logs (retaining the state of the last 24 hours).

[0143] (4) Dynamic demand perception algorithm

[0144] ① Technical principle: Establish a resource demand prediction model based on a long short-term memory network LSTM neural network:

[0145] h t =σ(W xh x t +W hh h t-1 +b h ),

[0146] Among them, the input layer contains 12-dimensional business indicators such as IOPS and throughput, and the output layer generates the optimal resource configuration scheme.

[0147] ② Technical effect: According to actual measurement, in the government cloud stress test, the resource configuration deviation rate is reduced from 41.7% of the traditional scheme to 4.3%.

[0148] 2. Breakthrough of characteristic technology

[0149] (1) Security enhancement design: A logically isolated anti-ransom storage area is built to prevent malicious tampering and reading of critical business data, ensuring that core backup data cannot be tampered with; a completely independent exercise isolation area is established as a disaster recovery resource facility, which realizes the simulation switching exercise and verification of the business system. Among them, the fault scenario generator adopts an abnormal injection algorithm based on reinforcement learning, which can realize the automatic combination test of more than 100 kinds of fault modes.

[0150] (2) Deep adaptation of China's information creation: The full-stack verification of localization is shown in Table 1:

[0151]

[0152] Table 1

[0153] Through the above disaster recovery management system for hybrid cloud environment, a heterogeneous resource pool scheduling module is built, realizing the unified management of X86 / ARM server computing, storage and network; an intelligent data bus is adopted, with 15 kinds of database parsing engines such as Dream / Oracle / GaussDB built-in; a dual-active topology perception engine is used to monitor more than 100 types of device health indicators in real time; an isolated sandbox exercise system is realized, supporting simulation of more than 100 kinds of fault scenarios.

[0154] The following takes the FR-DR based on the disaster recovery management integrated platform of Fengrui Communication as an example to illustrate another exemplary description of the intelligent integrated disaster recovery method and disaster recovery system of the present application.

[0155] 1. System architecture design of disaster recovery system:

[0156] FR-DR adopts a four-layer progressive architecture, and the core functions of each layer are as follows:

[0157] (1) The intelligent resource scheduling layer can realize the virtualization integration of physical resources based on hyper-converged architecture:

[0158]

[0159] Among them, represents the dynamic reorganization operator of software-defined resources.

[0160] The dynamic load perception algorithm can be used to monitor the business input / output (IO) mode in real time:

[0161] W t = α · CPU util + β · IOPS + γ · Latency

[0162] When W t is greater than the preset threshold, the elastic expansion strategy is triggered. Experimental data shows that this mechanism can improve the resource utilization rate from 38% to 82%.

[0163] (2) Data multi-dimensional protection layer can achieve multi-source heterogeneous data synchronization through the design of data synchronization bus DCB:

[0164] ① Protocol analysis engine: support syntax tree conversion of 15 kinds of database Redo Log, solve the DDL statement compatibility problem of Dream database and Oracle database.

[0165] ② Blockchain storage module: using improved Merkle-Patricia tree structure, realizing real-time hash chain of backup data. In the anti-ransomware test, this mechanism successfully blocked 100% of encryption attacks (ICSA Labs certification report). Among them, real-time hash chain is:

[0166] Hash Block =SHA3(SHA3(Data chunk )||Timestamp)。

[0167] (3) Cloud native adaptation layer can design cluster controller for K8s environment:

[0168] ① POD topology analysis: based on service grid Service Mesh to generate microservice dependency graph.

[0169] ② Cross-cloud migration engine: using CRIU(Checkpoint / Restore in Userspace) technology to realize container hot migration. The actual application can verify that the recovery time target RTO of securities system container service can be shortened from 8 minutes and 12 seconds to 43 seconds. Among them, container hot migration is:

[0170]

[0171] (4) Automatic operation and maintenance layer can build an AI-based operation and maintenance decision system:

[0172] ① Fault prediction model: using LSTM neural network to process historical alarm data:

[0173] h t =σ(W xh x t +W hh h t-1 +b h ),

[0174] In the storage fault prediction, the accuracy of F1-score 0.91 is achieved.

[0175] ② Process orchestration engine: support low-code visual orchestration of more than 200 atomic operations. Through actual application, it can be verified that the customer switching steps can be simplified from 58 to 3.

[0176] 2. Key technology verification:

[0177] (1) Heterogeneous database synchronization test Build a dual-active environment of Oracle 19c database and Dream DM8 database, and the verification results are shown in Table 2 as follows:

[0178] Indicator Traditional OGG scheme FR-DR scheme Improvement amplitude Synchronization latency (ms) 2180 92 23.7 times Transaction integrity 88.7% 99.999% 12.7% CPU occupancy 63% 17% 73%

[0179] Table 2

[0180] (2) Hybrid cloud disaster recovery drill Cross-cloud switching test is carried out in the provincial government cloud (Huawei cloud + local OpenStack), including:

[0181] ① Network topology reconstruction: The software-defined network SDN controller completes VxLAN tunnel reconstruction within 9 seconds.

[0182] ② Data consistency check: PB-level data comparison is realized by using the improved RSYNC algorithm:

[0183]

[0184] The test results show that the data difference rate is less than 10 -9 , which meets the requirements of financial-level disaster recovery.

[0185] Through the above-mentioned fusion super-converged architecture, semantic-level data synchronization and intelligent disaster recovery system of AIOps, a heterogeneous resource pooling scheduling model is proposed, which realizes the utilization rate of resources across X86 / ARM architecture to 82%; The design of data synchronization bus protocol is realized, which solves the problem of data synchronization delay between domestic database and cloud computing platform (for example, the synchronization efficiency of Dream database to Oracle database is improved by 22.5 times); The resource prediction algorithm is constructed, and the 72-hour early warning of storage space is realized.

[0186] In summary, as a new generation of intelligent disaster recovery management system, the present application is committed to solving the systematic defects of traditional disaster recovery system under the background of digital transformation, and realizing cross-platform, full-process and high-availability business continuity guarantee through technical innovation. In view of the current industry core pain points, the system constructs a three-dimensional solution covering resource integration, data synchronization, intelligent operation and maintenance and other dimensions:

[0187] In the aspect of heterogeneous resource integration, the traditional scheme needs to independently purchase computing, storage, network equipment and disaster recovery software, resulting in poor compatibility of cross-vendor equipment, low resource utilization (such as more than 30% of storage redundancy over-provisioning), and more than 60% of the project cycle is consumed in the joint debugging of multi-supplier coordination. The present application abstracts multi-source hardware devices into a unified resource pool through hyper-converged virtualization technology, realizes the dynamic disaster recovery scheduling capability of cross-brand and cross-architecture systems by using software-defined storage (SDS) and software-defined network (SDN), and through actual application, it can be verified that the storage resource utilization can be improved from 42% to 85%, the equipment joint debugging period is shortened by 72%, and the heterogeneous system interface bottleneck is effectively eliminated.

[0188] In the field of hybrid cloud disaster recovery, the existing cloud vendor scheme needs to build a region, resulting in incompatibility between cloud disaster recovery and local system, and single point failure risk (such as cloud platform failure leading to inavailability of disaster recovery end). The present application innovatively designs a cloud platform neutral architecture, which is compatible with multiple mainstream cloud platforms and local Internet data center (IDC) environment through a bidirectional synchronization gateway, establishes a disaster recovery system decoupled from cloud vendors, and through actual application, it can be verified that in the event of cloud platform regional service interruption, the system can automatically trigger cross-cloud switching, realize service recovery with RPO less than 30 seconds and RTO less than 10 minutes, and avoid huge economic loss.

[0189] In view of the problem of disconnection between disaster recovery design and implementation, the design scheme and implementation are disconnected in the traditional mode, resulting in more than 40% of the disaster recovery resource configuration not matching the actual load of the business. The present application adopts a dynamic demand perception algorithm, constructs a resource prediction model based on 12 performance indicators such as real-time IOPS and throughput of business systems, and combines sandbox pre-validation technology to make the resource configuration accuracy of the project reach 97.3%, avoiding 40% resource waste caused by design deviation in the traditional scheme.

[0190] In view of the problem of low efficiency of multi-node monitoring and switching, the existing system relies on manual inspection of the state of each component, and disaster recovery switching needs to manually perform 20+ operation steps, with an average time consumption of more than 2 hours. The present application establishes an intelligent monitoring system covering more than 200 types of device indicators through the synergistic effect of topology perception engine and automated workflow engine, simplifies more than 20 steps of switching operation with human participation into 3 automated strategies, and through actual application, it can be verified that in system drills, it can realize millisecond-level linkage of database cluster switching and DNS redirection, and compress the overall RTO from 4.2 hours to 2 minutes and 15 seconds.

[0191] In view of the limitation of heterogeneous data synchronization in the localization replacement trend, the traditional scheme only supports specific databases (such as Oracle / MySQL), and the coverage rate of the localization environment (Dream / People's University Golden Storehouse) and the file system (ext4 / xfs / GPFS) is less than 60%, the system of the present application is built-in intelligent data bus, which integrates 15 kinds of protocol conversion modules of domestic / international databases, and solves the real-time synchronization problem of structured data of different databases, and supports the unified scheduling of third-party tools (ADG / OGG), and through actual application, it can be verified that in the localization reconstruction project, cross-platform synchronization of 2TB data per day can be realized, and the success rate is 99.98%;

[0192] In view of the problem of missing disaster recovery verification, the industry generally lacks automatic verification means, and 70% of disaster recovery systems are not tested regularly, the unique isolated sandbox exercise system of the present application builds a simulation library containing 200 kinds of fault scenes such as network partition and storage degradation, through the automatic multi-dimensional evaluation report generation mechanism, through actual application, it can be verified that it can help the hospital to complete the quarterly disaster recovery exercise under the premise of zero business impact, find and repair 17 potential hidden dangers, and improve the actual availability of the system from 68% to 99.95%.

[0193] The large-scale application of the present application significantly promotes the paradigm shift of disaster recovery system construction from compliance to capacity driving, and provides key technical support for various industries to cope with new challenges such as hybrid cloud architecture, Leiyuan replacement and ransomware attacks.

[0194] In addition, it should be noted that the present application can be provided as a method, device or computer program product. Therefore, the embodiments of the present application can adopt a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer usable storage media containing computer usable program code.

[0195] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams according to the method, terminal device (system) and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device realize the functions specified in the flowcharts and / or block diagrams. Figure 1 The device for realizing the functions specified in one flow or multiple flows and / or blocks Figure 1 The device for realizing the functions specified in one flow or multiple flows and / or blocks

[0196] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow or flows and / or blocks Figure 1 function specified in the flow or flows and / or blocks Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow or flows and / or blocks

[0197] It is also noted that the illustrative language used herein is merely intended to educate the public and it is not intended to limit the scope of the present application. It is further noted that the various steps or functions outlined herein can be implemented in any order or in parallel, unless otherwise specified or unless the order clearly could not be performed in parallel. It is also noted that the various steps or functions outlined herein can be performed by any entity or set of entities, unless otherwise specified or unless the order clearly could not be performed by the same entity or set of entities. It is further noted that the terms "comprise", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Furthermore, unless otherwise indicated, the steps or functions outlined herein can be performed in any order or in parallel, unless otherwise specified or unless the order clearly could not be performed in parallel. Also, unless otherwise indicated, the various steps or functions outlined herein can be performed by any entity or set of entities, unless otherwise specified or unless the order clearly could not be performed by the same entity or set of entities.

[0198] Finally, it is to be understood that the above description is intended to be illustrative, and not restrictive. Many other embodiments will be apparent to those of skill in the art upon reviewing the above description. The scope of the application should, therefore, be determined not with reference to the above description, but instead with reference to the appended claims, along with their full scope of equivalents.

Claims

1. An intelligent integrated disaster recovery method, characterized in that: include: Connect to X86 architecture devices and ARM architecture devices, and call the control interface corresponding to each device in the preset heterogeneous device driver library; Access the computing resources of each device i through the control interface i , Storage resources i and network resources Network i , build a heterogeneous resource pool ResourcePool: Based on software-defined storage and software-defined networking technologies, each resource in the resource pool is virtualized into standardized resource units; Extract attribute parameters for each resource unit j to form a resource identification tuple R corresponding to the resource unit j j , and construct a unified resource abstract view table VRDL = [R1, R2, R3, ..., R j ,...]; wherein the attribute parameters include at least central processing unit (CPU) resources, the number of read and write operations per second (IOPS) and latency; Based on the constructed resources, business data synchronization is carried out, and the business load is monitored for each resource unit j and the resource load index W at time t is calculated. t : W t =a·CPU util +b·IOPS+c·Latency Among them, CPU util Indicates CPU usage, α, β, and γ are weight coefficients; In W t When the preset threshold value Threshold is exceeded, a disaster recovery prompt is triggered and a disaster recovery switching operation is performed using the VRDL.

2. The intelligent integrated disaster recovery method according to claim 1, wherein: The calling of the control interface corresponding to each of the devices in the preset heterogeneous device driver library includes: For the X86 architecture device or the ARM architecture device, parse the BIOS information and network card MAC address of the device; The BIOS information and the MAC address of the network card are used to match the device compatibility table items, and according to the matching result, the control interface corresponding to the device in the heterogeneous device driver library is called.

3. The intelligent integrated disaster recovery method according to claim 1, wherein: The software-defined storage and software-defined networking technologies are used to virtualize the resources in the resource pool into standardized resource units, including: Mapping the computing resources into virtual CPU units and registering them as computing resource units through a standard interface; Dividing the storage resource into labeled logical volume blocks and registering them as storage resource units through a standard interface; Create a virtual network interface for the network resource, configure bandwidth limits, forwarding rules and QoS parameters, and encapsulate it into a network resource unit through a standard interface.

4. The intelligent integrated disaster recovery method according to claim 1, wherein: The method further comprises: The weight coefficients α, β, and γ are dynamically updated based on historical monitoring data of the business load using a sliding window weighted average method.

5. The intelligent integrated disaster recovery method according to claim 1, wherein: The performing the disaster recovery switching operation by using the VRDL includes: The priority score of the resource to be scheduled in the VRDL is calculated using the following formula: P=δ*QoS score +μ*SLA weight +η*Latency Among them, δ, μ, η are weight coefficients, QoS score Indicates the quality of service (SLA) of a resource weight Indicates the importance of the resource.

6. The intelligent integrated disaster recovery method according to claim 1, wherein: The performing the disaster recovery switching operation by using the VRDL includes: Determining at least one candidate resource in the VRDL for performing a disaster recovery switching operation; Isolate resource copies to build a sandbox environment, execute fault scripts, and perform sandbox pre-exercises on the candidate resources; the fault scripts include: network interruption simulation, storage input and output congestion simulation, and / or system call blocking simulation; Based on the results of the sandbox pre-drill, resources for performing a disaster recovery switching operation are determined.

7. The intelligent integrated disaster recovery method according to claim 1, wherein: The method further comprises: Performing a benchmark test on the performance of each resource unit j and recording initial performance parameters of the resource unit j; Performing a performance check on the resource unit j according to a predetermined period, wherein the performance check includes: whether there is hardware failure, performance degradation, network delay, and / or abnormal resource consumption; When the result of the performance check is that the performance is abnormal, a disaster recovery prompt is triggered and a disaster recovery switching operation is performed using the VRDL.

8. The intelligent integrated disaster recovery method according to claim 7, characterized in that: The method further comprises: For the resource unit j whose performance check result indicates performance abnormality, determining a cause of the abnormality of the resource unit j; Take preset corresponding repair measures for the determined abnormal cause and perform fault recovery operations; the repair measures may include: restarting the device, changing the network channel, scheduling backup resources and / or initiating manual intervention.

9. The intelligent integrated disaster recovery method according to claim 1, wherein: The method further comprises: A resource demand prediction model is constructed using a time series analysis algorithm, wherein the samples of the model are the load data of each resource unit j in different time periods; The resource demand prediction model is used to predict resource demand, and the allocation of each resource unit in the resource pool is dynamically adjusted according to the prediction result.

10. An intelligent integrated disaster recovery system, characterized in that: include: Device access module, resource pool construction module, resource virtualization module, view construction module, load calculation module and resource scheduling module; The device access module is used to access X86 architecture devices and ARM architecture devices, and call the control interface corresponding to each of the devices in the preset heterogeneous device driver library; The resource pool construction module is used to access the computing resources Computer of each device i through the control interface i , Storage resources i and network resources Network i , build a heterogeneous resource pool ResourcePool: The resource virtualization module is used to virtualize each resource in the resource pool into a standardized resource unit based on software-defined storage and software-defined network technologies; The view construction module extracts attribute parameters for each resource unit j to form a resource identification tuple R corresponding to the resource unit j. j , and construct a unified resource abstract view table VRDL = [R1, R2, R3, ..., R j ,...]; wherein the attribute parameters include at least central processing unit (CPU) resources, the number of read and write operations per second (IOPS) and latency; The load calculation module is used to synchronize business data based on the constructed resources, monitor the business load for each resource unit j and calculate the resource load index W at time t. t : W t =a·CPU util +b·IOPS+c·Latency Among them, CPU util Indicates CPU usage, α, β, and γ are weight coefficients; The resource scheduling module is used to t When the preset threshold value Threshold is exceeded, a disaster recovery prompt is triggered and a disaster recovery switching operation is performed using the VRDL.

Citation Information

Patent Citations

  • Heterogeneous resource-oriented multi-dimensional scheduling system

    CN107066319A