Disaster recovery method, device and system

CN116302691BActive Publication Date: 2026-09-22ALIBABA CLOUD COMPUTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310181688.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-23
Publication Date
2026-09-22
Estimated Expiration
2043-02-23

AI Technical Summary

Technical Problem

[0002]随着计算机技术的高速发展,越来越多的数据采用电子化的形式进行存储和处理,数据管理系统己经成为企业运转的关键,而这种数据的存储和处理方式虽然能够提高便利性,一旦数据管理系统发生灾难,会很容易丢失和损坏数据

Benefits of technology

[0021]本说明书一个实施例提供的容灾方法,应用于调度平台,调度平台为第一机房和第二机房提供一组管理服务,在第一机房故障的情况下,调用第二机房的切换服务,对第二机房中的数据存储节点,以及第二机房中的元数据库实例节点进行更新;调用第二机房中管控服务单元,将第一机房对应的数据处理任务分配至第二机房,并对第二机房中数据库实例节点进行更新;基于第二机房中更新后的数据存储节点、元数据库实例节点以及数据库实例节点,向第二机房同步数据;根据数据同步结果,向第二机房发送执行元数据处理任务和数据处理任务的执行指令。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116302691B_ABST
    Figure CN116302691B_ABST
Patent Text Reader

Abstract

The embodiment of the present specification provides a disaster recovery method, device and system, wherein the disaster recovery method is applied to a scheduling platform, wherein the scheduling platform provides a group of management services for a first machine room and a second machine room, including: in the case that the first machine room fails, calling a switching service of the second machine room, updating data storage nodes in the second machine room and meta database instance nodes in the second machine room; calling a management and control service unit in the second machine room, distributing data processing tasks corresponding to the first machine room to the second machine room, and updating database instance nodes in the second machine room; synchronizing data to the second machine room based on the updated data storage nodes, meta database instance nodes and database instance nodes in the second machine room; and sending an execution instruction of executing a meta data processing task and a data processing task to the second machine room according to a data synchronization result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of computer technology, and in particular to a disaster recovery method, apparatus, and system. Background Technology

[0002] With the rapid development of computer technology, more and more data is stored and processed electronically. Data management systems have become crucial for enterprise operations. While this method of data storage and processing improves convenience, data can easily be lost or corrupted in the event of a disaster. Current technologies typically employ a dual-data center setup, with two independent data service providers deployed in the primary and backup data centers. In the event of a primary data center failure, the backup data center takes over its functions. However, this deployment method requires the deployment and management of two sets of data services, increasing deployment costs and operational complexity, and also resulting in a slow failover speed. Therefore, a disaster recovery method is urgently needed to address these issues. Summary of the Invention

[0003] In view of this, embodiments of this specification provide a disaster recovery method. One or more embodiments of this specification also relate to a disaster recovery device, a disaster recovery system, a computing device, a computer-readable storage medium, and a computer program, to address the technical deficiencies existing in the prior art.

[0004] According to a first aspect of the embodiments of this specification, a disaster recovery method is provided, applied to a scheduling platform, wherein the scheduling platform provides a set of management services for a first data center and a second data center, including:

[0005] In the event of a failure in the first data center, the switching service of the second data center is invoked to update the data storage nodes and metadata instance nodes in the second data center.

[0006] The management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center, and the database instance nodes in the second data center are updated.

[0007] Based on the updated data storage nodes, metadata instance nodes, and database instance nodes in the second data center, synchronize data to the second data center;

[0008] Based on the data synchronization results, the execution instructions for the metadata processing task and the data processing task are sent to the second computer room.

[0009] According to a second aspect of the embodiments of this specification, a disaster recovery device is provided, comprising:

[0010] The update module is configured to invoke the switching service of the second data center in the event of a failure in the first data center, and update the data storage nodes and metadata instance nodes in the second data center.

[0011] The allocation module is configured to call the management and control service unit in the second data center to allocate the data processing tasks corresponding to the first data center to the second data center and update the database instance nodes in the second data center;

[0012] The synchronization module is configured to synchronize data to the second data center based on the updated data storage nodes, metadata instance nodes, and database instance nodes in the second data center;

[0013] The sending module is configured to send the execution instructions for the metadata processing task and the data processing task to the second data center based on the data synchronization results.

[0014] According to a third aspect of the embodiments of this specification, a disaster recovery system is provided, comprising:

[0015] The system comprises a first data center, a second data center, and a scheduling platform. The scheduling platform stores executable data synchronization instructions. When these instructions are executed by the scheduling platform, they implement the steps of the disaster recovery method described above, and are used to allocate the data stored in the first data center and the data processing tasks to the second data center.

[0016] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising:

[0017] Memory and processor;

[0018] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the disaster recovery method described above.

[0019] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions that, when executed by a processor, implement the steps of the disaster recovery method described above.

[0020] According to a sixth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the disaster recovery method described above.

[0021] The disaster recovery method provided in one embodiment of this specification is applied to a scheduling platform. The scheduling platform provides a set of management services for a first data center and a second data center. In the event of a failure in the first data center, the platform invokes the switching service of the second data center to update the data storage nodes and metadata instance nodes in the second data center; it also invokes the management and control service unit in the second data center to allocate the data processing tasks corresponding to the first data center to the second data center and update the database instance nodes in the second data center; based on the updated data storage nodes, metadata instance nodes, and database instance nodes in the second data center, it synchronizes data with the second data center; and based on the data synchronization results, it sends execution instructions for metadata processing tasks and data processing tasks to the second data center.

[0022] The first and second data centers use the same set of management services provided by the scheduling platform, which reduces service deployment costs and operational complexity. In the event of a failure in the first data center, the backup and switchover service of the second data center can be directly invoked, enabling the second data center to replace the first data center and continue to provide services to the outside world. This meets the switching needs of the disaster recovery system while improving service switching speed and fault recovery efficiency. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of a disaster recovery method provided in one embodiment of this specification;

[0024] Figure 2 This is a flowchart illustrating a disaster recovery method provided in one embodiment of this specification;

[0025] Figure 3 This is a flowchart illustrating the process of a disaster recovery method provided in one embodiment of this specification.

[0026] Figure 4 This is an architecture diagram of a disaster recovery system provided in one embodiment of this specification;

[0027] Figure 5 This is a schematic diagram of the structure of a disaster recovery device provided in one embodiment of this specification;

[0028] Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0029] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0030] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0031] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0032] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0033] K8S (Kubernetes) is a portable, scalable, open-source platform for managing containerized workloads and services, facilitating declarative configuration and automation.

[0034] DBaaS (Database as a Service): This refers to database services provided by cloud providers. Compared to traditional self-built databases, cloud-provided DBaaS systems include the database kernel and a complete ecosystem, including lifecycle management, backup and recovery, and manual or automated operation and maintenance services.

[0035] RPO (Recovery Point Objective): Data recovery point target, measured in time. It represents the required point in time when the system and data must be recovered in the event of a disaster. RPO indicates the maximum amount of data loss the system can tolerate. The smaller the amount of data loss the system can tolerate, the smaller the RPO value.

[0036] RTO (Recovery Time Objective): Recovery Time Objective, expressed in time. It represents the time required for an information system to recover from a point of failure after a disaster. RTO indicates the maximum tolerable service downtime a system can tolerate. The higher the urgency of the system's service requirements, the lower the RTO value.

[0037] Paxos Protocol: The Paxos protocol is one of the few decentralized distributed protocols proven in engineering practice to offer strong consistency and high availability. In Paxos, there is a group of completely peer-to-peer participating nodes, each making a decision on a given event. A decision becomes effective if it receives the approval of more than half of the nodes. Paxos can function as long as more than half of the nodes are operational, and it effectively handles anomalies such as downtime and network fragmentation.

[0038] Three data centers in the same city: A private cloud deployment model characterized by: 1. Interconnectivity among the three data centers; 2. Low latency (network latency <1.5ms) due to their shared location; 3. A lightweight data center that does not deploy cloud infrastructure or ecosystem services, serving only the database. The third data center, together with the primary and backup data centers, forms a lightweight distributed disaster recovery environment, reducing disaster recovery construction costs and achieving zero data loss.

[0039] Active-active: An application deployment mode in which multiple process replicas run on different physical machines without distinguishing between primary and backup roles, and each application backup provides services.

[0040] Primary / standby: An application deployment mode in which one or more of the multiple process replicas are primary nodes, providing services, while the rest are standby nodes, serving as data backups for the primary nodes; the primary and standby roles can be switched between each other in some applications.

[0041] XDB: A product form of RDS-MySQL, a clustered MySQL instance that uses the Paxos protocol to ensure data consistency among multiple nodes, with at least 3 nodes under normal conditions.

[0042] Figure 1 This is a schematic diagram of a disaster recovery method provided in one embodiment of this specification, as shown below. Figure 1As shown, the resource scheduling platform provides the same set of management services for both the first and second data centers. The first data center corresponds to a first switching service, a management and control service layer, database instances, data storage nodes, and metadata database instance nodes. Correspondingly, the second data center corresponds to a second switching service, a management and control service layer, database instances, data storage nodes, and metadata database instance nodes. Data synchronization between database instances in the first and second data centers, between data storage nodes in the first and second data centers, and between metadata database instance nodes in the first and second data centers is achieved via the Paxos protocol. In the event of a failure in the first data center, the resource scheduling platform invokes the switching service in the second data center to update the data storage nodes and metadata database instance nodes in the second data center; it also invokes the management and control service unit in the second data center to allocate data processing tasks corresponding to the first data center to the second data center and update the database instance nodes in the second data center; based on the updated data storage nodes, metadata database instance nodes, and database instance nodes in the second data center, it synchronizes data to the second data center; and based on the data synchronization results, it sends execution instructions for metadata processing tasks and data processing tasks to the second data center.

[0043] A single management service manages all database instances and control components. Metadata is physically replicated, ensuring rapid data synchronization and minimizing data loss under low latency (often less than 1.5 milliseconds) in both primary and secondary data centers within the same city. It achieves kernel-level physical replication without relying on external components, ensuring security and efficiency. Using a single management service allows for active-active control components, minimizing service switching time and reducing operational costs.

[0044] This specification provides a disaster recovery method, and also relates to a disaster recovery device, a disaster recovery system, a computing device, a computer-readable storage medium, and a computer program, which are described in detail in the following embodiments.

[0045] See Figure 2 , Figure 2 A flowchart of a disaster recovery method according to an embodiment of this specification is shown. The disaster recovery method is applied to a scheduling platform that provides a set of management services for a first data center and a second data center, and specifically includes the following steps.

[0046] Step S202: In the event of a failure in the first data center, invoke the switching service of the second data center to update the data storage nodes and metadata instance nodes in the second data center.

[0047] Specifically, disaster recovery refers to maintaining uninterrupted system operation while minimizing data loss during disasters such as natural disasters, equipment failures, and human-caused damage. Disaster recovery systems typically consist of two or more functionally identical systems that can perform health checks and function switching. When one system stops working due to an accident (such as a fire or earthquake), the entire application system can switch to another, allowing that system to continue functioning normally. Correspondingly, the first and second data centers represent two functionally identical systems. Both data centers have the same data center structure, typically consisting of a cloud infrastructure layer, a management and control service layer, and a kernel service layer. The first data center can be either the primary or backup data center. The second data center can serve as a backup data center, providing services in the event of a failure in the first data center. In the disaster recovery method provided in this embodiment, a scheduling platform provides a set of management services for both the first and second data centers. This scheduling platform includes, but is not limited to, a Kubernetes platform. The switching service for the second data center refers to the service deployed in the second data center, which is invoked when the first data center fails, allowing the second data center to take over the tasks of the first data center. The cloud infrastructure layer of the second data center includes a metadata instance unit and a data storage unit. The metadata instance unit contains multiple metadata instance nodes for managing metadata and achieving data synchronization. The data storage unit contains multiple data storage nodes for achieving data synchronization.

[0048] Based on this, the scheduling platform can simultaneously provide the same set of management services for both the first and second data centers, used to manage the data storage units and metadata instance units in both data centers. In the event of a failure in the first data center, the platform invokes the failover service of the second data center to update the data storage nodes of the data storage units and the metadata instance nodes of the metadata instance units in the second data center, changing their attributes to enable them to provide services externally.

[0049] In practical applications, to improve data center disaster recovery capabilities, a third data center is included in addition to the first and second data centers, constructing a three-data center disaster recovery architecture within the same city. The third data center does not require the deployment of unnecessary cloud infrastructure and management components, reducing the material requirements and operational costs of the disaster recovery architecture. A cloud infrastructure is deployed in the first and second data centers to manage the database instances and management components in both data centers. The original data in the cloud infrastructure can be physically replicated, ensuring rapid data synchronization and preventing data loss under the low latency conditions of the first and second data centers. The first and second data centers use the same management service, enabling cross-data center deployment of database instances. Kernel-level physical replication can be achieved without relying on external components, improving the security and efficiency of data synchronization.

[0050] Furthermore, considering that the second data center contains at least one second data storage node and at least one second metadata database instance node, and that only one of the second data storage node and the second metadata database instance node can provide services externally, while the other nodes are only used for data synchronization, it is necessary to update the second data storage node and the second metadata database instance node contained in the second data center when the second data center is used as a data center to provide services externally. The specific implementation is as follows:

[0051] The system invokes the switching service of the second data center, selects a data storage node from the second data storage nodes contained in the second data center, and updates the data storage node. The updated data storage node is used to receive storage node data submitted for the second data center. The system also invokes the switching service of the second data center, selects a metadata instance node from the second metadata instance nodes in the second data center, and updates the metadata instance node. The updated metadata instance node is used to receive metadata instance data submitted for the second data center.

[0052] Specifically, a second data storage node refers to a data storage node within a second data center used for data synchronization. Correspondingly, a data storage node is one or more selected from multiple second data storage nodes to provide data services externally, receiving storage node data submitted to the second data center. Other nodes within the second data storage node are only used for data synchronization and do not need to provide data services externally. Similarly, a second metadata instance node refers to a metadata instance node within a second data center used for data synchronization. Correspondingly, a metadata instance node is one or more selected from multiple second metadata instance nodes to provide data services externally, receiving storage node data submitted to the second data center. Other nodes within the second metadata instance node are only used for data synchronization and do not need to provide data services externally.

[0053] Based on this, the switching service of the second data center is invoked, a data storage node is selected from at least two second data storage nodes contained in the second data center, and the data storage node is updated. The updated data storage node is used to receive storage node data submitted for the second data center. The second data storage nodes in the second data center are all in a standby state before the first data center fails, and are only used for data synchronization between the first and second data centers. When the first data center fails and the second data center takes over the service, a data storage node is selected from at least two second data storage nodes and updated to a service state to provide services externally. The second data center's switchover service is invoked, and a metadata database instance node is selected from at least two second metadata database instance nodes in the second data center and updated. The updated metadata database instance node is used to receive metadata database instance data submitted to the second data center. Correspondingly, the second metadata database instance nodes in the second data center are all in a standby state before the first data center fails, and are only used for data synchronization between the first and second data centers. When the first data center fails and the second data center takes over the service, a metadata database instance node is selected from at least two second metadata database instance nodes and updated to a service state to provide services externally.

[0054] For example, besides the database kernel, database services also include various peripheral management systems that implement functions such as instance lifecycle management. When a data center-level disaster occurs, it's necessary to ensure high availability of both the management system and the database instances, quickly and stably switching services to a second data center while maintaining data consistency and functional integrity. In a scenario with three data centers in the same city, if the first data center fails, the second data center takes over and provides services. The management metadata database in the second data center contains three XDB leader nodes, and the Kubernetes base corresponds to three etcd leader nodes (metadata service nodes). When the first data center fails and the second data center is activated, one of the three XDB leader nodes is arbitrarily selected and updated to an XDB leader node to provide external services, while the other two XDB leader nodes continue asynchronous data synchronization. Similarly, one of the three etcd leader nodes is arbitrarily selected and updated to an etcd leader node to provide external services, while the other two etcd leader nodes continue asynchronous data synchronization.

[0055] In summary, updating data storage nodes and metadata database instance nodes within the second data storage nodes in the second data storage room is achieved by selecting data storage nodes and updating metadata database instance nodes within the second metadata database instance nodes in the second data storage room. This enables the updating of some nodes within the second data storage room, thereby ensuring the continuity of services provided to the outside world.

[0056] Step S204: Call the management and control service unit in the second data center to allocate the data processing tasks corresponding to the first data center to the second data center, and update the database instance nodes in the second data center.

[0057] Specifically, both the first and second data centers deploy management and control service layers, each with multiple management and control service units. The second data center deploys the same number of management and control service units as the first data center; that is, each management and control service unit in the first data center has a corresponding replica management and control service unit in the second data center. These management and control service units provide management and control services to the data centers. When the first data center experiences a failure, all data processing tasks corresponding to the first data center are distributed to the second data center. Data processing tasks refer to tasks received during the operation of the first data center for data processing. Database instance nodes refer to the database instances in the second data center corresponding to the service layer, used for data processing. The database instance units corresponding to the first and second data centers use a strongly consistent data synchronization protocol, such as the Paxos protocol, for data synchronization.

[0058] Therefore, in the event of a failure in the first data center, the management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center. After the failover service in the second data center is invoked to update the data storage nodes and metadata instance nodes in the second data center, the management and control service in the second data center can operate normally, receiving and processing data tasks. The management and control service unit in the second data center is then invoked to update the database instance nodes in the second data center, enabling the database instance nodes in the processing ready state to provide services externally.

[0059] In practical applications, when the first data center is functioning normally and providing services, the database instance nodes in the first data center act as leader nodes, i.e., primary nodes. During normal operation of the first data center, the leader nodes provide write services, and the written data is synchronized to the database instance nodes (follower nodes) in the second and third data centers via the Paxos protocol. When the first data center fails and cannot provide services, the second data center is activated, and the database instance nodes (follower nodes) in the second data center are updated to become the leader nodes. The database instance nodes in the second data center then provide write services, and the written data is synchronized to the database instance nodes (follower nodes) in the third data center via the Paxos protocol. This ensures that even if the first data center fails, services can still be provided normally and uninterruptedly.

[0060] Furthermore, considering that the database instance nodes in the first data center were in a ready state before the failure of the first data center, and that the second data center needs to take over the service after the failure, the database instance nodes in the second data center need to provide services externally. The specific implementation is as follows:

[0061] The management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center, and to update the status and attribute information of the database instance nodes in the second data center. The updated database instance nodes are used to receive database instance data submitted to the second data center, and the management and control service unit is used to manage the status of the database instance nodes in the second data center.

[0062] Specifically, status information refers to the status information of database instance nodes, including but not limited to ready status and service status. When a database instance node in the second data center is in the ready status, it does not provide services to the outside world. When a database instance node in the second data center is in the service status, it provides services to the outside world. Corresponding to the status of the database instance node, the database instance node also has attribute information. When the database instance node in the second data center is a leader node, it can provide services to the outside world. When the database instance node in the second data center is a follower lower node, it does not provide services to the outside world and is only used to realize data synchronization between the first and second data centers.

[0063] Based on this, the management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center, and to update the status and attribute information of the database instance nodes in the second data center. Before the failure of the first data center, the database instance nodes in the second data center were in a ready state, which was a follow lower node. After the failure of the first data center, the second data center provides services to the outside world. The ready state of the database instance nodes in the second data center is updated to the service state, and the follow lower nodes are updated to leader nodes. The updated database instance nodes are used to receive database instance data submitted to the second data center. The management and control service unit is used to manage the status of the database instance nodes in the second data center.

[0064] Continuing the previous example, when the second data center takes over service from the first data center, the database instance node (XDB instance) in the second data center is updated from lower to leader to provide services externally. Management services are all deployed with a replica in the second data center, employing a multi-active architecture. In the event of a failure in the first data center, Kubernetes automatically transfers traffic to the second data center.

[0065] In summary, by updating the status and attribute information of the database instance nodes in the second data center, and then providing services to the outside world based on the database instance nodes in the second data center, the continuity of services can be guaranteed in the event of a failure in the first data center.

[0066] Furthermore, considering the security of data processing and data storage, a third data center can be deployed in addition to the first and second data centers. The third data center does not require unnecessary management, failover, or control services; it only contains database instance nodes for synchronizing data from the database instance units in the second data center. The specific implementation is as follows:

[0067] Based on the updated database instance node, the instance data in the second data center is synchronized to the third database instance node in the third data center, and the third database instance node is used to perform asynchronous storage of the instance data.

[0068] Specifically, the third data center refers to another data center provided in addition to the first and second data centers. The third data center only includes database instance nodes, which are used to synchronize the data corresponding to the database instance nodes in the second data center based on the data synchronization protocol. There is no need to deploy unnecessary control services, management services, etc. in the third data center; only the database instance nodes need to be deployed.

[0069] Based on this, in addition to the first and second data centers, a third data center is also included. When the second data center is able to provide services to the outside world normally, the instance data corresponding to the database instance nodes in the second data center is synchronized to the third database instance nodes in the third data center, and the data synchronization is performed through an asynchronous synchronization strategy.

[0070] Continuing with the previous example, in addition to the first and second data centers, a third data center is also deployed. The XDB instance in the third data center is a lower node, and the XDB instance in the third data center synchronizes data with the XDB instance in the second data center through the Paxos protocol.

[0071] In summary, in addition to the first and second data centers, a third data center is also included. The database instance nodes in the third data center are used to achieve asynchronous data synchronization. Furthermore, there is no need to deploy unnecessary control and management services in the third data center, thereby reducing material consumption and lowering the operating costs of the data center.

[0072] Furthermore, considering the disaster recovery capabilities of the data center architecture, in order to minimize data loss and ensure task continuity during the switchover between the first and second data centers, data center switching can be implemented through either proactive or reactive methods, as detailed below:

[0073] The system invokes the switching service of the second data center to send an update task to the management service unit in the second data center; it then invokes the management service unit in the second data center that receives the update task to update the database instance nodes in the second data center; or, it receives an update instruction for the second data center and updates the database instance nodes in the second data center based on the update instruction; the updated database instance nodes in the second data center are then used to receive data operation tasks.

[0074] Specifically, an update task refers to a task issued to the management service unit of the data center that will provide services when the first data center fails or when there is a need for data center switching. This task informs the management service unit in the data center that there is a data center switching task, and then calls the management service unit to assist in switching between the first and second data centers. Correspondingly, an update command refers to a computer command submitted by the user to the scheduling platform when the first data center fails, which is used to switch between the first and second data centers.

[0075] Therefore, in the event of a failure in the primary data center, the primary data center will report the failure to the scheduling platform. The scheduling platform will then invoke the switchover service of the secondary data center to send an update task to the management service unit in the secondary data center. After receiving the update task, the management service unit in the secondary data center will invoke the management service unit that received the update task to update the database instance nodes in the secondary data center. Alternatively, in the event of a failure in the primary data center, since users can perceive the failure, they can submit an update command to the scheduling platform for the secondary data center. Upon receiving the update command for the secondary data center, the scheduling platform will update the database instance nodes in the secondary data center based on the update command. The updated database instance nodes in the secondary data center will then be used to receive data operation tasks, allowing the secondary data center to take over the service provided by the primary data center.

[0076] Following the previous example, when switching database instances between the first and second data centers, the switch can be automatic or manual. Whether the cloud infrastructure's etcd (metadata service) and management metadata database are switched, and whether some database instances lacking automatic switching capabilities are switched, is handled uniformly by the disaster recovery switching service to ensure an orderly switch. In special cases, users are also given the option not to switch services to the second data center, choosing to sacrifice service continuity to ensure system data consistency. Database instance switching needs to be discussed in different categories: Since the XDB instances in the first and second data centers use a strongly consistent Paxos protocol, the RPO can be guaranteed to be zero even in the event of a failure in the first data center, and a leader node will be automatically determined in the second data center. In this case, both RPO and RTO are determined by the kernel. Some database instances lack automatic switching capabilities. In this case, a switch task needs to be issued through the disaster recovery switching service, and the management service will perform the primary / standby switch. In this case, the RPO and RTO of the database instance are determined by the recovery speed of the management service. Since the recovery of the management service depends on the recovery of the cloud infrastructure's management metadata database, the final RTO is determined by the Kubernetes infrastructure metadata database.

[0077] In summary, in the event of a failure in the primary data center, the dispatch platform can automatically switch between the primary and secondary data centers. Alternatively, users can manually switch according to their needs, thereby improving the flexibility of switching between the primary and secondary data centers. The automatic switching by the dispatch platform avoids the need for users to perform real-time fault checks, thus improving the fault response speed.

[0078] Furthermore, considering the uncertainty of when the first data center will fail, the metadata in the second data center will not be consistent with that in the first data center. Therefore, the management and control service unit needs to correct the metadata to restore it. The specific implementation is as follows:

[0079] Obtain historical data corresponding to the fault in the first data center; invoke the management and control service unit in the second data center to update the metadata of the metadata database instance node in the second data center based on the historical data within a preset recovery time.

[0080] Specifically, historical data refers to the metadata contained in the first data center before the failure of the first data center; the preset recovery time refers to the pre-set data recovery time, indicating that in the event of a disaster, data and services need to be restored within the preset recovery time (RTO). RTO marks the maximum tolerable service downtime; the higher the urgency of service recovery, the smaller the RTO value. When the first data center fails, after the management services in the second data center recover, metadata correction is performed first to ensure that subsequent services can continue normally.

[0081] Therefore, when the first data center fails, after the management and control service unit in the second data center recovers, it retrieves the historical data stored in the scheduling platform prior to the failure of the first data center. It then calls the management and control service unit in the second data center to update the metadata of the metadata database instance nodes in the second data center based on the retrieved historical data within a preset recovery time, so that the second data center can subsequently take over the service provided by the first data center.

[0082] Following the previous example, the master-slave switchover of management parameters in the DBaaS system is often very important. Inconsistency between management parameters and kernel state may lead to the incorrect distribution of switchover task streams, causing secondary disasters. Therefore, after the management service recovers, it will also correct the metadata as soon as possible. This time is determined by the RTO of the management service. Since the recovery of the management service depends on the recovery of the underlying management metadata database, it is ultimately determined by the RTO of the underlying metadata database.

[0083] In summary, after the management and control service unit in the second data center is restored, the metadata in the second data center is corrected based on the historical data corresponding to the failure in the first data center, so that the second data center can take over the service provided by the first data center in the future.

[0084] Step S206: Synchronize data to the second data center based on the updated data storage nodes, metadata instance nodes, and database instance nodes in the second data center.

[0085] Based on this, after the data storage nodes, metadata instance nodes, and database instance nodes in the second data center are updated, data can be synchronized to the second data center based on the updated data storage nodes, metadata instance nodes, and database instance nodes in the second data center. This is used to prepare data for the second data center to provide services to the outside world, so that the second data center can take over the service from the first data center.

[0086] In practical applications, the disaster recovery architecture consisting of a first, second, and third data center has the capability for primary / standby failover, ensuring high availability. In the event of a disaster, it can automatically switch from the first data center to the second, or users can manually decide to temporarily stop services from either the first or second data center to avoid data inconsistencies in some database instances. By adopting a three-data center disaster recovery architecture, database kernels with disaster recovery capabilities can achieve an RPO of 0, meeting financial-grade disaster recovery requirements.

[0087] Furthermore, since the amount of data processed in the first data center is large, data synchronization between the first and second data centers via external components would require additional deployment of external components. To improve data synchronization efficiency and consistency, data synchronization between the first and second data centers can be based on a data synchronization protocol, as implemented below:

[0088] Determine the data storage space, read the stored data in the data storage space, and synchronize the stored data to the updated data storage node in the second data center based on the data synchronization protocol; read the metadata in the data storage space, and synchronize the metadata to the updated metadata database instance node in the second data center based on the data synchronization protocol; read the instance data in the data storage space, and synchronize the instance data to the updated database instance node in the second data center based on the data synchronization protocol.

[0089] Specifically, the data storage space can be local storage or cloud storage. This space records data in the first data center in real time and can be backed up using snapshots for subsequent data recovery. The stored data is the backup data corresponding to the snapshot taken in the first data center before its failure. Metadata is used for data synchronization with the second data center. Instance data is used for data synchronization with the second data center. The main purpose of snapshots is to enable online data recovery. When application failures or file corruption occur on the storage device, timely data recovery can be performed, restoring the data to the state at the time the snapshot was taken. The data synchronization protocol can be a strongly consistent protocol such as Paxos for data synchronization.

[0090] Based on this, a data storage space for storing the snapshot of the first data center is determined, and the corresponding stored data of the first data center is read from the data storage space. The stored data is then synchronized to the updated data storage node in the second data center based on a data synchronization protocol. Metadata in the data storage space is read and synchronized to the updated metadata instance node in the second data center based on a data synchronization protocol. Instance data in the data storage space is read and synchronized to the updated database instance node in the second data center based on a data synchronization protocol. This synchronizes the data in the updated data storage node, metadata instance node, and database instance node in the second data center to the state before the failure of the first data center, enabling the second data center to take over and continue providing services.

[0091] Continuing with the previous example, data synchronization between the first and second data centers is implemented using the Paxos protocol. Specifically, data synchronization is achieved between the XDB instances in the first and second data centers, between the management metadata databases in the first and second data centers, and between the etcd (metadata service) nodes in the first and second data centers.

[0092] In summary, the data synchronization protocol enables data synchronization between the data storage nodes, metadata instance nodes, and database instance nodes in the first and second data centers, thereby improving the efficiency and accuracy of data synchronization.

[0093] Step S208: Based on the data synchronization result, send the execution instructions for the metadata processing task and the data processing task to the second computer room.

[0094] Based on this, in the event of a failure in the first data center, the failover service of the second data center is invoked to update the data storage nodes and metadata instance nodes in the second data center. The management service unit in the second data center is then invoked to allocate data processing tasks corresponding to the first data center to the second data center and update the database instance nodes there. After the update is complete, the second data center can replace the failed first data center to provide services. The data synchronization result refers to the data synchronization result obtained after synchronizing data from the first data center to the second data center, thus synchronizing data from all layers of the first data center to the second data center. Data processing tasks corresponding to the first data center are sent to the second data center, which then performs these tasks. The scheduling platform sends metadata processing tasks and execution instructions to the second data center, which then performs metadata processing and data processing. The metadata processing tasks refer to the data processing tasks corresponding to the metadata database instance nodes, and the corresponding data processing tasks are the data processing tasks corresponding to the database instance units.

[0095] In practical applications, before the failure of the first data center, metadata processing tasks are performed by the metadata instance nodes in the first data center, and data processing tasks are performed by the database instance units in the first data center. Data synchronization between the metadata instance nodes in the first data center and the metadata instance nodes in the second data center, as well as data synchronization between the database instance units in the first and second data centers, is achieved through the Paxos protocol. After the failure of the first data center, the database instance units and metadata instance nodes in the first data center cannot provide services normally. Therefore, after the database instance units and metadata instance nodes in the second data center complete their status updates, they can replace the first data center in providing services, performing metadata processing and data processing tasks. Simultaneously, the data corresponding to the database instance units in the second data center is synchronized to the database instance units in the third data center via the Paxos protocol.

[0096] Furthermore, considering that after a fault occurs in the first data center, staff will inspect and troubleshoot the fault in the first data center, and after the fault in the first data center is restored, the dispatch platform can switch the second data center back to the first data center, so that the first data center can continue to provide services, and the second data center can be restored to the state before the fault in the first data center. The specific implementation is as follows:

[0097] In the event of a failure recovery in the first data center, the system invokes the switching service of the first data center to update the data storage nodes of the database storage unit in the first data center and the metadata database instance nodes in the second data center; it also invokes the management service unit in the first data center to allocate the data processing tasks corresponding to the second data center to the first data center and update the database instance nodes in the first data center; based on the updated data storage nodes, metadata database instance nodes, and database instance nodes in the first data center, it synchronizes data with the first data center; and according to the data synchronization results, it sends the execution instructions for the metadata processing tasks and the data processing tasks to the first data center.

[0098] Based on this, once the first data center recovers from its failure and is able to provide services, the second data center can be switched back to the first data center, restoring it to its state before the first data center failed. The first data center will then take over providing services. After the first data center recovers, the switchover service of the first data center is invoked to update the data storage nodes of the database storage unit in the first data center and the metadata database instance nodes in the second data center. The management service unit in the first data center is invoked to allocate the corresponding data processing tasks from the second data center to the first data center and update the database instance nodes in the first data center. Based on the updated data storage nodes, metadata database instance nodes, and database instance nodes in the first data center, data is synchronized to the first data center. Based on the data synchronization results, execution instructions for metadata processing tasks and data processing tasks are sent to the first data center.

[0099] Following the previous example, after the fault in the first data center is recovered, the system can switch back to the first data center from the second data center. The management and control services, XDB instances, management and control metadata database, and Kubernetes base in the second data center will no longer provide services to the outside world. The second data center will be restored to the ready state, and the services of the first data center after the fault is recovered will be taken over by the second data center to provide services to the outside world.

[0100] In summary, after the first data center is restored from its fault, the second data center providing services can be switched to the first data center. That is, the first data center will replace the second data center in providing services, and the second data center will be restored to the state before the first data center failed. This will enable the switching between the first and second data centers after the first data center is restored from its fault, and allow the first data center to operate normally.

[0101] Furthermore, considering that only one of the first and second data centers needs to provide services externally, while the other is in standby mode, once the first data center recovers from a failure, the data center providing services can switch back to the first data center. Simultaneously, the second data center is set to standby mode, only synchronizing data with the first data center and not providing external services. The specific implementation is as follows:

[0102] The data storage nodes and metadata instance nodes in the second data center are restored; the management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center, and the database instance nodes in the second data center are restored.

[0103] Specifically, restoration refers to updating the service status of the data storage nodes that provide services to the outside world in the second data center, as well as the metadata instance nodes in the second data center, to the ready state, that is, restoring them to the state before the failure of the first data center. Only asynchronous data synchronization is performed between the first data center and no services are provided to the outside world.

[0104] Based on this, after the first data center is restored from failure, the service status of the second data center is updated to the state before the failure of the first data center. The data storage nodes, database instance nodes, and metadata database instance nodes in the second data center are restored. The management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center, and the database instance nodes in the second data center are restored.

[0105] Following the previous example, after the failure in the first data center is resolved, the management and control services, XDB instances, management and control metadata database, and Kubernetes base in the second data center will no longer provide services to the outside world. The XDB instances in the second data center will be updated to "fo l lower", the xdbleader will be updated to "xdb fo l lower", and the etcd leader will be updated to "etcd fo l lower". In other words, the second data center will be restored to the ready state and will be used to synchronize data from the first data center.

[0106] In summary, the disaster recovery method provided in one embodiment of this specification is applied to a scheduling platform. The scheduling platform provides a set of management services for a first data center and a second data center. In the event of a failure in the first data center, the platform invokes the switching service of the second data center to update the data storage nodes and metadata instance nodes in the second data center; it also invokes the management and control service unit in the second data center to allocate the data processing tasks corresponding to the first data center to the second data center and update the database instance nodes in the second data center; based on the updated data storage nodes, metadata instance nodes, and database instance nodes in the second data center, it synchronizes data with the second data center; and based on the data synchronization results, it sends execution instructions for metadata processing tasks and data processing tasks to the second data center.

[0107] The first and second data centers use the same set of management services provided by the scheduling platform, which reduces service deployment costs and operational complexity. In the event of a failure in the first data center, the backup and switchover service of the second data center can be directly invoked, enabling the second data center to replace the first data center and continue to provide services to the outside world. This meets the switching needs of the disaster recovery system while improving service switching speed and fault recovery efficiency.

[0108] The following is in conjunction with the appendix Figure 3 Taking the disaster recovery method provided in this manual as an example of its application in a data center within the same city, the disaster recovery method will be further explained. Figure 3 The present specification illustrates a process flowchart of a disaster recovery method provided in one embodiment, which specifically includes the following steps.

[0109] Step S302: In the event of a failure in the main data center, invoke the switchover service of the backup data center, select a data storage node from the second data storage nodes included in the backup data center, and update the data storage node.

[0110] Step S304: Call the switchover service of the standby data center, select the metadata database instance node in the second metadata database instance node in the standby data center, and update the metadata database instance node.

[0111] Step S306: Call the management and control service unit in the standby data center to allocate the data processing tasks corresponding to the main data center to the standby data center, and update the status and attribute information of the database instance nodes in the standby data center.

[0112] Step S308: Based on the updated database instance node, synchronize the instance data in the standby data center to the third database instance node in the third data center. The third database instance node is used for asynchronous storage of instance data.

[0113] Step S310: Obtain historical data corresponding to the main data center failure, call the management and control service unit in the backup data center, and update the metadata of the metadata database instance nodes in the backup data center based on the historical data within the preset recovery time.

[0114] Step S312: If the main server room is recovered from a failure, switch the backup server room to the main server room.

[0115] Figure 4 This is an architecture diagram of a disaster recovery system provided in one embodiment of this specification. The primary data center, backup data center, and third data center constitute the basic disaster recovery system. The primary and backup data centers are managed by the same Kubernetes system, and the cloud-native database system is architecturally layered, enabling the disaster recovery system to have primary / backup failover capabilities. In the event of a disaster, it can automatically switch between primary and backup data centers, or users can manually decide to temporarily suspend service to avoid data inconsistencies among some database instances. Figure 4As shown, the main data center deploys a Kubernetes base, a management metadata database, a database instance (XDB instance), a management service, and a disaster recovery failover service. Correspondingly, the backup data center deploys the same Kubernetes base, management metadata database, database instance (XDB instance), management service, and disaster recovery failover service-b. Global load balancing is achieved through Kubernetes services. The failover service-a in the main data center and the failover service-b in the backup data center can be managed by the user Actor. The third data center only deploys the database instance (XDB instance).

[0116] The Kubernetes (K8S) base in the primary data center provides K8S services. The disaster recovery system also provides internal or external Global Load Balancing (GLB). The K8S base contains three etcd nodes (K8S metadata services): one leader node and two follower nodes. The leader node provides services externally, and the two follower nodes perform data synchronization. The K8S base in the backup data center provides K8S services and contains three etcd follower nodes for asynchronous data synchronization. The management metadata database in the primary data center contains three xdb nodes: one leader node and two follower nodes. The leader node provides services externally, and the two follower nodes perform data synchronization. The management metadata database in the backup data center contains three xdb follower nodes for asynchronous data synchronization. The primary data center contains the leader database instance, while the backup and third data centers contain follower database instances. Data synchronization between nodes in the primary and backup data centers, as well as between database instances, is achieved via the Paxos protocol. Data synchronization between the backup data center and the lower database instances in the third data center is also achieved via the Paxos protocol. Both the primary and backup data centers provide highly available management and control services.

[0117] In other words, the six etcd nodes in the primary and backup data centers form a Paxos cluster. The three nodes in the primary data center act as leader, follower lower, and follower lower, ensuring single-node disaster recovery capabilities within the data center. The three etcd nodes in the backup data center act as learners, asynchronously synchronizing data from the primary data center. The management metadata database uses xdb, similar to etcd, forming a six-node Paxos cluster in both the primary and backup data centers. The three nodes in the primary data center act as leader, follower lower, and follower lower, ensuring single-node disaster recovery capabilities within the data center. The three nodes in the backup data center act as learners, asynchronously synchronizing data from the primary data center. The management service simply deploys a replica in the backup data center. Since the primary and backup data centers belong to the same Kubernetes cluster, services are provided externally through a Kubernetes service. At this point, the stateless management service is in a multi-active state. The stateful management service needs to be redesigned for distributed operation to ensure normal service even when both the primary and backup servers are simultaneously running in the service backend. Database instances are deployed in three availability zones to form a cluster. Figure 4 (Shown as a three-node XDB), through, as Figure 4 Data synchronization is performed using the Paxos protocol, a strong consistency protocol. A disaster recovery and failover service is set up independently in both the primary and backup data centers. The primary data center service has the ability to switch the base station and instances back to the primary data center; the backup data center service has the ability to switch the base station and instances back to the backup data center. In the event of a single data center failure, it has the capability for primary / backup failover.

[0118] When the main data center experiences a failure, the disaster recovery and switching service of the backup data center is invoked to promote the roles of the three etcd learner nodes in the backup data center to leader, follow lower, and follow lower, so that they can provide services to the outside world. The disaster recovery and switching service of the backup data center is invoked to promote the roles of the three xdb learner nodes of the backup data center's management metadata database to leader, follow lower, and follow lower, so that they can provide services to the outside world. Management service layer: Each management service has a replica deployed in the backup data center, adopting a multi-active architecture. When the main data center fails, all traffic is automatically routed to the backup data center. The metadata database that the management service depends on can be used after the three xdb learner nodes of the management metadata database can provide services to the outside world, so the management service in the backup data center can also run normally at this time. Kernel service layer: (1) such as Figure 4The XDB shown uses the strongly consistent Paxos protocol, which ensures that the RPO is zero even if the main data center fails, and automatically elects a master in the standby data center. At this time, both RPO and RTO are determined by the kernel. In addition, the master-slave switch of management parameters in DBaaS system is often very important. Inconsistency between management parameters and kernel state may lead to the issuance of incorrect switch task streams, causing secondary disasters. Therefore, after the management service recovers, it will correct the metadata as soon as possible. This time is determined by the RTO of the management service. Since the recovery of the management service depends on the recovery of the base management metadata database, it is ultimately determined by the RTO of the base metadata database. (2) Some database instance kernels do not have the function of automatic master switching. In this case, it is necessary to issue a switch task through the disaster recovery switch service and perform master-slave switch through the management service. At this time, the RPO and RTO of the database instance are determined by the recovery speed of the management service. Since the recovery of the management service depends on the recovery of the base management metadata database, it is ultimately determined by the RTO of the base metadata database. Whether the cloud-based etcd and management metadata database are switched, and whether the database instances that do not have automatic switching capabilities are switched, are handled uniformly by the disaster recovery switching service to ensure orderly switching. In special circumstances, users are also given the option not to switch services to the backup data center, choosing to sacrifice service continuity to ensure the consistency of system data.

[0119] In summary, by adopting a three-datacenter disaster recovery architecture, the database kernel with disaster recovery capabilities can achieve an RPO of 0, meeting financial-grade disaster recovery requirements. The third datacenter avoids deploying unnecessary foundational and management components, reducing material requirements and operational costs. A single cloud foundation manages all database instances and management components. The foundation's metadata is physically replicated, ensuring rapid synchronization between Kubernetes etcd and the management metadata database under low latency (often less than 1.5 milliseconds) across the primary and backup datacenters within the same city, minimizing data loss. Using a single cloud foundation allows database instances to be deployed across datacenters, truly achieving kernel-level physical replication without relying on external components, ensuring security and efficiency. A single cloud foundation manages management components, achieving true multi-active management and minimizing management service switching time. Managing all database instances and management components with a single cloud foundation eliminates the need to maintain two cloud foundations and management services, and avoids issues such as inconsistent metadata between primary and backup datacenters and inconsistent database instance states, reducing operational costs.

[0120] This specification presents an embodiment of a disaster recovery system that uses a single cloud infrastructure to manage the data plane and control plane services of three data centers. At the cloud infrastructure layer, kernel-level physical replication based on the Paxos protocol reduces overhead and latency, enabling high-performance disaster recovery across three data centers within the same city. At the control service layer, seamless switching between dual-data center active-active architecture and data center-level disaster recovery is achieved based on a single cloud infrastructure. At the database instance layer, the system meets the financial-grade disaster recovery requirement of zero RPO for physical replication of the database kernel. Combined with the disaster recovery capabilities of the control service, rapid switching between the kernel and control is achieved, thereby realizing disaster recovery for DBaaS services.

[0121] Corresponding to the above method embodiments, this specification also provides embodiments of disaster recovery devices. Figure 5 A schematic diagram of a disaster recovery device according to one embodiment of this specification is shown. Figure 5 As shown, the device includes:

[0122] The update module 502 is configured to call the switching service of the second data center in the event of a failure in the first data center, and update the data storage nodes and metadata instance nodes in the second data center.

[0123] The allocation module 504 is configured to call the management and control service unit in the second data center to allocate the data processing tasks corresponding to the first data center to the second data center and update the database instance nodes in the second data center;

[0124] Synchronization module 506 is configured to synchronize data to the second data center based on the updated data storage node, metadata instance node and database instance node in the second data center.

[0125] The sending module 508 is configured to send an execution instruction for the metadata processing task and the data processing task to the second computer room based on the data synchronization result.

[0126] In an optional embodiment, the update module 502 is further configured to:

[0127] The system invokes the switching service of the second data center, selects a data storage node from the second data storage nodes contained in the second data center, and updates the data storage node. The updated data storage node is used to receive storage node data submitted for the second data center. The system also invokes the switching service of the second data center, selects a metadata instance node from the second metadata instance nodes in the second data center, and updates the metadata instance node. The updated metadata instance node is used to receive metadata instance data submitted for the second data center.

[0128] In an optional embodiment, the allocation module 504 is further configured to:

[0129] The management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center, and to update the status and attribute information of the database instance nodes in the second data center. The updated database instance nodes are used to receive database instance data submitted to the second data center, and the management and control service unit is used to manage the status of the database instance nodes in the second data center.

[0130] In an optional embodiment, the synchronization module 506 is further configured to:

[0131] Determine the data storage space, read the stored data in the data storage space, and synchronize the stored data to the updated data storage node in the second data center based on the data synchronization protocol; read the metadata in the data storage space, and synchronize the metadata to the updated metadata database instance node in the second data center based on the data synchronization protocol; read the instance data in the data storage space, and synchronize the instance data to the updated database instance node in the second data center based on the data synchronization protocol.

[0132] In an optional embodiment, the allocation module 504 is further configured to:

[0133] Based on the updated database instance node, the instance data in the second data center is synchronized to the third database instance node in the third data center, and the third database instance node is used to perform asynchronous storage of the instance data.

[0134] In an optional embodiment, the allocation module 504 is further configured to:

[0135] The system invokes the switching service of the second data center to send an update task to the management service unit in the second data center; it then invokes the management service unit in the second data center that receives the update task to update the database instance nodes in the second data center; or, it receives an update instruction for the second data center and updates the database instance nodes in the second data center based on the update instruction; the updated database instance nodes in the second data center are then used to receive data operation tasks.

[0136] In an optional embodiment, the allocation module 504 is further configured to:

[0137] Obtain historical data corresponding to the fault in the first data center; invoke the management and control service unit in the second data center to update the metadata of the metadata database instance node in the second data center based on the historical data within a preset recovery time.

[0138] In an optional embodiment, the sending module 508 is further configured to:

[0139] In the event of a failure recovery in the first data center, the system invokes the switching service of the first data center to update the data storage nodes of the database storage unit in the first data center and the metadata database instance nodes in the second data center; it also invokes the management service unit in the first data center to allocate the data processing tasks corresponding to the second data center to the first data center and update the database instance nodes in the first data center; based on the updated data storage nodes, metadata database instance nodes, and database instance nodes in the first data center, it synchronizes data with the first data center; and according to the data synchronization results, it sends the execution instructions for the metadata processing tasks and the data processing tasks to the first data center.

[0140] In an optional embodiment, the sending module 508 is further configured to:

[0141] The data storage nodes and metadata instance nodes in the second data center are restored; the management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center, and the database instance nodes in the second data center are restored.

[0142] In summary, the disaster recovery device provided in one embodiment of this specification is applied to a scheduling platform. The scheduling platform provides a set of management services for a first data center and a second data center. In the event of a failure in the first data center, it invokes the switching service of the second data center to update the data storage nodes and metadata instance nodes in the second data center; it invokes the management and control service unit in the second data center to allocate the data processing tasks corresponding to the first data center to the second data center and updates the database instance nodes in the second data center; based on the updated data storage nodes, metadata instance nodes, and database instance nodes in the second data center, it synchronizes data to the second data center; and based on the data synchronization results, it sends execution instructions for metadata processing tasks and data processing tasks to the second data center.

[0143] The first and second data centers use the same set of management services provided by the scheduling platform, which reduces service deployment costs and operational complexity. In the event of a failure in the first data center, the backup and switchover service of the second data center can be directly invoked, enabling the second data center to replace the first data center and continue to provide services to the outside world. This meets the switching needs of the disaster recovery system while improving service switching speed and fault recovery efficiency.

[0144] The above is a schematic scheme of a disaster recovery device according to this embodiment. It should be noted that the technical solution of this disaster recovery device and the technical solution of the disaster recovery method described above belong to the same concept. For details not described in detail in the technical solution of the disaster recovery device, please refer to the description of the technical solution of the disaster recovery method described above.

[0145] Corresponding to the above method embodiments, this specification also provides a disaster recovery system embodiment, including: a first data center, a second data center, and a scheduling platform. The scheduling platform stores data synchronization executable instructions. When the data synchronization executable instructions are executed by the scheduling platform, they implement the steps of the above disaster recovery method, and are used to allocate the data stored in the first data center and the data processing tasks to the second data center.

[0146] The above is an illustrative scheme of a disaster recovery system according to this embodiment. It should be noted that the technical solution of this disaster recovery system and the technical solution of the disaster recovery method described above belong to the same concept. For details not described in detail in the technical solution of the disaster recovery system, please refer to the description of the technical solution of the disaster recovery method described above.

[0147] Figure 6 A structural block diagram of a computing device 600 according to one embodiment of this specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.

[0148] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. Access device 640 may include one or more of any type of wired or wireless network interface (e.g., network interface card (NIC)), such as IEEE 802.11 Wireless Local Area Network (WLAN) interface, Wi-MAX (Worldwide Interoperability for Microwave Access) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.

[0149] In one embodiment of this application, the aforementioned components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0150] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server. The processor 620 is configured to execute computer-executable instructions, which, when executed by the processor, implement the steps of the disaster recovery method described above.

[0151] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the disaster recovery method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the disaster recovery method described above.

[0152] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the disaster recovery method described above.

[0153] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the disaster recovery method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the disaster recovery method described above.

[0154] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the disaster recovery method described above.

[0155] The above is an illustrative scheme of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the disaster recovery method described above belong to the same concept. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the disaster recovery method described above.

[0156] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0157] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0158] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0159] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0160] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A disaster recovery method applied to a scheduling platform, wherein, The scheduling platform provides the same set of management services for both the first and second data centers. These management services manage database instances and control components, including: In the event of a failure in the first data center, the switching service of the second data center is invoked to update the data storage nodes and metadata instance nodes in the second data center. The management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center and update the database instance nodes in the second data center; each management and control service unit corresponding to the first data center has a corresponding replica management and control service unit in the second data center; Based on the updated data storage nodes, metadata instance nodes, and database instance nodes in the second data center, a strongly consistent data synchronization protocol is used to synchronize data to the second data center. Based on the data synchronization results, the execution instructions for the metadata processing task and the data processing task are sent to the second computer room.

2. The method according to claim 1, wherein invoking the switching service of the second data center to update the data storage nodes and metadata instance nodes in the second data center includes: The switching service of the second data center is invoked, a data storage node is selected in the second data storage node in the second data center, and the data storage node is updated. The updated data storage node is used to receive storage node data submitted for the second data center. The switching service of the second data center is invoked, a metadata instance node is selected in the second metadata instance node in the second data center, and the metadata instance node is updated. The updated metadata instance node is used to receive metadata instance data submitted for the second data center.

3. The method according to claim 1, wherein calling the management and control service unit in the second data center to allocate the data processing tasks corresponding to the first data center to the second data center and updating the database instance nodes in the second data center includes: The management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center, and to update the status and attribute information of the database instance nodes in the second data center. The updated database instance nodes are used to receive database instance data submitted to the second data center, and the management and control service unit is used to manage the status of the database instance nodes in the second data center.

4. The method according to claim 1, wherein synchronizing data to the second data center based on the updated data storage node, metadata instance node, and database instance node in the second data center includes: Determine the data storage space, read the stored data in the data storage space, and synchronize the stored data to the updated data storage node in the second data center based on the data synchronization protocol; Read the metadata in the data storage space, and synchronize the metadata to the updated metadata database instance node in the second data center based on the data synchronization protocol; Read the instance data in the data storage space, and synchronize the instance data to the updated database instance node in the second data center based on the data synchronization protocol.

5. The method according to claim 1, further comprising, after performing the step of updating the database instance node in the second data center: Based on the updated database instance node, the instance data in the second data center is synchronized to the third database instance node in the third data center, and the third database instance node is used to perform asynchronous storage of the instance data.

6. The method according to claim 1, wherein updating the database instance node in the second data center includes: Invoke the switching service of the second data center and send an update task to the management and control service unit in the second data center; The management and control service unit that received the update task in the second data center is invoked to update the database instance nodes in the second data center; or, The system receives an update instruction for the second data center and updates the database instance nodes in the second data center based on the update instruction; the updated database instance nodes in the second data center are used to receive data operation tasks.

7. The method according to claim 1, further comprising, after the step of updating the database instance node in the second data center, the method further comprising: Obtain historical data corresponding to the fault in the first computer room; The management and control service unit in the second data center is invoked to update the metadata of the metadata database instance nodes in the second data center based on the historical data within a preset recovery time.

8. The method according to claim 1, wherein sending an execution instruction for executing a metadata processing task and the data processing task to the second data center based on the data synchronization result further comprises: In the event that the first data center has recovered from a failure, the switching service of the first data center is invoked to update the data storage nodes of the database storage unit in the first data center and the metadata instance nodes in the second data center. The management and control service unit in the first data center is invoked to allocate the data processing tasks corresponding to the second data center to the first data center, and the database instance nodes in the first data center are updated. Based on the updated data storage nodes, metadata instance nodes, and database instance nodes in the first data center, synchronize data to the first data center; Based on the data synchronization results, the execution instructions for the metadata processing task and the data processing task are sent to the first computer room.

9. The method according to claim 8, wherein the step of calling the management and control service unit in the first data center to allocate the data processing task corresponding to the second data center to the first data center and updating the database instance node in the first data center further includes: The data storage nodes in the second computer room and the metadata instance nodes in the second computer room are restored. The management and control service unit in the second data center is invoked to allocate the data processing tasks corresponding to the first data center to the second data center, and the database instance nodes in the second data center are restored.

10. A disaster recovery system, comprising: The system comprises a first computer room, a second computer room, and a scheduling platform. The scheduling platform stores executable data synchronization instructions. When the data synchronization executable instructions are executed by the scheduling platform, they implement the steps of the disaster recovery method according to any one of claims 1 to 9, and are used to allocate the data stored in the first computer room and the data processing tasks to the second computer room.

11. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the disaster recovery method according to any one of claims 1 to 9.

12. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the disaster recovery method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Database disaster tolerance method and device, server and storage medium

    CN111737043A

  • Data processing system and method

    CN112417043A