Disaster recovery switching method and device, storage medium and electronic device

By identifying the synchronization mode between the primary and backup clusters of the distributed database OceanBase, and adopting an automated switching method, the inefficiency and inaccuracy of disaster recovery switching caused by manual operation in existing technologies are solved, and an efficient disaster recovery switching process is achieved.

CN115599600BActive Publication Date: 2026-05-01CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA CONSTRUCTION BANK
Filing Date
2022-10-10
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing database disaster recovery switching methods mainly rely on manual operation, resulting in low accuracy and efficiency, and are time-consuming and labor-intensive.

Method used

By identifying the synchronization mode between the primary and backup clusters of the distributed database OceanBase, especially in the remote synchronization mode, an automated preset switching method is adopted for disaster recovery switching to ensure uninterrupted business operations. This includes steps such as simulation operation, green light testing, and inspection to achieve automated disaster recovery switching.

Benefits of technology

It improves the accuracy and efficiency of disaster recovery switchover, reduces the input of manpower and material resources, and realizes a highly efficient disaster recovery switchover process without human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115599600B_ABST
    Figure CN115599600B_ABST
Patent Text Reader

Abstract

The application discloses a disaster recovery switching method and device, a storage medium and an electronic device. A synchronization mode between a master cluster and a standby cluster of a distributed database OceanBase is identified to obtain a synchronization mode identification result. When the master cluster is affected by disaster recovery and the synchronization mode identification result is a remote synchronization mode, a preset switching mode is used to perform switching operation on the master cluster affected by disaster recovery and the standby cluster, so that the service of the master cluster affected by disaster recovery is ensured not to be interrupted. Disaster recovery switching of the database of the master cluster affected by disaster recovery is not required to be performed by an artificial switching mode. Only the synchronization mode between the master cluster and the standby cluster of the distributed database OceanBase needs to be identified. Different modes are considered in all aspects to consider possible abnormal conditions in the disaster recovery switching process, so that the accuracy of disaster recovery switching is improved. In the remote synchronization mode, automatic disaster recovery switching is realized by using the preset switching mode, so that manpower and material resources are reduced, and the efficiency of disaster recovery switching is improved.
Need to check novelty before this filing date? Find Prior Art

Description

A disaster recovery switching method, apparatus, storage medium, and electronic device Technical Field

[0001] This application relates to the field of database technology, and more specifically, to a disaster recovery switching method, apparatus, storage medium, and electronic device. Background Technology

[0002] With the rapid growth of data volume, database operation and maintenance has become a top priority in Internet Technology (IT) operations and maintenance. However, due to natural disasters, equipment failures, or human factors, data loss and business interruptions may occur. Therefore, disaster recovery failover for databases is necessary.

[0003] The existing disaster recovery failover methods for various databases typically use manual failover. Manual failover requires a significant investment of time and manpower, as well as coordinating multiple operations and maintenance personnel, which leads to low accuracy and efficiency in disaster recovery failover. Summary of the Invention

[0004] In view of this, this application discloses a disaster recovery switching method, apparatus, storage medium and electronic device, aiming to improve the accuracy and efficiency of disaster recovery switching.

[0005] To achieve the above objectives, the disclosed technical solution is as follows:

[0006] The first aspect of this application discloses a disaster recovery switchover method, the method comprising:

[0007] The synchronization mode between the primary cluster and the backup cluster of the distributed database OceanBase is identified, and the synchronization mode identification result is obtained; the backup cluster is a standby cluster of the primary cluster; there are multiple backup clusters; the synchronization mode is the data transmission mode between the primary cluster and the backup cluster.

[0008] When the primary cluster is affected by disaster recovery and the synchronization mode identification result is remote synchronization mode, in the remote synchronization mode, the primary cluster affected by disaster recovery and the backup cluster are switched through a preset switching method to ensure that the services of the primary cluster affected by disaster recovery are not interrupted.

[0009] Preferably, the synchronization mode includes a remote synchronization mode or a strong synchronization mode. The process of identifying the synchronization mode between the primary and backup clusters of the distributed database OceanBase, and obtaining the synchronization mode identification result, includes:

[0010] Determine the transmission distance between the primary and backup clusters of the distributed database OceanBase;

[0011] When the transmission distance is greater than the preset transmission distance, the identification result of the synchronization mode between the primary cluster and the backup cluster of the distributed database OceanBase is determined to be the remote synchronization mode; the remote synchronization mode is a synchronization mode that is not affected by network latency and backup database persistent log time.

[0012] When the transmission distance is less than or equal to the preset transmission distance, the identification result of the synchronization mode between the primary cluster and the backup cluster of the distributed database OceanBase is determined to be a strong synchronization mode; the strong synchronization mode is a synchronization mode affected by network latency and backup database persistent log time.

[0013] Preferably, when the primary cluster is affected by disaster recovery and the synchronization mode identification result is a remote synchronization mode, in the remote synchronization mode, a switch operation is performed between the primary cluster affected by disaster recovery and the backup cluster through a preset switching method to ensure that the services of the primary cluster affected by disaster recovery are not interrupted, including:

[0014] When the primary cluster is affected by disaster recovery, and the synchronization mode identification result is the remote synchronization mode, the backup cluster is simulated by a simulated switchover method under the remote synchronization mode to obtain the simulation result; the simulation operation is used to test whether the backup cluster can normally take over the operation of all services of the primary cluster affected by disaster recovery.

[0015] When the simulation results represent the simulation results of the backup cluster normally taking over all services of the primary cluster affected by disaster recovery, the primary cluster affected by disaster recovery and the backup cluster are switched through a tangent switching method, so that the backup cluster takes over the services of the primary cluster affected by disaster recovery and ensures that the services of the primary cluster affected by disaster recovery are not interrupted; the tangent switching method is to convert the database observer cluster, the OCP cluster, and the OMS cluster of the primary cluster affected by disaster recovery into the database observer cluster, the OCP cluster, and the OMS cluster of the backup cluster, respectively.

[0016] Preferably, the process of switching the primary cluster affected by disaster recovery to the backup cluster via a tangent switching method includes:

[0017] A first application stop operation is performed on the primary database of the primary cluster affected by disaster recovery and the backup database of the backup cluster; the first application stop operation is used to ensure that no data is lost during the switchover process between the primary cluster affected by disaster recovery and the backup cluster;

[0018] Perform a first green light test operation on the primary database and the backup database after the first application stop operation; the first green light test operation is used to query whether the network of the primary database and the configuration file exist after the first application stop operation, and whether the network of the backup database and the configuration file exist.

[0019] The first check operation is performed on the primary database and the standby database that have passed the first green light test operation; the first check operation is used to check whether the database status of the primary database and the standby database is normal, whether the primary database and the standby database are both in a switchable state, and whether the primary database and the standby database have stopped the full database backup check operation; the full database backup means that all database data managed by the server is backed up within a preset time.

[0020] When both the primary database and the backup database pass the first check operation, the primary cluster affected by disaster recovery and the backup cluster are switched to each other to obtain a new primary cluster; the new primary cluster is the backup cluster before the switch; the new backup cluster is the primary cluster affected by disaster recovery before the switch.

[0021] Check the database status of the new primary cluster and the new standby cluster after the switchover;

[0022] When both the database status of the switched primary cluster and the database status of the new backup cluster are normal, the application of the new primary cluster is started.

[0023] Preferred options also include:

[0024] When the primary cluster affected by disaster recovery returns to normal, the new primary cluster is switched to the recovered primary cluster through a switchback method, so that the recovered primary cluster can resume all services.

[0025] Preferably, when the primary cluster affected by disaster recovery returns to normal, the new primary cluster is switched to the recovered primary cluster via a switchback method, so that the recovered primary cluster resumes all services, including:

[0026] When the primary cluster affected by disaster recovery returns to normal, a second application stop operation is performed on the new primary cluster and the recovered primary cluster; the second application stop operation is used to ensure that no data is lost during the switchover process between the new primary cluster and the recovered primary cluster.

[0027] A second green light test operation is performed on the database of the new primary cluster after the second application stop operation and the database of the primary cluster that has recovered to normal. The second green light test operation is used to check whether the network of the new primary cluster after the second application stop operation is unobstructed and whether the configuration file exists, and whether the network of the database of the primary cluster that has recovered to normal after the second application stop operation is unobstructed and whether the configuration file exists.

[0028] A second check operation is performed on the databases of the new primary cluster and the restored primary cluster that have passed the second green light test. The second check operation is used to check whether the database status of the new primary cluster database and the restored primary cluster database is normal, whether both the new primary cluster database and the restored primary cluster database are in a switchable state, and whether both the new primary cluster database and the restored primary cluster database have stopped the full database backup check operation. The full database backup means that all database data managed by the server is backed up within a preset time.

[0029] When both the database of the new primary cluster and the database of the recovered primary cluster pass the second check operation, the new primary cluster and the recovered primary cluster will be switched over to each other.

[0030] Check the database status of the new primary cluster after the switchover and the database status of the restored primary cluster;

[0031] When both the database of the new primary cluster after the switchover and the database of the restored primary cluster are normal, start the application on the restored primary cluster.

[0032] A second aspect of this application discloses a disaster recovery switching device, the device comprising:

[0033] An identification unit is used to identify the synchronization mode between the primary cluster and the backup cluster of the distributed database OceanBase, and obtain the synchronization mode identification result; the backup cluster is a standby cluster of the primary cluster; there are multiple backup clusters; the synchronization mode is the data transmission mode between the primary cluster and the backup cluster;

[0034] The first switching unit is used to switch the primary cluster affected by disaster recovery to the backup cluster in the remote synchronization mode when the primary cluster is affected by disaster recovery and the synchronization mode identification result is remote synchronization mode, so as to ensure that the services of the primary cluster affected by disaster recovery are not interrupted.

[0035] Preferably, the identification unit includes:

[0036] The first determination module is used to determine the transmission distance between the primary cluster and the backup cluster of the distributed database OceanBase.

[0037] The second determining module is used to determine that the identification result of the synchronization mode between the primary cluster and the backup cluster of the distributed database OceanBase is a remote synchronization mode when the transmission distance is greater than the preset transmission distance; the remote synchronization mode is a synchronization mode that is not affected by network latency and backup database persistent log time.

[0038] The third determining module is used to determine that the identification result of the synchronization mode between the primary cluster and the backup cluster of the distributed database OceanBase is a strong synchronization mode when the transmission distance is less than or equal to the preset transmission distance; the strong synchronization mode is a synchronization mode affected by network latency and backup database persistent log time.

[0039] A third aspect of this application discloses a storage medium comprising stored instructions, wherein, when the instructions are executed, the device in which the storage medium resides is controlled to perform a disaster recovery switchover method as described in any one of the first aspects.

[0040] The fourth aspect of this application discloses an electronic device including a memory and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described in any of the first aspects of the disaster recovery switching method.

[0041] As described above, this application discloses a disaster recovery switchover method, apparatus, storage medium, and electronic device. It identifies the synchronization mode between the primary and backup clusters of the distributed database OceanBase, obtaining a synchronization mode identification result. The backup clusters are standby clusters of the primary cluster, and there are multiple backup clusters. The synchronization mode is the data transmission mode between the primary and backup clusters. When the primary cluster is affected by disaster recovery, and the synchronization mode identification result is a remote synchronization mode, in this remote synchronization mode, a preset switchover method is used to switch the affected primary cluster and backup clusters, ensuring uninterrupted service for the affected primary cluster. This solution eliminates the need for manual switching of the database in the affected primary cluster. It only requires identifying the synchronization mode between the primary and backup clusters of the distributed database OceanBase. For different modes, various possible anomalies during the disaster recovery switchover process are considered, improving the accuracy of the switchover. In the remote synchronization mode, a preset switchover method enables automated disaster recovery switchover, reducing manpower and resources and improving efficiency. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0043] Figure 1 is a flowchart illustrating a disaster recovery switching method disclosed in an embodiment of this application;

[0044] Figure 2 is a diagram of the database disaster recovery switching system architecture disclosed in an embodiment of this application;

[0045] Figure 3 is a schematic diagram of the switching operation between the primary cluster and the backup cluster affected by disaster recovery through the tangent switching method disclosed in the embodiments of this application;

[0046] Figure 4 is a schematic diagram of the switching operation between the primary cluster and the backup cluster affected by disaster recovery through the back-switch method disclosed in the embodiments of this application;

[0047] Figure 5 is a schematic diagram of the structure of a disaster recovery switching device disclosed in an embodiment of this application;

[0048] Figure 6 is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. Detailed Implementation

[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0050] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0051] As the background technology shows, the disaster recovery switching methods of various existing databases usually use manual switching. Using manual switching requires a lot of time, manpower, and coordination of multiple operation and maintenance personnel, which leads to low accuracy and low efficiency in disaster recovery switching.

[0052] To address the aforementioned issues, this application discloses a disaster recovery switchover method, apparatus, storage medium, and electronic device. It determines the synchronization mode between the primary and backup clusters of the distributed database OceanBase, obtaining a synchronization mode identification result. The backup clusters are standby clusters of the primary cluster, and there are multiple backup clusters. The synchronization mode is the data transmission mode between the primary and backup clusters. When the primary cluster is affected by disaster recovery, and the synchronization mode identification result is a remote synchronization mode, under this mode, a preset switchover method is used to switch the affected primary cluster and backup clusters, ensuring uninterrupted service for the affected primary cluster. This solution eliminates the need for manual switchover of the database in the affected primary cluster. It only requires determining the synchronization mode between the primary and backup clusters of the distributed database OceanBase. For different modes, various possible anomalies during the disaster recovery switchover process are considered, improving the accuracy of the switchover. In the remote synchronization mode, a preset switchover method enables automated disaster recovery switchover, reducing manpower and resources and improving efficiency. Specific implementation methods are described in the following embodiments.

[0053] Referring to Figure 1, a disaster recovery switching method disclosed in an embodiment of this application is shown. The disaster recovery switching method mainly includes the following steps:

[0054] S101: Identify the synchronization mode between the primary and backup clusters of the distributed database OceanBase and obtain the synchronization mode identification result; the backup cluster is a standby cluster of the primary cluster; there are multiple backup clusters; the synchronization mode is the data transmission mode between the primary and backup clusters.

[0055] The backup cluster and the primary cluster are located in different locations.

[0056] The synchronization modes between the primary and backup clusters of the distributed database OceanBase include remote synchronization mode and strong synchronization mode.

[0057] To facilitate understanding of the remote synchronization mode and the strong synchronization mode, Figure 2 is used as a reference. Figure 2 shows a schematic diagram of the OceanBase distributed database disaster recovery switchover system. Figure 2 is for illustrative purposes only.

[0058] In the remote synchronization mode, as shown in Figure 2 (marked "Asynchronous Log Transmission"), each engine-layer operation (Redo) log does not need to wait for simultaneous disk persistence on both the primary and standby clusters before a transaction is considered successful. The primary cluster's Redo logs do not need to be strongly synchronized to the standby cluster. This remote synchronization mode is unaffected by network latency and the standby database's log persistence time.

[0059] In strong synchronization mode, as shown in Figure 2 under the "Synchronize Transfer Log" section, each Redo log entry is considered successful only after both the primary and standby clusters have simultaneously persisted it to disk. The primary cluster's Redo logs must be strongly synchronized to the standby cluster. This strong synchronization mode is affected by network latency and the standby database's log persistence time.

[0060] Redo logs are the content transmitted between the primary and standby databases, and synchronization between the primary and standby databases is achieved by transmitting redo logs.

[0061] In Figure 2, Region SH represents Shanghai; Region HZ represents Hangzhou.

[0062] The process of identifying the synchronization mode between the primary and backup clusters of the distributed database OceanBase and obtaining the synchronization mode identification results is shown in A1-A3.

[0063] A1: Determine the transmission distance between the primary and backup clusters of the distributed database OceanBase.

[0064] A2: When the transmission distance is greater than the preset transmission distance, the synchronization mode identification result between the primary cluster and the backup cluster of the distributed database OceanBase is determined to be the remote synchronization mode; the remote synchronization mode is a synchronization mode that is not affected by network latency and the backup database persistence log time.

[0065] In this solution, because the transmission distance is greater than the preset transmission distance, and to ensure maximum performance of the primary database, the backup cluster adopts a remote synchronization mode. This solution targets OceanBase, a domestically developed distributed emerging database. The special feature of this database is that both the primary and backup databases exist as clusters, and each cluster achieves data consistency through replicas. Compared with centralized databases such as Oracle and MySQL, the switching process is more complex and the operation steps are more complicated.

[0066] The preset transmission distance can be 500 kilometers, 1000 kilometers, etc. The preset transmission distance is determined by technical personnel based on the actual situation, and this application does not impose specific limitations.

[0067] A3: When the transmission distance is less than or equal to the preset transmission distance, the synchronization mode identification result between the primary cluster and the backup cluster of the distributed database OceanBase is determined to be strong synchronization mode; strong synchronization mode is a synchronization mode affected by network latency and backup database persistent log time.

[0068] Among them, since the strong synchronization mode is a synchronization mode affected by network latency and the persistence log time of the backup database, the disaster recovery switching efficiency between the primary cluster and the backup cluster of the distributed database OceanBase is high when the transmission distance between the primary cluster and the backup cluster is less than or equal to the preset transmission distance.

[0069] S102: When the primary cluster is affected by disaster recovery and the synchronization mode identification result is remote synchronization mode, in remote synchronization mode, the primary cluster affected by disaster recovery and the backup cluster will be switched through a preset switching method to ensure that the services of the primary cluster affected by disaster recovery are not interrupted.

[0070] In Figure 2 above, in the database disaster recovery switchover scenario of the distributed database OceanBase, the primary and standby database configuration supports one primary cluster and up to 31 standby clusters. Typically, the primary and standby clusters are managed through Structured Query Language (SQL) or Oracle Certified Professional (OCP) management components.

[0071] In real-world scenarios, the primary cluster is typically the production cluster. Its characteristics include the ability to accept read and write operations from applications and the ability to achieve strong data consistency. Its role is the PRIMATY role, meaning the database corresponding to the PRIMATY role is the primary database.

[0072] The backup cluster is usually located in a different location from the primary cluster, ensuring that the backup cluster can take over when the primary cluster suffers irreversible damage such as natural disasters. The backup cluster is a data backup of the primary cluster, used to ensure data consistency, and its role is PHYSICAL STANDBY.

[0073] The transmission service of the primary and standby clusters is through Redo logs. That is, the primary cluster will automatically transmit Redo logs to the standby cluster, thereby achieving data synchronization between the primary and standby clusters.

[0074] The default failover mode is the forward failover mode. The forward failover mode means that the standby database switches to become the primary database, and the primary database becomes the standby database.

[0075] Specifically, when the primary cluster is affected by disaster recovery and the synchronization mode identification result is remote synchronization mode, in remote synchronization mode, the primary cluster affected by disaster recovery and the backup cluster are switched through a preset switching method to ensure that the business of the primary cluster affected by disaster recovery is not interrupted, as shown in B1-B2.

[0076] B1: When the primary cluster is affected by disaster recovery and the synchronization mode identification result is remote synchronization mode, the backup cluster is simulated by means of simulated switching under remote synchronization mode to obtain simulation results; the simulation operation is used to test whether the backup cluster can normally take over the operation of all services of the primary cluster affected by disaster recovery.

[0077] The simulated switching method can be switchover, failover, etc. This solution prefers the switchover method.

[0078] This solution utilizes a switchover approach, involving planned role reshuffling to ensure the backup database is built and functioning correctly. Switchover is typically lossless, meaning no data loss occurs. The system leverages the expertise of OceanBase database specialists for automated switching, enabling rapid and accurate transitions.

[0079] By using switchover, we simulate switching to a backup cluster to determine if the backup cluster can take over the primary cluster normally.

[0080] B2: When the simulation results represent the simulation results of the standby cluster normally taking over all services of the primary cluster affected by disaster recovery, the primary cluster affected by disaster recovery and the standby cluster are switched through a tangential switching method, so that the standby cluster takes over the services of the primary cluster affected by disaster recovery, ensuring that the services of the primary cluster affected by disaster recovery are not interrupted. The tangential switching method is to switch the database observer cluster, the cloud platform (OceanBase Cloud Platform, OCP) cluster, and the OMS cluster of the primary cluster affected by disaster recovery to the database kernel observer cluster, the OCP cluster, and the data migration service (OceanBase Migration Service, OMS) cluster of the standby cluster, respectively.

[0081] The Observer is the OceanBase component. The Observer is the core of the OceanBase database, and its high availability is ensured through the Paxos distributed consensus protocol.

[0082] OCP stands for OceanBase Component. OCP is a tool for operations and maintenance personnel, providing a user-friendly, transparent management and maintenance interface for OceanBase.

[0083] OMS stands for OceanBase component. OMS provides tools for online data migration and real-time incremental data replication.

[0084] The specific tangential switching method is shown in Figure 3, which illustrates the switching between the primary cluster and the backup cluster affected by disaster recovery through the tangential switching method.

[0085] In Figure 3, a forward switchover refers to the standby database becoming the primary database, and the primary database becoming the standby database. The reverse is the fallback switchover. During a forward switchover, the observer cluster, OCP cluster, and OMS cluster need to be switched separately.

[0086] In addition, the system being switched to a customer information system is divided into a private South cluster, a private North cluster, and a public cluster. The tangent switchover process involves the separate switching of the database observer clusters, OCP, and OMS for the private South cluster, North cluster, and public cluster.

[0087] Specifically, the primary cluster and backup cluster affected by disaster recovery are switched using a tangent switching method, as shown in C1-C6.

[0088] C1: Perform a first stop application operation on the primary database of the primary cluster affected by disaster recovery and the backup database of the backup cluster; the first stop application operation is used to ensure that no data is lost during the switchover process between the primary cluster and the backup cluster affected by disaster recovery.

[0089] This solution is a planned switchover drill, so the switchover needs to be implemented while the application is stopped to ensure that no data is lost.

[0090] C2: Perform the first green light test operation on the primary database and the standby database after the first application stop operation; the first green light test operation is used to check whether the network of the primary database and the configuration file exist after the first application stop operation, and whether the network of the standby database and the configuration file exist after the first application stop operation.

[0091] As shown in Figure 3, the first green light test operation includes green light tests for the private OB South cluster, green light tests for the private OB North cluster, and green light tests for the public OB cluster.

[0092] The first green light test includes checking network connectivity and the existence of configuration files. Only if the first green light test is successful can the next step be taken.

[0093] C3: Perform the first check operation on the primary and standby databases that have passed the first green light test operation; the first check operation is used to check whether the database status of the primary and standby databases is normal, whether the primary and standby databases are both in a switchable state, and whether the primary and standby databases have stopped the full database backup check operation; the full database backup indicates that all database data managed by the server is backed up within a preset time.

[0094] The preset time can be 5 minutes, 30 minutes, etc. The specific preset time is determined by technical personnel based on the actual situation, and this application does not impose any specific limitations.

[0095] As shown in Figure 3, the first inspection operation includes inspection of the private OB South cluster, inspection of the private OB North cluster, and inspection of the public OB cluster.

[0096] A full database backup refers to a full backup of the distributed storage service (Cloud Object Storage, COS).

[0097] Before disaster recovery failover, the first check operation (check, correct, and recheck) is performed. The first check operation includes checking whether the primary and standby databases are in normal status (including cluster roles, status, protection mode, and protection level, etc.), checking whether the primary and standby databases are in a failover state, and checking whether the primary database has stopped COS full backup.

[0098] If the COS full backup was not stopped, this step will correct it and automatically stop the COS full backup. Once everything is ready, the database will re-check a series of pre-switch checks until all checks are error-free and everything is in order before proceeding to the next step. This step is one of the highlights of this solution. Traditional disaster recovery switchovers require manual checks, corrections, and rebuilds before the switchover, which is time-consuming and labor-intensive. The one-click switchover solution can essentially achieve zero manual intervention, autonomously checking the pre-switch status.

[0099] The specific process for performing a self-check of the state before switching is as follows:

[0100] First, check if the cluster status is active.

[0101] If the status is active, query whether the cluster's start time (start_time) is normal;

[0102] If start_time is normal, then query whether the cluster's end time (stop_time) is normal.

[0103] If stop_time is normal, check if the database's full backup status is in the "full backup stopped" state.

[0104] If the database's full backup status is in the "full backup stopped" state, then execute C4. If the database's full backup status is in the "full backup not stopped" state, then execute the correction, automatically stop the COS full backup, and return to execute the step of checking whether the database's full backup status is in the "full backup stopped" state.

[0105] C4: When both the primary and backup databases pass the first check operation, the primary cluster affected by disaster recovery and the backup cluster will be switched to each other to obtain a new primary cluster; the new primary cluster is the backup cluster before the switch; the new backup cluster is the primary cluster affected by disaster recovery before the switch.

[0106] C5: Check the database status of the new primary cluster and the new standby cluster after the switch.

[0107] The process includes checking the database status of the new primary and backup clusters after the switchover. It confirms that the database status of the new primary and backup clusters is normal after the switchover.

[0108] C6: When the database status of both the new primary cluster and the new backup cluster are normal, start the application on the new primary cluster.

[0109] Once the database is functioning correctly, start the application. If the transactions are running normally on the new primary cluster, the switchover is successful.

[0110] The scripts for C1-C6 above are all implemented using shell programming. They are used to create a user interface through front-end and back-end interaction, providing a user-friendly interface for users.

[0111] When the primary cluster affected by disaster recovery returns to normal, the new primary cluster is switched to the recovered primary cluster through a switchback method, so that the recovered primary cluster can resume all services.

[0112] Among them, the switchback method refers to the primary database switching to the backup database, and the backup database switching to the primary database.

[0113] Specifically, when the primary cluster affected by disaster recovery returns to normal, the new primary cluster is switched to the restored primary cluster through a switchback method, so that the restored primary cluster can resume all services. The process is shown in D1-D6 and illustrated in Figure 4. Figure 4 shows a schematic diagram of switching the primary cluster affected by disaster recovery to the backup cluster through a switchback method.

[0114] D1: When the primary cluster affected by disaster recovery returns to normal, a second application stop operation is performed on both the new primary cluster and the recovered primary cluster. The second application stop operation is used to ensure that no data is lost during the switchover process between the new primary cluster and the recovered primary cluster.

[0115] D2: Perform a second green light test on the database of the new primary cluster after the second application stop operation and the database of the primary cluster that has recovered to normal. The second green light test is used to check whether the network of the new primary cluster after the second application stop operation is unobstructed and whether the configuration file exists, and whether the network of the database of the primary cluster that has recovered to normal after the second application stop operation is unobstructed and whether the configuration file exists.

[0116] As shown in Figure 4, the second green light test operation includes green light tests for the private OB South cluster, green light tests for the private OB North cluster, and green light tests for the public OB cluster.

[0117] The second green light test includes checking network connectivity and the existence of configuration files. Only after the second green light test is successful can D3 proceed.

[0118] D3: Perform a second check operation on the databases of the new primary cluster and the restored primary cluster that have passed the second green light test. The second check operation is used to check whether the database status of the new primary cluster and the restored primary cluster is normal, whether the databases of the new primary cluster and the restored primary cluster are both in a switchable state, and whether the database backup check operation of the new primary cluster and the restored primary cluster has been stopped. The database backup indicates that all database data managed by the server is backed up within a preset time.

[0119] The preset time can be 7 minutes, 25 minutes, etc. The specific preset time is determined by technical personnel based on the actual situation, and this application does not impose any specific limitations.

[0120] As shown in Figure 4, the second inspection operation includes inspection of the private OB South cluster, inspection of the private OB North cluster, and inspection of the public OB cluster.

[0121] D4: When both the database of the new primary cluster and the database of the recovered primary cluster pass the second check operation, the new primary cluster and the recovered primary cluster will be switched over to each other.

[0122] D5: Check the database status of the new master cluster after the switchover and the database status of the master cluster after it has returned to normal.

[0123] D6: When both the database of the new primary cluster after the switch and the database of the restored primary cluster are normal, start the application of the restored primary cluster.

[0124] This solution aims to achieve disaster recovery failover of the OceanBase database by replacing manual operation with automated failover. OceanBase's primary-standby high-availability architecture is a crucial supplement to its high availability capabilities. When the primary cluster becomes unavailable, the standby cluster can take over, achieving a maximum data loss probability (RPO) of 0, i.e., lossless failover disaster recovery capability. Here, RPO is the maximum data loss that can be measured in the event of a disaster.

[0125] Through testing, the practicality, standardization, and universality of this solution have been fully verified. Practicality is reflected in the smooth implementation of the forward and reverse switching process for the OceanBase database, reducing manpower and material resources. It considers various potential anomalies during the process, automating verification for each, and efficiently and practically achieves disaster recovery switching. Standardization is reflected in the streamlined, visualized, and automated process, accelerating the switching process. Universality is reflected in the fact that the concepts in this solution can be applied to switching processes for various databases, such as Oracle databases and relational distributed databases (GoldendB), requiring only the switching entity to be modified to the responding database.

[0126] In this embodiment, there is no need for manual switching of the primary cluster's database affected by disaster recovery. It only requires determining the synchronization mode between the primary and backup clusters of the distributed database OceanBase. For different modes, various possible anomalies during the disaster recovery switchover process are considered to improve accuracy. In the remote synchronization mode, automated disaster recovery switchover is achieved through preset switching methods, reducing manpower and material resources and improving efficiency.

[0127] Based on the disaster recovery switching method disclosed in FIG1 of the above embodiment, this application embodiment also discloses a disaster recovery switching device, as shown in FIG5. The disaster recovery switching device includes an identification unit 501 and a first switching unit 502.

[0128] The identification unit 501 is used to identify the synchronization mode between the primary cluster and the backup cluster of the distributed database OceanBase, and obtain the synchronization mode identification result; the backup cluster is the standby cluster of the primary cluster; there are multiple backup clusters; the synchronization mode is the data transmission mode between the primary cluster and the backup cluster.

[0129] The first switching unit 502 is used to switch the primary cluster affected by disaster recovery to the backup cluster in the remote synchronization mode when the primary cluster is affected by disaster recovery and the synchronization mode identification result is remote synchronization mode, so as to ensure that the services of the primary cluster affected by disaster recovery are not interrupted.

[0130] Furthermore, the identification unit 501 includes a first determining module, a second determining module, and a third determining module.

[0131] The first determination module is used to determine the transmission distance between the primary cluster and the backup cluster of the distributed database OceanBase.

[0132] The second determining module is used to determine that the synchronization mode identification result between the primary cluster and the backup cluster of the distributed database OceanBase is the remote synchronization mode when the transmission distance is greater than the preset transmission distance; the remote synchronization mode is a synchronization mode that is not affected by network latency and the persistence log time of the backup database.

[0133] The third determining module is used to determine the synchronization mode identification result between the primary cluster and the backup cluster of the distributed database OceanBase as strong synchronization mode when the transmission distance is less than or equal to the preset transmission distance; strong synchronization mode is a synchronization mode affected by network latency and backup database persistent log time.

[0134] Furthermore, the first switching unit 502 includes an operation module and a switching module.

[0135] The operation module is used to simulate operations on the standby cluster in the remote synchronization mode when the primary cluster is affected by disaster recovery and the synchronization mode identification result is remote synchronization mode. The simulation operation is used to test whether the standby cluster can normally take over the operation of all services of the primary cluster affected by disaster recovery.

[0136] The switching module is used to switch the primary cluster affected by the disaster recovery to the backup cluster when the simulation results indicate that the backup cluster has taken over all services of the primary cluster affected by the disaster recovery. This is done through a tangent switching method, allowing the backup cluster to take over the services of the primary cluster affected by the disaster recovery and ensuring that the services of the primary cluster affected by the disaster recovery are not interrupted. The tangent switching method is to convert the database observer cluster, OCP cluster, and OMS cluster of the primary cluster affected by the disaster recovery into the database observer cluster, OCP cluster, and OMS cluster of the backup cluster, respectively.

[0137] Furthermore, the switching module includes a first operation submodule, a second operation submodule, a third operation submodule, a fourth operation submodule, a first check submodule, and a first startup submodule.

[0138] The first operation submodule is used to perform a first stop application operation on the primary database of the primary cluster affected by disaster recovery and the backup database of the backup cluster. The first stop application operation is used to ensure that no data is lost during the switchover process between the primary cluster and the backup cluster affected by disaster recovery.

[0139] The second operation submodule is used to perform a first green light test operation on the primary database and the backup database after the first application stop operation. The first green light test operation is used to query whether the network of the primary database and the configuration file exist after the first application stop operation, and whether the network of the backup database and the configuration file exist after the first application stop operation.

[0140] The third operation submodule is used to perform a first check operation on the primary database and the standby database that have passed the first green light test operation. The first check operation is used to check whether the database status of the primary database and the standby database is normal, whether the primary database and the standby database are both in a switchable state, and whether the primary database and the standby database have stopped the full database backup check operation. The full database backup indicates that all database data managed by the server is backed up within a preset time.

[0141] The fourth operation submodule is used to perform a mutual switch between the primary cluster affected by disaster recovery and the backup cluster when both the primary database and the backup database pass the first check operation, resulting in a new primary cluster; the new primary cluster is the backup cluster before the mutual switch; the new backup cluster is the primary cluster affected by disaster recovery before the mutual switch.

[0142] The first inspection submodule is used to check the database status of the new primary cluster and the new backup cluster after the switch.

[0143] The first startup submodule is used to start the application on the new primary cluster when both the database status of the switched primary cluster and the database status of the new backup cluster are normal.

[0144] Furthermore, the disaster recovery switching device also includes a second switching unit.

[0145] The second switching unit is used to switch the new primary cluster to the restored primary cluster when the primary cluster affected by disaster recovery returns to normal, so that the restored primary cluster can resume all services.

[0146] Furthermore, the second switching unit includes a fifth operation submodule, a sixth operation submodule, a seventh operation submodule, an eighth operation submodule, a second check submodule, and a second startup module.

[0147] The fifth operation submodule is used to perform a second stop application operation on the new primary cluster and the restored primary cluster when the primary cluster affected by disaster recovery returns to normal. The second stop application operation is used to ensure that no data is lost during the switchover process between the new primary cluster and the restored primary cluster.

[0148] The sixth operation submodule is used to perform a second green light test operation on the database of the new primary cluster after the second application stop operation and the database of the primary cluster that has recovered to normal. The second green light test operation is used to query whether the network of the new primary cluster after the second application stop operation is unobstructed and whether the configuration file exists, and whether the network of the database of the primary cluster that has recovered to normal after the second application stop operation is unobstructed and whether the configuration file exists.

[0149] The seventh operation submodule is used to perform a second check operation on the databases of the new primary cluster and the restored primary cluster that have passed the second green light test operation. The second check operation is used to check whether the database status of the new primary cluster database and the restored primary cluster database is normal, whether the databases of the new primary cluster database and the restored primary cluster database are both in a switchable state, and whether the database backup check operation of the new primary cluster database and the restored primary cluster database has been stopped. The database backup indicates that all database data managed by the server is backed up within a preset time.

[0150] The eighth operation submodule is used to switch the new primary cluster and the recovered primary cluster to each other when both the database of the new primary cluster and the database of the recovered primary cluster pass the second check operation.

[0151] The second inspection submodule is used to check the database status of the new master cluster after the switch and the master cluster after it has recovered to normal.

[0152] The second startup submodule is used to start the application of the restored primary cluster when both the database of the new primary cluster after the switch and the database of the restored primary cluster are normal.

[0153] In this embodiment, there is no need for manual switching of the primary cluster's database affected by disaster recovery. It only requires determining the synchronization mode between the primary and backup clusters of the distributed database OceanBase. For different modes, various possible anomalies during the disaster recovery switchover process are considered to improve accuracy. In the remote synchronization mode, automated disaster recovery switchover is achieved through preset switching methods, reducing manpower and material resources and improving efficiency.

[0154] This application embodiment also provides a storage medium, which includes stored instructions, wherein, when the instructions are executed, the device where the storage medium is located is controlled to perform the above-described disaster recovery switching method.

[0155] This application also provides an electronic device, the structural schematic of which is shown in FIG6. Specifically, it includes a memory 601 and one or more instructions 602. One or more instructions 602 are stored in the memory 601 and are configured to be executed by one or more processors 603 to perform the above-mentioned disaster recovery switching method.

[0156] The specific implementation processes and derivative methods of the above embodiments are all within the protection scope of this application.

[0157] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0158] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0159] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0160] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A disaster recovery switchover method, characterized in that, The method includes: identifying the synchronization mode between the primary cluster and backup cluster of the distributed database OceanBase, and obtaining a synchronization mode identification result; the backup cluster is a standby cluster of the primary cluster; there are multiple backup clusters; the synchronization mode is the data transmission mode between the primary cluster and the backup cluster; when the primary cluster is affected by disaster recovery and the synchronization mode identification result is a remote synchronization mode, in the remote synchronization mode, a switch operation is performed between the primary cluster affected by disaster recovery and the backup cluster through a preset switching method to ensure that the services of the primary cluster affected by disaster recovery are not interrupted; when the primary cluster is affected by disaster recovery and the synchronization mode identification result is a remote synchronization mode, in the remote synchronization mode, a switch operation is performed between the primary cluster affected by disaster recovery and the backup cluster through a preset switching method to ensure that the services of the primary cluster affected by disaster recovery are not interrupted; Through a preset switching method, the primary cluster affected by disaster recovery and the backup cluster are switched over to ensure that the services of the primary cluster affected by disaster recovery are not interrupted. This includes: when the primary cluster is affected by disaster recovery and the synchronization mode identification result is the remote synchronization mode, the backup cluster is simulated under the remote synchronization mode through a simulated switching method to obtain simulation results; the simulation operation is used to test whether the backup cluster can normally take over all services of the primary cluster affected by disaster recovery; when the simulation result indicates that the backup cluster can normally take over all services of the primary cluster affected by disaster recovery, the primary cluster affected by disaster recovery and the backup cluster are switched over through a tangent switching method, so that the backup cluster can take over the services of the primary cluster affected by disaster recovery. To ensure uninterrupted service on the primary cluster affected by disaster recovery, the handover method involves converting the database observer cluster, OCP cluster, and OMS cluster of the primary cluster affected by disaster recovery into the database observer cluster, OCP cluster, and OMS cluster of the backup cluster, respectively. The process of switching the primary cluster affected by disaster recovery to the backup cluster via the handover method includes: performing a first application stop operation on the primary database of the primary cluster affected by disaster recovery and the backup database of the backup cluster; the first application stop operation is used to ensure uninterrupted service on the primary cluster affected by disaster recovery and the backup cluster. The process of switching the backup cluster ensures no data loss; a first green light test is performed on the primary database and the backup database after the first application stop operation; the first green light test is used to check whether the network of the primary database and the configuration file exist after the first application stop operation, and whether the network of the backup database and the configuration file exist after the first application stop operation are accessible; a first check operation is performed on the primary database and the backup database that pass the first green light test; the first check operation is used to check whether the database status of the primary database and the backup database is normal, whether both the primary database and the backup database are in a switchable state, and whether both the primary database and the backup database have stopped the full database backup check operation.The database full backup process involves backing up all database data managed by the server within a preset time period. When both the primary and backup databases pass the first check operation, the primary cluster affected by the disaster recovery is switched to the backup cluster to obtain a new primary cluster. The new primary cluster is the backup cluster before the switchover. The new backup cluster is the primary cluster affected by the disaster recovery before the switchover. The database status of the new primary cluster and the new backup cluster after the switchover is checked. When the database status of both the new primary cluster and the new backup cluster is normal, the application on the new primary cluster is started.

2. The method according to claim 1, characterized in that, The synchronization mode includes a remote synchronization mode or a strong synchronization mode. Identifying the synchronization mode between the primary and backup clusters of the OceanBase distributed database and obtaining the synchronization mode identification result includes: determining the transmission distance between the primary and backup clusters of the OceanBase distributed database; when the transmission distance is greater than a preset transmission distance, determining the synchronization mode identification result between the primary and backup clusters of the OceanBase distributed database as a remote synchronization mode; the remote synchronization mode is a synchronization mode unaffected by network latency and backup database persistence log time; when the transmission distance is less than or equal to the preset transmission distance, determining the synchronization mode identification result between the primary and backup clusters of the OceanBase distributed database as a strong synchronization mode; the strong synchronization mode is a synchronization mode affected by network latency and backup database persistence log time.

3. The method according to claim 1, characterized in that, Also includes: When the primary cluster affected by disaster recovery returns to normal, the new primary cluster is switched to the recovered primary cluster through a switchback method, so that the recovered primary cluster can resume all services.

4. The method according to claim 3, characterized in that, When the primary cluster affected by disaster recovery recovers to normal, a switchback is performed to connect the new primary cluster to the recovered primary cluster, enabling the recovered primary cluster to resume all services. This includes: performing a second application stop operation on both the new primary cluster and the recovered primary cluster when the primary cluster affected by disaster recovery recovers to normal; the second application stop operation ensures no data loss during the switchover process; performing a second green light test operation on the databases of the new primary cluster and the recovered primary cluster after the second application stop operation; the second green light test operation checks the network connectivity and configuration file existence of the new primary cluster and the database of the recovered primary cluster after the second application stop operation; and checking the network connectivity and configuration file existence of the database of the new primary cluster after the second application stop operation. The system performs a second check on the databases of both the new and restored primary clusters. This second check verifies whether the database status of both the new and restored primary clusters is normal, whether both are in a switchable state, and whether both have stopped the full database backup check. The full database backup indicates that all database data managed by the server is backed up within a preset time period. When both the new and restored primary clusters pass the second check, the new and restored primary clusters are switched over. The status of the databases of both the new and restored primary clusters is then checked. If both are normal, the application on the restored primary cluster is started.

5. A disaster recovery switching device, characterized in that, The device includes: an identification unit, used to identify the synchronization mode between the primary cluster and the backup cluster of the distributed database OceanBase, and obtain a synchronization mode identification result; the backup cluster is a standby cluster of the primary cluster; there are multiple backup clusters; the synchronization mode is the data transmission mode between the primary cluster and the backup cluster; a first switching unit, used to, when the primary cluster is affected by disaster recovery and the synchronization mode identification result is a remote synchronization mode, perform a switching operation between the primary cluster affected by disaster recovery and the backup cluster in the remote synchronization mode through a preset switching method, to ensure that the services of the primary cluster affected by disaster recovery are not interrupted; the first switching unit includes: an operation module, used to, when the primary cluster is affected by disaster recovery... When disaster recovery is affected and the synchronization mode identification result is the remote synchronization mode, the backup cluster is simulated using a simulated switching method under the remote synchronization mode to obtain simulation results. The simulation operation is used to test whether the backup cluster can normally take over all services of the primary cluster affected by disaster recovery. The switching module is used to switch the primary cluster affected by disaster recovery to the backup cluster using a tangential switching method when the simulation result indicates that the backup cluster can normally take over all services of the primary cluster affected by disaster recovery. This ensures that the backup cluster takes over the services of the primary cluster affected by disaster recovery and that the services of the primary cluster affected by disaster recovery are not interrupted. The tangential switching method is used to switch the primary cluster affected by disaster recovery to the backup cluster using a tangential switching method. The switching module converts the primary cluster's database observer cluster, the primary cluster's OCP cluster affected by disaster recovery, and the primary cluster's OMS cluster affected by disaster recovery into the backup cluster's database observer cluster, backup cluster's OCP cluster, and backup cluster's OMS cluster. The switching module includes: a first operation submodule, used to perform a first application stop operation on the primary database of the primary cluster affected by disaster recovery and the backup database of the backup cluster; the first application stop operation is used to ensure that no data is lost during the switching process between the primary cluster affected by disaster recovery and the backup cluster; a second operation submodule processes the primary database and the backup database after the first application stop operation. The system performs a first green light test operation. This first green light test operation checks whether the network of the primary database and the configuration file exist after the first application stop operation, and also checks whether the network of the backup database and the configuration file exist after the first application stop operation. The third operation submodule performs a first check operation on the primary and backup databases that have passed the first green light test operation. This first check operation checks whether the database status of the primary and backup databases is normal, whether both the primary and backup databases are in a switchable state, and whether both the primary and backup databases have stopped the full database backup check operation. The full database backup indicates that all database data managed by the server is backed up within a preset time period.The fourth operation submodule is used to switch the primary cluster affected by disaster recovery to the backup cluster when both the primary and backup databases pass the first check operation, resulting in a new primary cluster; the new primary cluster is the backup cluster before the switch; the new backup cluster is the primary cluster affected by disaster recovery before the switch. The first check submodule is used to check the database status of the new primary cluster and the new backup cluster after the switch. The first startup submodule is used to start the application on the new primary cluster when both the database status of the switched primary cluster and the database status of the new backup cluster are normal.

6. The apparatus according to claim 5, characterized in that, The identification unit includes: a first determining module, used to determine the transmission distance between the primary cluster and the backup cluster of the distributed database OceanBase; a second determining module, used to determine that when the transmission distance is greater than a preset transmission distance, the identification result of the synchronization mode between the primary cluster and the backup cluster of the distributed database OceanBase is a remote synchronization mode; the remote synchronization mode is a synchronization mode unaffected by network latency and backup database persistence log time; and a third determining module, used to determine that when the transmission distance is less than or equal to the preset transmission distance, the identification result of the synchronization mode between the primary cluster and the backup cluster of the distributed database OceanBase is a strong synchronization mode; the strong synchronization mode is a synchronization mode affected by network latency and backup database persistence log time.

7. A storage medium, characterized in that, The storage medium includes stored instructions, wherein, when the instructions are executed, the device containing the storage medium is controlled to perform the disaster recovery switchover method as described in any one of claims 1 to 4.

8. An electronic device, characterized in that, It includes a memory, and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Hybrid cloud disaster recovery system and control method thereof

    CN111741135A

  • Access flow forwarding method, cluster management method and related devices

    CN112565327A