Method, apparatus and device for verifying recovery point objective of dual-active disaster recovery system

By executing test scripts and arbitration nodes on the virtual machine to verify the data loss of the dual-live disaster recovery system after the failure migration, it solves the scenario coverage problem that is difficult to verify in the existing technology and achieves efficient data loss verification.

CN114546589BActive Publication Date: 2025-06-03NEW H3C BIG DATA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210147129.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-17
Publication Date
2025-06-03
Estimated Expiration
2042-02-17

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively verify the data loss caused by failures in the virtual machine's business data writing process, and the underlying storage technologies of different manufacturers are different, so the coverage scenarios are incomplete.

Method used

By executing test scripts on selected virtual machines, simulate business data writing and continuously detect virtual machine accessibility at the arbitration node. When a failure occurs and a failure is migrated, compare the last timestamp recorded in the data file with the timestamp of the last detection result before migration recorded in the accessibility detection result file to verify whether the recovery point target meets the requirements.

Benefits of technology

It realizes data loss verification of dual-live disaster recovery system in case of failure during the virtual machine writing business data, covering different underlying storage technologies scenarios, is simple and easy to use and has certain versatility, and improves testing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114546589B_ABST
    Figure CN114546589B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device and equipment for verifying the recovery point objective of a dual-active disaster recovery system, which are used to solve the technical problems of the efficiency and generality of verifying the recovery point objective (RPO) of a dual-active disaster recovery system. In the present invention, business data is simulated and recorded into a storage system through a script on a selected virtual machine (VM), the reachability of the selected VM is continuously detected at an arbitration node, then a host or site failure is simulated to enable the system to perform a failover of the selected VM, and finally, by analyzing and comparing the timestamp of the last piece of data written by the selected VM before the failure migration and the timestamp of the last reachability detection result of the selected VM recorded by the arbitration node before the migration, it is verified whether the RPO of the dual-active disaster recovery system meets the requirements. The technical solution of the present invention does not need to focus on the underlying storage technology, can cover the scenario where a failure suddenly occurs during the process of the virtual machine writing business data, is simple and easy to use, and has a certain generality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of communication and cloud computing, and particularly to a method, device and equipment for verifying the recovery point objective of a dual-active disaster recovery system. Background Art

[0002] With the popularization and development of informatization construction, the data risks and threats faced by data centers are increasing. To ensure the all-weather operation of data centers, disaster recovery systems need to focus on comprehensive business continuity. As a solution, dual-active disaster recovery in data centers means deploying two identical data centers in the same city to achieve dual-active business and data. When Data Center Site #1 fails, the business system can be quickly restored at Data Center Site #2 through technologies such as dual-active business and dual-active storage.

[0003] The recovery point objective (RPO) is one of the key indicators of disaster recovery capabilities. It refers to the maximum amount of data loss allowed in a business system during a disaster and is used to measure the data redundancy backup ability of a disaster recovery system. Dual-active disaster recovery in data centers can achieve a data loss of 0, that is, RPO = 0.

[0004] In the market, relevant manufacturers continuously launch their own dual-active disaster recovery solutions for data centers, and the underlying technologies used are not the same. For different manufacturers' dual-active disaster recovery systems, a verification solution is needed to simply and effectively verify the RPO. Especially when virtual machines in Data Center Site #1 are writing business data, due to unplanned reasons (such as earthquakes, fires, etc.), if the entire Data Center Site #1 fails or a certain physical host accidentally loses power, after the virtual machine business data is successfully migrated to Data Center Site #2, it is necessary to verify that the data volume has not been lost, that is, RPO = 0.

[0005] One current detection method is to first write a certain amount of data in the virtual machines of Data Center Site #1 and then simulate the overall failure of Data Center Site #1. After the virtual machine is restored in Data Center Site #2, check the correctness and integrity of the previously written data. It is also possible to consider pausing the logical unit LUN where the virtual machine disk is located. When Data Center Site #2 restores the LUN, verify whether the data volume has decreased.

[0006] In the above detection methods, the virtual machine simulates the failure of the data center or physical host only after writing the business data. For the situation where the virtual machine suddenly fails during the process of writing business data, this technical verification solution is powerless. The underlying storage technologies used in the dual-active solutions launched by different manufacturers are not the same, and LUN is only one of the storage technologies. Only considering the detection of the data volume of LUN, the covered scenarios are incomplete and there are certain omissions. Summary of the Invention

[0007] In view of this, the present invention provides a method, apparatus and device for verifying the recovery point objective of a dual-active disaster recovery system, which are used to solve the technical problems of the efficiency and generality of verifying the recovery point objective (RPO) of a dual-active disaster recovery system.

[0008] Based on the embodiments of the present invention, the present invention provides a method for verifying the recovery point objective of a dual-active disaster recovery system, and the method includes:

[0009] Execute a test script within a selected virtual machine (VM) at the first site, and the test script is used to periodically write simulated service data carrying timestamps into a specified data file at preset time intervals;

[0010] The arbitration node continuously and periodically detects the reachability of the selected VM, and records the detection results and detection times into a local reachability detection result file;

[0011] If the host or site where the selected VM is located fails, causing the dual-active disaster recovery system to perform a fault migration of the selected VM, then after the fault migration of the selected VM is completed, compare the last timestamp recorded in the specified data file with the timestamp of the result data of the last reachable detection result recorded in the reachability detection result file before the migration of the selected VM, and verify whether the recovery point objective (RPO) of the dual-active disaster recovery system meets the requirements according to the consistency between the two.

[0012] Further, the arbitration node realizes the reachability detection of the selected VM through a network reachability detection tool or through a detection component deployed on the selected VM.

[0013] Further, the verifying whether the recovery point objective (RPO) of the dual-active disaster recovery system meets the requirements according to the consistency between the two specifically is:

[0014] Judge whether the last timestamp recorded in the specified data file is equal to the timestamp of the result data of the last reachable detection result recorded in the reachability detection result file before the migration of the selected VM. If they are equal, it is determined that the recovery point objective (RPO) is 0.

[0015] Further, the test script is a non-automatically started script, and the test script automatically terminates running after the selected VM starts a fault migration; the specified data file is stored in the storage cluster of the dual-active disaster recovery system.

[0016] On the other hand based on the embodiments of the present invention, the present invention also provides a device for verifying the recovery point objective of a dual-active disaster recovery system, and the device includes:

[0017] A business data simulation writing module, which is used to execute a test script in a selected virtual machine (VM) at the first site, and the test script is used to periodically write simulated business data carrying timestamps into a specified data file at preset time intervals;

[0018] A reachability detection module, which is used to continuously and periodically detect the reachability of the selected VM at the arbitration node, and record the detection results and detection times into a local reachability detection result file;

[0019] An RPO verification module, which is used to, if the host or site where the selected VM is located fails, causing the active-active disaster recovery system to perform a fault migration of the selected VM, then after the fault migration of the selected VM is completed, compare the last timestamp recorded in the specified data file with the timestamp of the result data whose last detection result before the migration of the selected VM is reachable recorded in the reachability detection result file, and verify whether the recovery point objective (RPO) of the active-active disaster recovery system meets the requirements according to the consistency between the two.

[0020] Further, the reachability detection module realizes the reachability detection of the selected VM through a network reachability detection tool or through a detection component deployed on the selected VM.

[0021] Further, the RPO verification module determines whether the last timestamp recorded in the specified data file is equal to the timestamp of the result data whose last detection result before the migration of the selected VM is reachable recorded in the reachability detection result file. If the results are equal, it is determined that the recovery point objective RPO is 0.

[0022] Based on the embodiments of the present invention, the present invention also provides an electronic device, including a processor, a communication interface, a storage medium, and a communication bus. Among them, the processor, the communication interface, and the storage medium complete mutual communication through the communication bus;

[0023] The storage medium is used to store a computer program;

[0024] The processor is used to implement the method steps in the method for verifying the recovery point objective of the active-active disaster recovery system provided by the present invention when executing the computer program stored on the storage medium.

[0025] In the present invention, business data is simulated and recorded into a storage system through a script on a selected VM. The arbitration node continuously detects the reachability of the selected VM, then simulates a host or site failure to enable the system to perform a failover of the selected VM. Finally, the RPO of the active-active disaster recovery system is verified by analyzing and comparing the timestamp of the last record written by the selected VM before the failure migration and the timestamp of the last reachability detection result of the selected VM recorded by the arbitration node before the migration. The technical solution of the present invention does not need to focus on the underlying storage technology, can cover the scenario where a failure suddenly occurs during the process of the virtual machine writing business data, is simple and easy to use, and has a certain degree of generality. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments of the present invention or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings of the embodiments of the present invention.

[0027] Figure 1 Structural schematic diagram of an active-active disaster recovery system provided by an embodiment of the present invention;

[0028] Figure 2 Example of a selected VM executing a test script to write a timestamp to a specified data file in an embodiment of the present invention;

[0029] Figure 3 Example of an arbitration node continuously and periodically detecting a selected VM in an embodiment of the present invention;

[0030] Figure 4 Example of a selected VM writing a timestamp to a specified data file by a test script before a failure migration in an embodiment of the present invention;

[0031] Figure 5 Example of a reachability detection result file recorded by an arbitration node in an embodiment of the present invention;

[0032] Figure 6 Structural schematic diagram of an electronic device for implementing the method for verifying the recovery point objective of an active-active disaster recovery system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0033] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and do not limit the embodiments of the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention are also intended to include the plural forms, unless the context clearly indicates otherwise. The term "and / or" used in the present invention means any or all possible combinations including one or more of the associated listed items.

[0034] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the embodiments of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, in addition, the word "if" used may be interpreted as "when" or "while" or "in response to a determination".

[0035] The object of the present invention is to provide a method for verifying the recovery point objective (RPO) of a dual-active disaster recovery system, as well as a device and equipment for implementing the method. The basic idea of the present invention is as follows: on a selected virtual machine (VM), business data is simulated and recorded into the storage system through a script. The reachability of the selected VM is continuously detected at the arbitration node. Then, a host or site failure is simulated to enable the system to perform a failover of the selected VM. Finally, the RPO of the dual-active disaster recovery system is verified by analyzing and comparing the timestamp of the last piece of data written by the selected VM before the failure migration and the timestamp of the last reachability detection result of the selected VM recorded by the arbitration node before the migration. The technical solution of the present invention does not need to focus on the underlying storage technology, can cover the scenario where a failure suddenly occurs during the process of a virtual machine writing business data, is simple and easy to use, and has a certain degree of generality.

[0036] Figure 1 FIG. 9 is a schematic structural diagram of a dual-active disaster recovery system provided by an embodiment of the present invention. The method for verifying the recovery point objective of the dual-active disaster recovery system provided by the present invention can be applied to this dual-active disaster recovery system. The two sites in this dual-active disaster recovery system are Data Center Site #1 and Data Center Site #2 respectively. Each site runs 2 physical servers. The data disks of all servers form a distributed storage system, providing dual-active capabilities for the storage cluster. The computing cluster high availability (HA) function is enabled at both sites, providing the ability to migrate virtual machines between data center sites. When a failure occurs in a certain data center site or a certain server, the arbitration node plays a role in preventing issues such as a split brain of the management platform and the storage cluster. The arbitration node runs in the form of a physical server, and its system running time is consistent with the system time of all servers in Data Center Site #1 and Data Center Site #2. The power supplies of the arbitration node, Data Center Site #1, and Data Center Site #2 are independent of each other and do not affect each other.

[0037] To describe this solution simply and effectively, this example selects a virtual machine VM, that is, the selected VM, in Data Center Site #1, that is, the first site, as an example to describe in detail the Figure 1 specific process of verifying whether the RPO of the exemplary dual-active disaster recovery system is 0 during a simulated disaster. The steps include:

[0038] Step 1. Execute a test script within a selected VM at the first site. The test script is used to periodically write simulated service data carrying timestamps into a specified data file in the storage cluster at preset time intervals.

[0039] In this embodiment, a VM is selected in Site #1 as the test virtual machine, i.e., the selected VM. Assume the hostname of this virtual machine is centos30. Run the test script on this test virtual machine. The test script periodically writes simulated service data into a specified data file in the distributed storage cluster at preset time intervals. The preset time interval can be set according to the characteristics of service data writing. For example, for services with frequent writing, the time interval can be set shorter, such as 1 second, 10 seconds, etc. For service scenarios where service data writing is not very frequent, the time interval can be set longer, such as 1 minute, 5 minutes, etc.

[0040] Figure 2 As an example of the test script for the selected VM writing timestamps into the specified data file, the test script running on the test virtual machine centos30 obtains the current timestamp every 1 second and records the timestamp into the specified data file time - test.txt stored in the distributed storage cluster, thereby simulating the process of continuous writing of service data into the distributed storage cluster.

[0041] Step 2, the arbitration node continuously and periodically detects the reachability of the selected VM, and records the detection result and detection time into the reachability detection result file located locally at the arbitration node.

[0042] The arbitration node can use a network reachability detection tool, such as the ping program, to detect the reachability of the selected VM located at the first site; it can also achieve the reachability detection of the selected VM by deploying a detection component on the selected VM or using an existing protocol component that can implement reachability detection. When the arbitration node receives a response to the detection message from the detection component, it indicates that the selected VM is network - reachable. If there is no response after a timeout, it indicates that it is unreachable. The arbitration node records the detection result and detection time into the local reachability detection result file.

[0043] Figure 3 This is an example of the arbitration node continuously and periodically detecting the selected VM in an embodiment of the present invention. In this example, the arbitration node cvk - A23 continuously pings the selected VM by running a script and records the ping operation result information and timestamp information into the local file ping - centos30.txt.

[0044] Step 3, simulate a failure of the host or site where the selected VM is located, so that the active - active disaster recovery system performs a failure migration of the selected VM.

[0045] The method of simulating the failure of the host or site where the selected VM is located can be to unplug the power cord of the host where the selected VM is located to simulate the failure of the host physical server, or to turn off the power switch of Data Center Site #1 to simulate the failure at the entire data center level. In the above two cases, the selected VM will failover to a physical server at Data Center Site #2.

[0046] Step 4, after the failover of the selected VM is completed, compare the last timestamp recorded in the specified data file with the timestamp of the result data indicating reachability of the last detection result before the migration of the selected VM recorded in the reachability detection result file, and determine whether they are consistent. If they are consistent, it means that the RPO is 0.

[0047] The selected VM will lose connection with the arbitration node during the failover process and will resume communication with the arbitration node after the migration is successful. Figure 4 This is an example of the test script writing a timestamp to the specified data file before the failover of the selected VM in an embodiment of the present invention. The last timestamp recorded in the specified data file time - test.txt is "2020 - 03 - 10 19:22:11", and this timestamp is the timestamp written at the last moment before the failover of the selected VM. The test script automatically terminates after the selected VM starts the failover, so it will terminate writing timestamps to the specified data file. Since the test script is not an automatically started script, after the selected VM fails over to Site 2 and starts, the test script will not automatically run and continue writing timestamps.

[0048] Figure 5 This is an example of the reachability detection result file recorded by the arbitration node in an embodiment of the present invention. The arbitration node continuously and periodically detects the reachability of the selected VM. Since the arbitration node keeps performing ping operations on the selected VM, the information recorded in the ping - centos30.txt file can generally be divided into three parts:

[0049] The first part is the timestamp indicating reachability of the ping information recorded when the selected VM has not started the failover, such as Figure 5 The part above the first underline.

[0050] The second part is the detection result and timestamp information of the ping information being unreachable (Destination Host Unreachable) during the failover process of the selected VM;

[0051] The third part is the timestamp indicating reachability of the ping information recorded after the selected VM fails over successfully and resumes communication with the arbitration node, such as Figure 5 The part below the second underline and below.

[0052] Figure 5 Among them, the first underscore indicates that the selected VM becomes unreachable after the timestamp "19:22:11" because it enters the fault migration process. The last timestamp record saved before the fault migration of the selected VM is "19:22:11". As Figure 5 shown, the timestamp record of the first underscore part in the reachability detection result file recorded by the arbitration node is also "19:22:11", indicating that the business data written by the selected VM before the fault migration is not lost after the fault migration, which proves that RPO = 0.

[0053] The present invention realizes a verification implementation scheme with RPO = 0 in the data center dual-active disaster recovery system, which can cover the scenario where a fault suddenly occurs during the running of business data and does not rely on the underlying storage technology to design a targeted verification method. This scheme is simple and effective, does not require additional test tool assistance, has a certain generality, and improves the test efficiency of the data center dual-active disaster recovery system.

[0054] Figure 6 FIG. is a schematic structural diagram of an electronic device for implementing the method for verifying the recovery point objective of the dual-active disaster recovery system provided by an embodiment of the present invention. The device 600 includes: a processor 610 such as a central processing unit (CPU), a communication bus 620, a communication interface 640, and a storage medium 630. Among them, the processor 610 and the storage medium 630 can communicate with each other through the communication bus 620. The storage medium 630 stores a computer program, and when the computer program is executed by the processor 610, the functions of one or more steps in the method for verifying the recovery point objective of the dual-active disaster recovery system provided by the present invention can be realized.

[0055] Figure 6In the exemplary device, the storage medium may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Additionally, the storage medium may also be at least one storage device located far from the aforementioned processor. The processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0056] It should be recognized that the embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory memory. The method can be implemented in a computer program using standard programming techniques, including a non-transitory storage medium configured with the computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if needed, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. Additionally, for this purpose, the program can run on a dedicated integrated circuit programmed for this purpose. Moreover, the operations of the processes described in the present invention can be performed in any suitable order, unless the present invention otherwise indicates or is otherwise clearly inconsistent with the context. The processes described in the present invention (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed commonly on one or more processors, by hardware, or a combination thereof. The computer program includes multiple instructions executable by one or more processors.

[0057] Further, the method can be implemented in any type of computing platform operatively connected to a suitable one, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer and, when the storage medium or device is read by the computer, can be used to configure and operate the computer to perform the processes described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the above-described steps in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.

[0058] The above are only embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for verifying the recovery point objective of a dual-active disaster recovery system, characterized in that, the method comprises: executing a test script within a selected virtual machine VM at the first site, the test script being used to periodically write simulated service data carrying timestamps into a specified data file at preset time intervals; the arbitration node continuously and periodically detects the reachability of the selected VM, and records the detection results and detection times into a local reachability detection result file; if the host or site where the selected VM is located fails, causing the dual-active disaster recovery system to perform a fault migration of the selected VM, then after the fault migration of the selected VM is completed, compare the last timestamp recorded in the specified data file with the timestamp of the result data whose last detection result before the migration of the selected VM is reachable recorded in the reachability detection result file, and verify whether the recovery point objective RPO of the dual-active disaster recovery system meets the requirements according to the consistency between the two.

2. The method according to claim 1, characterized in that, the arbitration node realizes the reachability detection of the selected VM through a network reachability detection tool or through a detection component deployed on the selected VM.

3. The method according to claim 1, characterized in that, the verifying whether the recovery point objective RPO of the dual-active disaster recovery system meets the requirements according to the consistency between the two specifically is: judging whether the last timestamp recorded in the specified data file is equal to the timestamp of the result data whose last detection result before the migration of the selected VM is reachable recorded in the reachability detection result file, and if they are equal, determining that the recovery point objective RPO is 0.

4. The method according to claim 1, characterized in that, the test script is a non-automatically started script, and the test script automatically terminates running after the selected VM starts a fault migration; the specified data file is stored in the storage cluster of the dual-active disaster recovery system.

5. A device for verifying the recovery point objective of a dual-active disaster recovery system, characterized in that, the device comprises: a service data simulation writing module, used for executing a test script within a selected virtual machine VM at the first site, the test script being used to periodically write simulated service data carrying timestamps into a specified data file at preset time intervals; a reachability detection module, used for the arbitration node to continuously and periodically detect the reachability of the selected VM, and record the detection results and detection times into a local reachability detection result file; an RPO verification module, used for if the host or site where the selected VM is located fails, causing the dual-active disaster recovery system to perform a fault migration of the selected VM, then after the fault migration of the selected VM is completed, compare the last timestamp recorded in the specified data file with the timestamp of the result data whose last detection result before the migration of the selected VM is reachable recorded in the reachability detection result file, and verify whether the recovery point objective RPO of the dual-active disaster recovery system meets the requirements according to the consistency between the two.

6. The device according to claim 5, characterized in that, The reachability detection module implements reachability detection for the selected VM through a network reachability detection tool or through a detection component deployed on the selected VM.

7. The apparatus according to claim 5, wherein, the RPO verification module determines whether the last timestamp recorded in the specified data file is equal to the timestamp of the result data of the last reachable detection result before the selected VM migration recorded in the reachability detection result file. If the results are equal, it is determined that the recovery point objective RPO is 0.

8. The apparatus according to claim 6, wherein, the test script is a non-automatically started script, and the test script automatically terminates when the selected VM starts a fault migration; the specified data file is stored in the storage cluster of the active-active disaster recovery system.

9. An electronic device, wherein, it includes a processor, a communication interface, a storage medium, and a communication bus. Among them, the processor, the communication interface, and the storage medium complete mutual communication through the communication bus; the storage medium is used to store a computer program; the processor, when executing the computer program stored on the storage medium, implements the method steps described in any one of claims 1-4.

10. A storage medium, on which a computer program is stored, wherein, the computer program, when executed by a processor, implements the method steps described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Database management method and related equipment

    CN118069662A