Double-site disaster recovery system and method, electronic equipment and readable storage medium
By using a dual-site disaster recovery system, incremental data is monitored and stored at the production site to the target storage pool, and then transmitted to the disaster recovery site. By combining backup storage pools and shared memory to optimize data transmission, the problems of long recovery time and high cost in traditional disaster recovery solutions are solved, and rapid data recovery and efficient transmission are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JINAN INSPUR DATA TECH CO LTD
- Filing Date
- 2026-01-13
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional single-site storage solutions are at risk of data loss. Current mainstream disaster recovery solutions have long recovery times, high storage costs, poor data consistency, and large network latency.
A dual-site disaster recovery system is adopted. By monitoring the incremental data generated by virtual machines at the production site and storing it in the target storage pool, the data is then sent to the placeholder virtual machine at the disaster recovery site. The backup storage pool and shared memory are used to optimize data transmission, thereby achieving rapid data recovery and efficient transmission.
It enables rapid data recovery, reduces storage resource consumption and costs, ensures business continuity, and improves the efficiency and reliability of data transmission.
Smart Images

Figure CN122019265A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data transmission technology, and in particular to a dual-site disaster recovery system, method, electronic device, and readable storage medium. Background Technology
[0002] With the rapid development of information technology and the continuous growth of data volume, traditional single-site storage solutions are at risk of data loss. Dual-site disaster recovery solutions achieve real-time data backup and recovery by establishing two sites in two different geographical locations. Current mainstream disaster recovery technologies include tape backup, disk backup, and cloud backup. However, these technologies suffer from long recovery times and high storage costs. Summary of the Invention
[0003] This application provides a dual-site disaster recovery system, method, electronic device, and readable storage medium to at least address the problems of long data recovery time and high storage cost in mainstream disaster recovery solutions in related technologies.
[0004] This application provides a dual-site disaster recovery system, which includes: Production site and disaster recovery site. The production site contains the target storage pool and multiple hosts, each host contains virtual machines, and the disaster recovery site contains placeholder virtual machines. The production site is used to monitor incremental data generated by virtual machines within the host and store the incremental data to the target storage pool. The production site is used to retrieve incremental data from the target storage pool and send it to the placeholder virtual machine corresponding to the virtual machine within the disaster recovery site.
[0005] This application provides a dual-site disaster recovery method applied to a dual-site disaster recovery system, the method comprising: Monitor incremental data generated by virtual machines within the host and store the incremental data to the target storage pool; The incremental data is retrieved from the target storage pool and sent to the placeholder virtual machine corresponding to the virtual machine within the disaster recovery site.
[0006] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described dual-site disaster recovery methods.
[0007] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described dual-site disaster recovery methods.
[0008] In the embodiments of this application, by monitoring incremental data generated by virtual machines within the host at the production site and storing the incremental data in a target storage pool, the production site can retrieve the incremental data from the target storage pool and send it to a placeholder virtual machine corresponding to the virtual machine at the disaster recovery site. Therefore, by monitoring incremental data generated by virtual machines within the host at the production site and storing it in the target storage pool, and then retrieving the incremental data from the storage pool and sending it to the placeholder virtual machine at the disaster recovery site, when the production site host fails, the placeholder virtual machine can use the stored incremental data for fault recovery of the production site host. This solves the technical problems of long data recovery time and high storage costs in related technologies, achieving rapid data recovery to ensure business continuity and significantly reducing storage resource consumption and costs. Attached Figure Description
[0009] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A schematic diagram of the structure of a dual-site disaster recovery system provided in this application embodiment; Figure 2 A schematic diagram of the structure of another dual-site disaster recovery system provided in an embodiment of this application; Figure 3 A schematic diagram illustrating the process of a dual-site disaster recovery method provided in an embodiment of this application; Figure 4 A schematic diagram illustrating the process of yet another dual-site disaster recovery method provided in this application embodiment; Figure 5 A schematic diagram illustrating the process of a dual-site disaster recovery scheme based on a backup pool, provided in an embodiment of this application; Figure 6 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0012] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0013] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0014] With the rapid development of information technology and the continuous growth of data volume, data security has become increasingly important. Traditional single-site storage solutions can no longer meet current needs because data faces the risk of loss in the event of a disaster. Dual-site disaster recovery solutions achieve real-time data backup and recovery by establishing two sites in two different geographical locations: one for production and the other for disaster recovery. However, current mainstream disaster recovery solutions place high demands on storage, including storage capacity, data transfer speed, and data recovery time. At the same time, dual-site disaster recovery also faces some challenges, such as weak data consistency, high network latency, and high storage costs. To address these issues, a new disaster recovery solution is needed to meet current requirements.
[0015] In the current technological context, the concept of backup storage pools has been widely adopted. A backup storage pool is a technology that combines multiple storage devices together to provide a unified storage space.
[0016] This technology enables the expansion of storage capacity, the improvement of storage performance, and the reduction of storage costs. Dual-site disaster recovery solutions based on backup storage pools are currently a popular research area. This solution achieves real-time data backup and recovery by establishing a backup storage pool between two sites.
[0017] Backup storage pools can store incremental data from the production site and then send it to the disaster recovery site. Dual-site disaster recovery solutions are particularly important in the financial sector, where data security and reliability are paramount, directly impacting the reputation and interests of financial institutions. Therefore, an efficient, reliable, and secure disaster recovery solution is needed to meet the demands of the financial industry.
[0018] Current mainstream disaster recovery solutions include tape backup, disk backup, and cloud backup. However, these solutions all have their own drawbacks, such as long recovery times for tape backup, high storage costs for disk backup, and low security for cloud backup. A dual-site disaster recovery solution based on backup storage pools can solve these problems, providing an efficient, reliable, and secure solution. At the same time, a dual-site disaster recovery solution also needs to consider issues such as data consistency and network latency. Data consistency refers to the need for data to remain consistent between the two sites, while network latency refers to the delay time during data transmission. A dual-site disaster recovery solution based on backup storage pools can solve these problems by enabling real-time data backup and recovery.
[0019] In summary, dual-site disaster recovery solutions based on backup storage pools are a popular research area. This approach offers an efficient, reliable, and secure disaster recovery solution that meets the needs of the financial sector.
[0020] To address the aforementioned problems, embodiments of this application provide a dual-site disaster recovery system, such as... Figure 1 As shown, Figure 1 The present application provides a dual-site disaster recovery system, the specific contents of which are as follows: The dual-site disaster recovery system includes a production site and a disaster recovery site. The production site contains a target storage pool and multiple hosts, each host contains a virtual machine, and the disaster recovery site contains a placeholder virtual machine.
[0021] The production site is used to monitor incremental data generated by virtual machines within the host and store the incremental data to the target storage pool.
[0022] Specifically, the production site is responsible for monitoring the data generated by the virtual machines within the production site host. Since the virtual machines on the production site continuously generate incremental data while providing services, the production site starts a monitoring process for each virtual disk of the working virtual machines. This monitoring process intercepts the incremental data from the virtual disks and redirects it to a target storage pool. This target storage pool is used to store the incremental data from the production site.
[0023] The production site is used to retrieve incremental data from the target storage pool and send it to the placeholder virtual machine corresponding to the virtual machine within the disaster recovery site.
[0024] Specifically, the production site reads incremental data from the backup storage pool for virtual machines and transmits it to the disaster recovery site via the disaster recovery network. The incremental data is transmitted in the form of data packets, each containing a header and data blocks. The header contains metadata information for the data blocks, while the data blocks contain the actual data content. The disaster recovery site can verify the accuracy of the data content based on this structure. After receiving the data from the production site, the disaster recovery site parses the header information and writes the data content to the virtual disk of the placeholder virtual machine based on the header information.
[0025] In the embodiments of this application, the production site monitors the incremental data generated by the production site's virtual machine, stores the incremental data in the target storage device, and then transmits the incremental data to the disaster recovery site. The disaster recovery site stores the incremental data in the virtual disk of the placeholder virtual machine corresponding to the production site's virtual machine. Therefore, this system solves the technical problems of long data recovery time and high storage costs in related technologies, achieving rapid data recovery to ensure business continuity and significantly reducing storage resource consumption and costs.
[0026] Embodiments of this application provide yet another dual-site disaster recovery system, such as Figure 2 As shown, Figure 2 This application provides another dual-site disaster recovery system, the production site of which also includes a monitoring module. The specific contents of the system are as follows: The monitoring module is used to obtain the resource utilization of each host in the production site and determine the performance of each host based on the resource utilization.
[0027] Specifically, such as Figure 2 The monitoring module is independently applied to the administrator host to monitor and schedule production hosts in the production site. The module monitors the resource utilization (including CPU (Central Processing Unit) utilization, memory utilization, and network performance) and data transmission status of each host in real time and evaluates host performance. When evaluating data transmission, the monitoring module can use a weighted formula; for example, the weight of CPU utilization can be set to 0.7, and the weights of memory utilization and network performance can both be set to 0.1. This weighting is used to determine the host's performance level. When low host performance is detected, a notification is sent to prepare for the selection of a suitable host. It should be noted that the above weights can be freely set according to the specific operating environment and are not limited to the specific weight data mentioned above.
[0028] In the embodiments of this application, the monitoring module monitors the resource usage and data transmission of each host in real time and evaluates the host performance. When the host performance is low, the monitoring module sends a notification to prepare for selecting a suitable host for transmitting incremental data.
[0029] As an optional implementation, the production site also includes a data transmission scheduling module: The data transmission scheduling module is connected to the monitoring module. It is used to receive the resource utilization rate of each host sent by the monitoring module, determine the target host based on the resource utilization rate, and obtain the incremental data of the virtual machines in the target host. The target host is a host whose resource utilization rate meets the performance threshold among multiple hosts.
[0030] Specifically, such as Figure 2 The data transmission scheduling module communicates with the monitoring module. When the host performance is low, the monitoring module sends a notification to the data transmission scheduling module (e.g., via an interface). After receiving the notification from the monitoring module, the data transmission scheduling module selects a host with higher performance based on the current resource usage of each host. It then reissues the data transmission command on the host whose resource usage meets the performance threshold. In other words, it reissues the data transmission command on the target host with the highest performance (meeting the performance threshold) to transmit incremental data to the disaster recovery site.
[0031] In the embodiments of this application, since the target storage pool can achieve resource sharing on multiple hosts, the data transmission scheduling module can continuously maintain the state of issuing data transmission instructions on the host with higher performance based on the host's resource utilization rate, thereby improving the efficiency and reliability of data transmission.
[0032] As an optional implementation, the production site also includes a data transmission module: The data transmission module is connected to the data transmission scheduling module. It is used to receive the identification information sent by the data transmission scheduling module to identify the target host, select the incremental data corresponding to the target host from the target storage pool based on the identification information, and transmit the incremental data to the placeholder virtual machine corresponding to the virtual machine of the target host in the disaster recovery site.
[0033] Specifically, the target host's identification information includes a score characterizing the target host's resource utilization rate, obtained using the weighting formula applied in the monitoring module described above. Optionally, the target host's identification information is used to identify the target host itself. For example, if the target host is host A, then host A's identifier is A; or, by obtaining the MAC address (Media Access Control Address) of a host that meets the performance threshold, that host can be located and designated as the target host. Figure 2The data transmission module is connected to the data transmission scheduling module. After receiving instructions from the data transmission scheduling module, the data transmission module determines the target host based on the aforementioned identification information, reads the incremental data of the virtual machine in the target host from the target storage pool, encrypts the data, and transmits it to the disaster recovery host in the disaster recovery site through the disaster recovery network. Since the disaster recovery host contains a placeholder virtual machine corresponding to the target virtual machine, the disaster recovery site, after receiving the encrypted data, decrypts the data and transmits the incremental data to the virtual disk in the placeholder virtual machine. The incremental data is transmitted in the form of data packets, which include a data packet header and data blocks. The data packet header contains metadata information of the data blocks, and the data blocks are the actual data content. The disaster recovery site can verify the accuracy of the data content based on this structure. In the embodiments of this application, the data transmission module transmits data to the disaster recovery site, data encryption ensures the data security of network transmission, and the dedicated disaster recovery network ensures the efficiency of data transmission, thus guaranteeing the data security of incremental data during transmission and improving the efficiency of data transmission.
[0034] As an optional implementation, the target storage pool includes a backup storage pool, which exists independently of the host: The backup storage pool is used to store incremental data when the storage resource utilization rate is less than a preset threshold. The preset threshold is used to determine whether incremental data should be written to the backup storage pool.
[0035] Specifically, such as Figure 2 As shown, when the target storage pool is a backup storage pool, if the storage resource utilization rate of the storage resources allocated to virtual machines in the backup storage pool is less than a preset threshold, i.e., the storage resource utilization rate is less than 95% (this value can be determined by the user and is not strictly limited to 95%; when the storage resource utilization rate is less than this value, it is considered that the storage resource utilization rate is normal), incremental data is written to the backup storage pool. When incremental data of multiple virtual machines on different hosts is written to the backup pool, the data of different virtual machines is written to different storage locations, that is, the incremental data of virtual machines will be directly written to the specified space allocated to that virtual machine in the backup storage pool.
[0036] It's important to know that the backup storage pool is built using SSDs (Solid State Drives). Because it supports concurrent writes and low-latency reads, it can be shared by multiple hosts simultaneously and can also dynamically allocate storage resources for each virtual machine, avoiding capacity bottlenecks.
[0037] In the embodiments of this application, the utilization rate of the storage resources of the backup storage pool is judged by using a preset threshold. When the utilization rate of the storage resources of the backup storage pool is normal, the production site writes incremental data to the backup storage pool. Since it supports concurrent writing and low-latency reading, it can be shared by multiple hosts at the same time, thereby realizing the dynamic allocation of storage resources used by each virtual machine and avoiding capacity bottlenecks.
[0038] As an optional implementation, the target storage pool includes shared memory, which is stored on a host at the production site: Shared memory is used to store the address information of incremental data when the storage resource utilization of the backup storage pool exceeds a preset threshold.
[0039] Specifically, such as Figure 2 As shown, when the target storage pool is shared memory within the host, if the storage resource utilization rate of the storage resources allocated to the virtual machine in the backup storage pool is less than a preset threshold, for example, when the storage resource utilization rate of the backup storage pool is greater than 95%, the address information of the incremental data is written into the shared memory.
[0040] Optionally, incremental data can be written to shared memory in the form of a bitmap. The bitmap only records the address of data changes, not the actual data, thus requiring very little space to record incremental data changes. The data transfer module retrieves the incremental data based on the bitmap in shared memory and the disk data for transmission. Furthermore, the address information of the incremental data is written to shared memory instead of the backup storage pool in the following three situations: 1. If the disaster recovery network bandwidth is low or other reasons cause the data transmission speed to be lower than the data writing speed, the data in the backup pool will continue to grow, and the resource utilization rate in the backup pool will reach a high level.
[0041] Specifically, the production site generates incremental data faster than the disaster recovery site acquires incremental data, and the backup storage pool, acting as a buffer, is continuously filled. Based on this situation, the address information of the incremental data can be stored in shared memory to reduce the load on the backup storage pool.
[0042] 2. The backup storage pool is experiencing a problem and is therefore unusable.
[0043] Specifically, hardware or software issues can render a backup storage pool unusable. The hardware of a backup storage pool is the physical carrier of data storage (e.g., solid-state drives). A failure in critical hardware will directly lead to inaccessibility or data loss, which is the most common cause of unavailability. Backup storage pools rely on software systems (e.g., storage operating systems, backup management software) for data management. Software malfunctions can prevent data from being written correctly, resulting in anomalies, commonly caused by configuration errors or inherent software defects.
[0044] 3. When multiple virtual machines simultaneously write incremental data to the backup storage pool, the write I / O (Input / Output) exceeds the storage pool's maximum I / O.
[0045] Specifically, because multiple virtual machines simultaneously write incremental data to the backup storage pool, the total rate of incremental data production far exceeds the maximum write speed of the backup storage pool, causing the backup storage pool's storage space to be quickly filled. To avoid this problem, when write I / O exceeds the backup storage pool's maximum I / O, the address information of the incremental data is stored in shared memory to alleviate the storage pressure on the backup storage pool.
[0046] In the embodiments of this application, by applying shared memory, when the resource rate is abnormal, data is written to the shared memory in bitmap form. Finally, regardless of the data's location, the data transmission module reads the corresponding data content for transmission. That is, when the data is in the backup pool, it directly reads and transmits the data; when the data is in shared memory, it reads the address information, finds the address of the incremental data using the address information, reads the incremental data, and then transmits it. The application of shared memory in this embodiment solves the problem in related technologies where abnormalities or failures in the backup storage pool lead to difficulties in incremental data transmission, thereby shortening the incremental data transmission time and improving the transmission efficiency. Furthermore, the combined application of the backup storage pool and shared memory further enhances the reliability and stability of disaster recovery.
[0047] The embodiments of this application provide a dual-site disaster recovery method, such as... Figure 3 As shown, Figure 3 The flowchart of a dual-site disaster recovery method provided in this application embodiment is as follows: Step S301: Monitor the incremental data generated by the virtual machines in the host and store the incremental data to the target storage pool.
[0048] Specifically, the production site is responsible for monitoring the data generated by the virtual machines within the production site host. Since the virtual machines on the production site continuously generate incremental data while providing services, the production site starts a monitoring process for each virtual disk of the working virtual machines. This monitoring process intercepts incremental data from the virtual disks and redirects it to a target storage pool. This target storage pool is used to store the incremental data from the production site.
[0049] Step S302: The incremental data is retrieved from the target storage pool and sent to the placeholder virtual machine corresponding to the virtual machine in the disaster recovery site.
[0050] Specifically, the production site reads incremental data from the target storage pool for virtual machines and transmits it to the disaster recovery site via the disaster recovery network. The incremental data is transmitted in the form of data packets, each containing a header and data blocks. The header contains metadata information for the data blocks, while the data blocks contain the actual data content. The disaster recovery site can verify the accuracy of the data content based on this structure. After receiving the data from the production site, the disaster recovery site parses the header information and writes the data content to the virtual disk in the placeholder virtual machine based on the header information.
[0051] In the embodiments of this application, the production site monitors the incremental data generated by the production site's virtual machine, stores the incremental data in the target storage pool, and then transmits the incremental data to the disaster recovery site. The disaster recovery site stores the incremental data in the virtual disk of the placeholder virtual machine corresponding to the production site's virtual machine. Therefore, this method solves the technical problems of long data recovery time and high storage costs in related technologies, achieving rapid data recovery to ensure business continuity and significantly reducing storage resource consumption and costs.
[0052] This embodiment provides yet another dual-site disaster recovery method. Figure 4 This is a flowchart of another dual-site disaster recovery method according to an embodiment of this application, wherein the target storage pool includes a backup storage pool and shared memory, such as... Figure 4 As shown, the process includes the following steps: Step S401: Monitor incremental data generated by virtual machines within the host and store the incremental data to the target storage pool. For details, please refer to [link to relevant documentation]. Figure 1 Step S301 of the illustrated embodiment will not be described again here.
[0053] Step S402: The incremental data is retrieved from the target storage pool and sent to the placeholder virtual machine corresponding to the virtual machine in the disaster recovery site.
[0054] Specifically, step S402 includes: Step S4021: If the address information of incremental data does not exist in the shared memory, then the incremental data is obtained from the backup storage pool.
[0055] Specifically, the utilization rate of storage resources allocated to virtual machines in the backup storage pool is considered normal. Specifically, when the utilization rate of storage resources allocated to virtual machines in the backup storage pool is less than or equal to a preset threshold (i.e., less than or equal to 95%), it is considered that the utilization rate of storage resources allocated to virtual machines in the backup storage pool is normal. When the utilization rate of storage resources allocated to virtual machines in the backup storage pool is normal, there is no address information for incremental data in the shared memory. This means that the incremental data of the virtual machines is directly written into the designated space allocated in the backup pool, and all incremental data is stored in the backup storage pool. The backup tool compares the baseline information (such as data fingerprint, LSN (Log Sequence Number), etc.) of the previous backup data (incremental data) to identify differences and retrieves newly added and modified incremental data from the backup storage pool.
[0056] Step S4022: If address information of incremental data exists in the shared memory, then based on the address information, extract the first incremental data corresponding to the address information, and extract the second incremental data from the backup storage pool. The incremental data includes the first incremental data and the second incremental data.
[0057] Specifically, the first incremental data is the incremental data corresponding to the address information stored in shared memory, and the second incremental data is the incremental data stored in the backup storage pool. When the utilization rate of the storage resources allocated to the virtual machine in the backup storage pool is abnormal, the virtual machine incremental data will no longer be directly written to the backup pool. Instead, it will be temporarily stored in shared memory through address information (e.g., address information in the form of a bitmap). This address information is used to record the location of the data change in the first incremental data, rather than the actual first incremental data. It can use very little space to record the changes in incremental data. Afterwards, incremental data will be obtained and transmitted based on the address information in shared memory and disk data.
[0058] An abnormal utilization rate of storage resources allocated to virtual machines in the backup storage pool is specifically defined as follows: when the utilization rate of storage resources allocated to virtual machines in the backup storage pool exceeds a preset threshold (i.e., the utilization rate exceeds 95%), it is considered an abnormal utilization rate of storage resources allocated to virtual machines in the backup storage pool. Since there may still be second incremental data awaiting transmission in the backup storage pool when the utilization rate of storage resources allocated to virtual machines in the backup storage pool is abnormal, the second incremental data in the backup storage pool must be transmitted first, followed by the first incremental data corresponding to the address information in the shared memory.
[0059] Step S4023: If the backup storage pool fails or the storage space in the backup storage pool exceeds a preset threshold, if the address information of incremental data already exists in the shared memory and the backup storage pool has stored the new incremental data, then first extract the incremental data corresponding to the address information, and then obtain the new incremental data from the backup storage pool. The new incremental data is the new incremental data generated by the virtual machine when the backup storage pool recovers from the failure or the storage space does not exceed the preset threshold.
[0060] Specifically, backup storage pool failures manifest in two ways: hardware or software issues can render the backup storage pool unusable. The hardware of the backup storage pool is the physical carrier of data storage (e.g., solid-state drives). A failure in critical hardware will directly lead to inaccessibility or data loss, which is the most common cause of unavailability. Backup storage pools rely on software systems (e.g., storage operating systems, backup management software) for data management. Software malfunctions can prevent data from being written correctly, resulting in anomalies, commonly caused by configuration errors or inherent software defects.
[0061] Therefore, when the backup storage pool fails or the storage resource utilization rate allocated to virtual machines in the backup storage pool is abnormal (i.e., the aforementioned storage resource utilization rate exceeds the preset threshold), the address information of the first incremental data begins to be stored in shared memory. Once the storage resource utilization rate of the backup storage pool returns to normal, the first incremental data corresponding to the address information in shared memory will be transferred first. It's important to understand that the write and read operations of the backup storage pool are similar to a "queue," following a first-in, first-out (FIFO) principle. This means that the second incremental data stored in the backup storage pool first will be processed first. Only after processing the second incremental data in the backup storage pool will the system process the first incremental data corresponding to the address information in shared memory. If the second incremental data in the backup storage pool has not been fully processed, and the storage resource utilization rate of the backup storage pool returns to normal, and new incremental data (i.e., newly added incremental data) is written to the backup storage pool, the new incremental data will not be processed first. Instead, the system will prioritize processing the second incremental data already in the backup storage pool and the first incremental data corresponding to the address information stored in shared memory. Only after all the old data (i.e., the first incremental data and the second incremental data) has been processed will the new incremental data be processed.
[0062] In the embodiments of this application, based on whether the utilization rate of the storage resources allocated to the virtual machine in the backup storage pool is in a normal or abnormal state, it is ensured that all incremental data can be transmitted in an orderly and regular manner, achieving data consistency and further improving the stability of data transmission.
[0063] As an optional embodiment, this application also provides a process for a dual-site disaster recovery scheme based on a backup pool, applicable to dual-site disaster recovery methods and systems, such as... Figure 5 As shown: A dual-site disaster recovery solution based on a backup pool includes: the production site intercepts incremental data of virtual machine services, writes the incremental data to the backup pool, the data transmission scheduling module dynamically schedules data transmission based on host resource utilization, the data transmission module transmits data from the backup storage pool to the disaster recovery site, and the disaster recovery site receives the data and writes it to its local storage device.
[0064] 1) When performing off-site data disaster recovery for virtual machines at the production site, a placeholder virtual machine with the same configuration as the virtual machine at the remote disaster recovery site is created. A monitoring process is started for each virtual disk of the working virtual machine at the production site. The monitoring process intercepts incremental data of the virtual disk and redirects the incremental data to the backup storage pool.
[0065] It's important to understand that incremental data is transmitted in the form of data packets. Each data packet contains a header and data blocks. The header contains metadata about the data blocks, while the data blocks contain the actual data content. The disaster recovery site can use this structure to verify the accuracy of the data content. During off-site disaster recovery, after receiving data from the production site, the disaster recovery site parses the data packet header information and writes the data content to the virtual disk in the placeholder virtual machine based on the header information.
[0066] 2) The backup storage pool is built using SSDs, supporting concurrent writes and low-latency reads. It can be shared by multiple hosts simultaneously and can dynamically allocate storage resources for each virtual machine, avoiding capacity bottlenecks. When incremental data from multiple virtual machines on different hosts is written to the backup pool, the data from different virtual machines is written to different storage locations. The SSD architecture enables elastic expansion of storage resources, allowing for expansion when storage resources are insufficient.
[0067] It's important to know that when storage resources are insufficient, SSDs can expand storage capacity by adding new solid-state drives, thus physically increasing storage capacity.
[0068] 3) When the storage resource utilization rate allocated to the virtual machine in the backup pool is normal, the virtual machine incremental data will be directly written to the designated space allocated in the backup pool. However, under the following circumstances, the virtual machine incremental data will no longer be directly written to the backup pool, but will be temporarily stored in shared memory in the form of a bitmap. Since the bitmap only records the location of data changes, not the actual data, it can use very little space to record incremental data changes. The data transmission module will obtain the incremental data for transmission based on the bitmap in shared memory and the disk data: 1. If the disaster recovery network bandwidth is low or other reasons cause the data transmission speed to be lower than the data writing speed, the data in the backup pool will continue to grow, and the resource utilization rate in the backup pool will reach a high level.
[0069] Specifically, the production site generates incremental data faster than the disaster recovery site acquires incremental data, and the backup storage pool, acting as a buffer, is continuously filled. Based on this situation, the address information of the incremental data can be stored in shared memory to reduce the load on the backup storage pool.
[0070] 2. The backup pool is experiencing a problem and is therefore unusable.
[0071] Specifically, hardware or software issues can render a backup storage pool unusable. The hardware of a backup storage pool is the physical carrier of data storage (e.g., solid-state drives). A failure in critical hardware will directly lead to inaccessibility or data loss, which is the most common cause of unavailability. Backup storage pools rely on software systems (e.g., storage operating systems, backup management software) for data management. Software malfunctions can prevent data from being written correctly, resulting in anomalies, commonly caused by configuration errors or inherent software defects.
[0072] 3. When multiple virtual machines simultaneously write incremental data to the backup pool, the write I / O exceeds the storage pool's maximum I / O.
[0073] Specifically, because multiple virtual machines simultaneously write incremental data to the backup storage pool, the total rate of incremental data production far exceeds the maximum write speed of the backup storage pool, causing the backup storage pool's storage space to be quickly filled. To avoid this problem, when write I / O exceeds the backup storage pool's maximum I / O, the address information of the incremental data is stored in shared memory to alleviate the storage pressure on the backup storage pool.
[0074] Incremental capture based on kernel modules reduces performance overhead; when the backup storage pool reaches a preset threshold, incremental data can be stored in shared memory in bitmap form to achieve lightweight dynamic data interception.
[0075] 4) The monitoring module will monitor the resource usage (including CPU, memory usage and network performance) and data transmission status of each host in real time and evaluate the host performance. When the host performance is low, the monitoring module will notify the data transmission scheduling module.
[0076] Specifically, the monitoring module is independently applied to the administrator's host to monitor and schedule the production hosts on the production site. The monitoring module monitors the resource utilization and data transmission of each host in real time and evaluates host performance. When evaluating data transmission, the monitoring module can introduce a weighted formula to determine host performance, selecting the host with the highest relative performance among all hosts. When low performance is detected on a host, a notification is sent to prepare for the subsequent selection of a suitable host. It should be noted that...
[0077] 5) After receiving the command from the monitoring module, the data transmission scheduling module will select a high-performance host based on the current resource usage of each host and reissue the data transmission command on the high-performance host; the data transmission scheduling module selects the host with excellent performance for data transmission based on the host performance, thus realizing dynamic resource scheduling.
[0078] Specifically, the data transmission scheduling module communicates with the monitoring module. When the host performance is low, the monitoring module sends a notification to the data transmission scheduling module. After receiving the notification from the monitoring module, the data transmission scheduling module selects a high-performance host based on the current resource usage of each host, and then reissues the data transmission command on the selected high-performance host. That is, the data transmission command is reissued on the host with the highest performance (meeting the performance threshold) to transmit incremental data to the disaster recovery site.
[0079] 6) After receiving the instruction from the data transmission scheduling module, the host will read the incremental data of the virtual machine from the backup storage pool, encrypt the incremental data, and then transmit it to the disaster recovery site through the disaster recovery network, thus balancing security and efficiency.
[0080] Specifically, when encrypting incremental data, the incremental data is encrypted by combining it with AES (Advanced Encryption Standard) and metadata verification.
[0081] 7) After receiving the data sent by the production site, the disaster recovery site first decrypts the data and parses the header information of the decrypted data packet. Based on the header information, the data content is written to the virtual disk of the corresponding virtual machine.
[0082] Specifically, after encrypting the incremental data, it is transmitted to the disaster recovery host in the disaster recovery site through the disaster recovery network. Since the disaster recovery host contains a placeholder virtual machine corresponding to the target virtual machine, the disaster recovery site decrypts the encrypted data upon receiving it and transmits the incremental data to the virtual disk in the placeholder virtual machine.
[0083] This application proposes a dual-site disaster recovery scheme based on a backup pool, which stores incremental data in the backup storage pool. The introduction of the backup storage pool enables dynamic expansion of storage capacity, improved storage performance, and reduced storage costs. This scheme allows for real-time data backup and recovery, ensuring data consistency between the two sites, providing a highly reliable disaster recovery solution, and ensuring data security and reliability. By introducing the backup pool as a caching device, this scheme reduces the requirements for storage devices, improves data transmission efficiency, and ensures the reliability and security of data transmission through data encryption.
[0084] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0085] Embodiments of this application also provide an electronic device, such as... Figure 6 As shown, it includes a memory 10 and a processor 20. The memory 10 stores a computer program, and the processor 20 is configured to run the computer program to perform the steps in any of the above-described dual-site disaster recovery method embodiments.
[0086] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the dual-site disaster recovery method when it is run.
[0087] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0088] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0089] The foregoing has provided a detailed description of a dual-site disaster recovery system, method, electronic device, and readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A dual-site disaster recovery system, characterized in that, The disaster recovery system includes: a production site and a disaster recovery site. The production site contains a target storage pool and multiple hosts. Each host contains a virtual machine. The disaster recovery site contains a placeholder virtual machine. The production site is used to monitor incremental data generated by the virtual machine within the host and store the incremental data to the target storage pool. The production site is used to retrieve the incremental data from the target storage pool and send it to the placeholder virtual machine corresponding to the virtual machine within the disaster recovery site.
2. The system according to claim 1, characterized in that, The production site also includes: a monitoring module; The monitoring module is used to obtain the resource utilization rate of each host in the production site, and determine the performance of each host based on the resource utilization rate.
3. The system according to claim 2, characterized in that, The production site also includes: a data transmission scheduling module; The data transmission scheduling module is connected to the monitoring module and is used to receive the resource utilization rate of each host sent by the monitoring module, determine the target host based on the resource utilization rate, and obtain incremental data of the virtual machine in the target host. The target host is a host whose resource utilization rate meets the performance threshold among multiple hosts.
4. The system according to claim 3, characterized in that, The production station also includes: a data transmission module; The data transmission module is connected to the data transmission scheduling module and is used to receive the identification information sent by the data transmission scheduling module to characterize the target host, select incremental data corresponding to the target host from the target storage pool based on the identification information, and transmit the incremental data to the placeholder virtual machine corresponding to the virtual machine of the target host in the disaster recovery site.
5. The system according to claim 1, characterized in that, The target storage pool includes a backup storage pool, which exists independently of the host. The backup storage pool is used to store the incremental data when the storage resource utilization rate is less than a preset threshold, wherein the preset threshold is used to determine whether the incremental data should be written to the backup storage pool.
6. The system according to claim 5, characterized in that, The target storage pool includes: shared memory, which is stored in the host; The shared memory is used to store the address information of the incremental data when the storage resource utilization rate of the backup storage pool is greater than the preset threshold.
7. A dual-site disaster recovery method, characterized in that, The method utilizes the dual-site disaster recovery system described in claim 1 to perform dual-site disaster recovery, and the method includes: Monitor incremental data generated by virtual machines within the host and store the incremental data to the target storage pool; The incremental data is retrieved from the target storage pool and sent to the placeholder virtual machine corresponding to the virtual machine within the disaster recovery site.
8. The method according to claim 7, characterized in that, The target storage pool includes a backup storage pool and shared memory. Retrieving the incremental data from the target storage pool includes: If the address information of the incremental data does not exist in the shared memory, the incremental data is extracted from the backup storage pool; If the address information of the incremental data exists in the shared memory, then based on the address information, the first incremental data corresponding to the address information is extracted, and the second incremental data is extracted from the backup storage pool, wherein the incremental data includes the first incremental data and the second incremental data; If the backup storage pool fails or the storage space in the backup storage pool exceeds a preset threshold, and if the address information of the incremental data already exists in the shared memory and the backup storage pool has stored the new incremental data, then the incremental data corresponding to the address information is extracted first, and then the new incremental data is obtained from the backup storage pool. The new incremental data is new incremental data generated by the virtual machine when the backup storage pool recovers from the failure or the storage space does not exceed the preset threshold.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the dual-site disaster recovery method as described in any one of claims 7 to 8 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the dual-site disaster recovery method as described in any one of claims 7 to 8.