Disaster recovery backup method and device, electronic equipment and medium

By dividing the data synchronization process into initial data and incremental data steps, and by implementing a programmatic fault response mechanism, the real-time and reliability issues of cross-regional data synchronization in traditional disaster recovery and backup solutions within a hyperconverged architecture are resolved, enabling efficient data backup and rapid business recovery.

CN121880101APending Publication Date: 2026-04-17CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA TELECOM CLOUD TECH CO LTD
Filing Date
2025-12-05
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional IT architecture disaster recovery and backup solutions are difficult to apply to hyperconverged architectures, and cannot guarantee the real-time performance and reliability of cross-regional data synchronization.

Method used

The data synchronization process is broken down into steps of first transmitting initial data and then transmitting incremental data. The fault response mechanism is programmed, and cross-site migration operations are implemented through synchronization, detection, and execution modules to ensure high reliability of data synchronization and high availability of services.

Benefits of technology

It effectively mitigates the negative impact of network latency on real-time data synchronization, overcomes the migration complexity challenges brought about by the lack of centralized shared storage in hyperconverged architecture, and achieves highly reliable data backup and highly available business takeover in a distributed environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880101A_ABST
    Figure CN121880101A_ABST
Patent Text Reader

Abstract

The invention discloses a disaster recovery backup method and device, electronic equipment and a medium, and belongs to the technical field of disaster recovery backup. The disaster recovery backup method is applied to a local hyper-converged site, the local hyper-converged site comprises a plurality of hyper-converged all-in-one machines and a virtual machine, each hyper-converged all-in-one machine comprises a plurality of computing nodes, the local hyper-converged site is in communication connection with at least one standby hyper-converged site, and the disaster recovery backup method comprises the following steps: synchronizing initial data to the standby hyper-converged site, obtaining incremental data in the synchronization process of the initial data; synchronizing the incremental data to the standby hyper-fusion site after detecting that the initial data is synchronized; when the fault state is detected, a fault switching strategy is triggered, the fault switching strategy comprises executing a cross-station migration operation, and the cross-station migration operation is used for controlling the standby hyper-convergence station to execute a migration strategy. According to the embodiment of the invention, the problems that a disaster recovery backup method in the prior art is difficult to adapt to a hyper-converged architecture and the real-time performance and the reliability of cross-regional data synchronization are difficult to guarantee are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of disaster recovery and backup technology, and specifically relates to a disaster recovery and backup method, device, electronic device and medium. Background Technology

[0002] With the continuous development of cloud computing and virtualization technologies, hyperconverged infrastructure (HCI) appliances are playing an increasingly important role in data center construction and operation due to their advantages such as high integration, simplified deployment and management, reduced operation and maintenance difficulty and cost, space and energy saving, and flexible expansion.

[0003] Traditional IT architecture disaster recovery and backup solutions often face problems such as high network latency and difficulty in ensuring the real-time performance and reliability of cross-regional data synchronization when dealing with cross-regional data center disaster recovery and backup, and cannot be fully applied to hyperconverged architecture. Summary of the Invention

[0004] The purpose of this invention is to provide a disaster recovery backup method, apparatus, electronic device, and medium that can solve the problems that existing disaster recovery backup methods are difficult to apply to hyperconverged architectures and difficult to guarantee the real-time performance and reliability of cross-regional data synchronization.

[0005] To solve the above-mentioned technical problems, the present invention is implemented as follows: In a first aspect, embodiments of the present invention provide a disaster recovery backup method, the method comprising: The disaster recovery backup method is applied to a local hyperconverged site, which includes multiple hyperconverged appliances and virtual machines. The hyperconverged appliance includes multiple compute nodes. The local hyperconverged site is communicatively connected to at least one backup hyperconverged site. The disaster recovery backup method includes: The initial data is synchronized to the backup hyperconverged site, and incremental data is acquired during the synchronization of the initial data; After the initial data synchronization is completed, the incremental data is synchronized to the backup hyperconverged site. When a fault condition is detected, a fault switching strategy is triggered. The fault switching strategy includes performing a cross-site migration operation, which is used to cooperate with the backup hyperconverged site to perform the migration strategy.

[0006] Optionally, the step of synchronizing initial data to the backup hyperconverged site and acquiring incremental data during the synchronization of the initial data includes: The initial data is divided into multiple data blocks; Traverse multiple data blocks and synchronize multiple data blocks to the backup hyperconverged site; Acquire the incremental data generated during the synchronization of the multiple data blocks.

[0007] Alternatively, disaster recovery backup methods may also include: Detect the data changes of multiple data blocks contained in the initial data; The data changes of multiple data blocks are encapsulated into multiple events containing hash identifiers. The data blocks and the hash identifiers correspond one-to-one, and the changes in the hash identifiers are used to characterize the data changes of the data blocks. The events are synchronized to the backup hyperconverged site.

[0008] Optionally, synchronizing multiple events to the backup hyperconverged site includes: The order of the events is determined, and the data blocks corresponding to the events are written into the storage unit of the backup hyperconverged site in the order of the events.

[0009] Optionally, the fault status includes the compute node fault, the virtual machine fault, and the network fault.

[0010] Optionally, when a fault state is detected, triggering a fault switching strategy includes: Obtain the status information of the computing node, and determine the fault status of the computing node based on the status information; When a failure of one of the compute nodes is detected, a virtual machine migration operation is performed to migrate the memory state of the virtual machine to another normally functioning compute node. When multiple compute node failures are detected, the cross-site migration operation is performed to migrate the virtual machine to the backup hyperconverged site.

[0011] Optionally, obtaining the status information of the computing node and determining the fault status of the computing node based on the status information includes: The computing nodes are divided into multiple autonomous groups, and the computing nodes in each autonomous group exchange heartbeat data at regular intervals. The heartbeat data is used to generate the status information of each computing node. Within a preset time period, the exchange status of the heartbeat data is determined. If the computing node does not receive the heartbeat data corresponding to the computing node, it is determined that the corresponding computing node is in a fault state and the fault information is reported.

[0012] Optionally, the step of triggering a fault switching strategy when a fault state is detected further includes: Detect the network connectivity status between each computing node, including the management network connectivity status, the service network connectivity status, and the storage network connectivity status; When the management network connectivity status is detected to be faulty, the management network connectivity status is forwarded to other computing nodes through the storage network; If a failure is detected in the connectivity status of the service network or the connectivity status of the storage network, the corresponding compute node is determined to be in a fault state, and the virtual machine migration operation is executed.

[0013] Optionally, the step of triggering a fault switching strategy when a fault state is detected further includes: When the virtual machine is in a fault state, control the virtual machine to restart; If the virtual machine remains in a faulty state after restarting, the virtual machine migration operation is performed.

[0014] Alternatively, disaster recovery backup methods may also include: Retrieve multiple local image data blocks; By comparing multiple local mirror data blocks with multiple data blocks from the backup fusion site, the local mirror data block that needs to be transmitted is determined.

[0015] Alternatively, disaster recovery backup methods may also include: Filter out the local mirror data blocks that are inconsistent with the data blocks of the backup fusion site, and mark these local mirror data blocks as to be transmitted; Select the portion of the local mirror data blocks that are consistent with the data blocks of the backup fusion site, and mark this portion of the local mirror data blocks as reusable.

[0016] Alternatively, disaster recovery backup methods may also include: The local mirror data blocks marked as to be transmitted are combined and transmitted to the backup fusion site.

[0017] Secondly, embodiments of the present invention provide a disaster recovery backup apparatus, comprising: The synchronization module is used to synchronize initial data to the backup hyperconverged site and acquire incremental data during the synchronization process of the initial data; The first detection module is used to synchronize the incremental data to the backup hyperconverged site after detecting that the initial data synchronization is complete. An execution module is used to trigger a failover strategy when a fault condition is detected. The failover strategy includes performing a cross-site migration operation. The cross-site migration operation is used to cooperate with the backup hyperconverged site to perform the migration strategy.

[0018] Optionally, the synchronization module includes; A partitioning submodule is used to divide the initial data into multiple data blocks; The first synchronization submodule is used to traverse multiple data blocks and synchronize the multiple data blocks to the backup hyperconverged site; The incremental data submodule is used to acquire the incremental data generated during the synchronization of multiple data blocks.

[0019] Optionally, the disaster recovery backup device also includes: The second detection module is used to detect the data changes of multiple data blocks contained in the initial data; An encapsulation module is used to encapsulate the data changes of multiple data blocks into multiple events containing hash identifiers, wherein the data blocks and the hash identifiers correspond one-to-one, and the changes in the hash identifiers are used to characterize the data changes of the data blocks; The second synchronization module is used to synchronize multiple events to the backup hyperconverged site.

[0020] Optionally, the second synchronization module includes: A sorting unit is used to determine the order of the multiple events and write the multiple data blocks corresponding to the multiple events into the storage unit of the backup hyperconverged site according to the order of the events.

[0021] Optionally, the fault status includes the compute node fault, the virtual machine fault, and the network fault.

[0022] Optionally, the execution module includes: The judgment submodule is used to obtain the status information of the computing node and judge the fault status of the computing node based on the status information; The first migration submodule is used to perform a virtual machine migration operation when a failure of one of the compute nodes is detected. The virtual machine migration operation is used to migrate the memory state of the virtual machine to other normally operating compute nodes. The second migration submodule is used to perform the cross-site migration operation when multiple compute node failures are detected. The cross-site migration operation is used to migrate the virtual machine to the backup hyperconverged site.

[0023] Optionally, the judgment submodule includes: A data exchange unit is used to divide the multiple computing nodes into multiple autonomous groups, and the computing nodes in each autonomous group exchange heartbeat data at regular intervals. The heartbeat data is used to generate the status information of each computing node. The detection unit is used to determine the exchange status of the heartbeat data within a preset time. If the computing node does not receive the heartbeat data corresponding to the computing node, it determines that the corresponding computing node is in a fault state and reports the fault information.

[0024] Optionally, the execution module also includes: The network detection submodule is used to detect the network connectivity status between each computing node, including the management network connectivity status, the service network connectivity status, and the storage network connectivity status. The forwarding submodule is used to forward the management network connectivity status to other computing nodes through the storage network when the management network connectivity status is detected to be in a fault state. The first execution submodule is used to determine that the corresponding compute node is in a fault state when the service network connectivity status or the storage network connectivity status is detected to be faulty, and to execute the virtual machine migration operation.

[0025] Optionally, the execution module also includes: The restart submodule is used to control the restart of the virtual machine when the virtual machine is in a fault state; The second execution submodule is used to perform the virtual machine migration operation when the virtual machine is still in a fault state after restarting.

[0026] Optionally, the disaster recovery backup device further includes a third detection module, which is used to acquire multiple local mirror data blocks, compare the multiple local mirror data blocks with the multiple data blocks of the backup converged site, and determine the local mirror data blocks that need to be transmitted.

[0027] Optionally, the third detection module is further configured to filter out the local mirror data blocks that are inconsistent with the data blocks of the backup fusion site, and mark these local mirror data blocks as to be transmitted; and to filter out the local mirror data blocks that are consistent with the data blocks of the backup fusion site, and mark these local mirror data blocks as reusable.

[0028] Optionally, the execution module is further configured to combine the portions of the local mirror data blocks marked as to be transmitted and transmit them to the backup fusion site.

[0029] Thirdly, embodiments of the present invention provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0030] Fourthly, embodiments of the present invention provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0031] In this embodiment of the invention, the disaster recovery backup method is applied to a local hyperconverged site. The local hyperconverged site includes multiple hyperconverged appliances and virtual machines. The hyperconverged appliance includes multiple compute nodes. The local hyperconverged site is communicatively connected to at least one standby hyperconverged site. The disaster recovery backup method first performs an initial full data synchronization, continuously acquiring incremental data during this period. This ensures that the standby hyperconverged site obtains a complete basic data copy, while minimizing the data loss window caused by the long duration of full synchronization. Subsequently, after the initial data synchronization is completed, the acquired incremental data is synchronized, thereby avoiding congestion that may be caused by simultaneous transmission of incremental and full data over the network. This improves network bandwidth utilization efficiency and the reliability of the incremental data synchronization process, ensuring the consistency and accuracy of the disaster recovery backup. When a fault condition is detected, a failover strategy including cross-site migration operations is automatically triggered. The complex migration process and business takeover process are implemented through preset automated operations, significantly reducing the delay and error risks introduced by manual diagnosis, decision-making, and operation in traditional disaster recovery solutions, thereby effectively shortening business interruption time. This invention addresses the issue by breaking down the data synchronization process into steps of transmitting initial data and then incremental data, and by programming the fault response mechanism. This mitigates the negative impact of network latency on real-time data synchronization and overcomes the migration complexity challenges posed by the lack of centralized shared storage in hyperconverged architectures. Ultimately, it achieves highly reliable data backup and highly available business takeover in a distributed environment.

[0032] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0033] Figure 1 This is a flowchart of a disaster recovery backup method according to an embodiment of the present invention; Figure 2 This is a flowchart of a disaster recovery backup method in another embodiment of the present invention; Figure 3 This is a schematic diagram of the communication connection between a local hyperconverged site and a backup hyperconverged site in one embodiment of the present invention; Figure 4 This is a flowchart of a disaster recovery backup method in another embodiment of the present invention; Figure 5 This is a structural block diagram of a disaster recovery backup device according to an embodiment of the present invention. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] The terms "first," "second," etc., used in the specification and claims of this invention are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0036] With the continuous development of cloud computing and virtualization technologies, hyperconverged infrastructure (HCI) appliances are playing an increasingly important role in data center construction and operation due to their advantages such as high integration, simplified deployment and management, reduced operation and maintenance difficulty and cost, space and energy saving, and flexible expansion.

[0037] Traditional IT architecture disaster recovery and backup solutions often face problems such as high network latency and difficulty in ensuring the real-time performance and reliability of cross-regional data synchronization when dealing with cross-regional data center disaster recovery and backup, and cannot be fully applied to hyperconverged architecture.

[0038] The inventors also discovered that with the widespread application of cloud computing, hyperconverged infrastructure (HCI) appliances will become one of the important trends in data center construction. Key technologies include resource computing, network planning, data storage, computer virtualization, and performance-related technologies such as redundant data processing, data compression, and snapshots. HCI architecture adopts a software-defined approach, abstracting hardware resources through virtualization technology to form virtual pools of computing, storage, and network resources for upper-layer applications. This can bring optimal efficiency, flexibility, data protection, and cost-effectiveness to virtualized data centers. Unlike traditional IT infrastructure, HCI infrastructure typically uses distributed storage systems, with data stored on the physical server's hard drive, and virtual machine nodes providing services externally. HCI appliances are converged products that integrate storage into compute servers, resulting in a high degree of hardware and software centralization. Therefore, disaster recovery and backup solutions for HCI appliances need to address the challenges of high availability and data consistency for virtual machines in cross-regional data centers.

[0039] Therefore, one of the core concepts of this invention is to decompose the data synchronization process into steps of transmitting initial data first and then incremental data, and to program the fault response mechanism, thereby mitigating the negative impact of network latency on real-time data synchronization, overcoming the migration complexity problem caused by the lack of centralized shared storage in hyperconverged architecture, and ultimately achieving highly reliable data backup and highly available business takeover in a distributed environment.

[0040] This invention provides a disaster recovery backup method, apparatus, electronic device, and medium that can solve the problems that existing disaster recovery backup methods are difficult to apply to hyperconverged architectures and difficult to guarantee the real-time performance and reliability of cross-regional data synchronization.

[0041] The disaster recovery backup method provided by the embodiments of the present invention will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0042] Figure 1 This is a flowchart of a disaster recovery backup method in one embodiment of the present invention, as follows: Figure 1 As shown, the disaster recovery backup method is applied to a local hyperconverged site, which includes multiple hyperconverged appliances and virtual machines. The hyperconverged appliances include multiple compute nodes, and the local hyperconverged site is communicatively connected to at least one backup hyperconverged site.

[0043] Specifically, hyperconverged appliances are based on standard general-purpose servers (such as x86 and ARM servers) and deploy software-defined cloud computing platform software to achieve the convergence of computing, storage, and networking, forming a software-defined cloud service center centered on virtualization with integrated hardware and software IT infrastructure. Furthermore, since hyperconverged appliances are integrated computing and storage systems, they are integrated into a unified storage pool through distributed storage software technology. It is also necessary to back up the disk image files of virtual machines. These data volumes are often very large, so the requirements for backup storage space and backup efficiency are higher.

[0044] Figure 1 The illustrated embodiment includes at least the following steps 101-103.

[0045] In step 101, initial data is synchronized to the backup hyperconverged site, and incremental data is acquired during the synchronization of the initial data.

[0046] A communication connection is established between the local hyperconverged site and the backup hyperconverged site, and initial data is synchronously transmitted to the backup hyperconverged site. This allows for data backup transmission before a failure occurs, ensuring that at least one backup hyperconverged site can obtain a basic data sample and preventing data loss due to sudden failure of the local hyperconverged site. Understandably, during the initial data synchronization process, the local hyperconverged site continues to process business operations, thus generating some incremental data. This incremental data is then acquired during the data synchronization process.

[0047] In step 102, after the initial data synchronization is detected to be complete, the incremental data is synchronized to the backup hyperconverged site. Once all initial data has been synchronized to the corresponding hyperconverged site, the acquired incremental data is then transmitted. On the one hand, uploading incremental data while the initial data is being transmitted will consume the network channel for uploading the initial data, thus reducing the synchronization speed of the initial data. On the other hand, while the initial data and incremental data are being uploaded simultaneously, incremental data is still being generated continuously. In some embodiments, the incremental data may change. Therefore, synchronizing incremental data only after all the initial data has been synchronized can improve the utilization efficiency of network bandwidth and the effectiveness of the incremental data synchronization process, ensuring the consistency of disaster recovery backup.

[0048] In step 103, when a fault state is detected, a fault switching strategy is triggered. The fault switching strategy includes performing a cross-site migration operation, which is used to cooperate with the backup hyperconverged site to perform the migration strategy.

[0049] In step 103, if the local hyperconverged site experiences a failure, a failover strategy is triggered. It should be noted that the failure status of the local hyperconverged site includes various scenarios, and the corresponding failover strategies differ. Specifically, the failover strategy includes a cross-site migration operation, which is used to migrate all functions or data of the local converged site to a backup converged site. In this embodiment of the invention, the corresponding failover strategy is triggered and automatically executed based on the corresponding failure status, eliminating the need for manual judgment, thereby reducing delays and error risks, and effectively shortening service interruption time.

[0050] In this embodiment of the invention, the disaster recovery backup method is applied to a local hyperconverged site. The local hyperconverged site includes multiple hyperconverged appliances and virtual machines. The hyperconverged appliance includes multiple compute nodes. The local hyperconverged site is communicatively connected to at least one standby hyperconverged site. The disaster recovery backup method first performs an initial full data synchronization, continuously acquiring incremental data during this period. This ensures that the standby hyperconverged site obtains a complete basic data copy, while minimizing the data loss window caused by the long duration of full synchronization. Subsequently, after the initial data synchronization is completed, the acquired incremental data is synchronized, thereby avoiding congestion that may be caused by simultaneous transmission of incremental and full data over the network. This improves network bandwidth utilization efficiency and the reliability of the incremental data synchronization process, ensuring the consistency and accuracy of the disaster recovery backup. When a fault condition is detected, a failover strategy including cross-site migration operations is automatically triggered. The complex migration process and business takeover process are implemented through preset automated operations, significantly reducing the delay and error risks introduced by manual diagnosis, decision-making, and operation in traditional disaster recovery solutions, thereby effectively shortening business interruption time. This invention addresses the issue by breaking down the data synchronization process into steps of transmitting initial data and then incremental data, and by programming the fault response mechanism. This mitigates the negative impact of network latency on real-time data synchronization and overcomes the migration complexity challenges posed by the lack of centralized shared storage in hyperconverged architectures. Ultimately, it achieves highly reliable data backup and highly available business takeover in a distributed environment.

[0051] In some embodiments, step 101 includes the following sub-steps: In sub-step 111, the initial data is divided into multiple data blocks.

[0052] In the above steps, the initial data is divided into multiple data blocks. The synchronization task of the originally large amount of initial data is decomposed into multiple independent and manageable data block synchronization tasks. This enables scheduling and monitoring of the transmission process, which facilitates parallel transmission or interrupted transmission. In some embodiments, when facing unstable wide area network links, dividing the initial data into multiple data blocks can effectively improve the success rate of the overall synchronization task and alleviate the problem of full synchronization failure or excessive time consumption caused by network latency or interruption.

[0053] In sub-step 112, multiple data blocks are traversed and the multiple data blocks are synchronized to the backup hyperconverged site.

[0054] Traversing and synchronizing multiple data blocks enables controllability of the synchronization process. Specifically, in some embodiments, synchronization tasks can be initiated sequentially or according to a strategy (such as by data block priority), and the synchronization status of each data block can be accurately recorded. If the synchronization process is interrupted for any reason, only the data blocks that have not been synchronized can be resumed, avoiding the resource waste caused by retransmitting the entire initial data in traditional solutions, and significantly reducing the average time and network overhead required to complete the full data baseline synchronization.

[0055] In sub-step 113, the incremental data generated during the synchronization of the multiple data blocks is obtained.

[0056] By acquiring incremental data in real time during the traversal of synchronized data blocks, it ensures that no data updates occurring on the local hyperconverged site are missed from the start of the initial data synchronization task. This allows the synchronization of multiple data blocks and the acquisition of incremental data to overlap in time, effectively reducing the risk of data loss or the need to pause business writes while waiting for the initial data synchronization to complete. This improves synchronization speed while also ensuring data consistency during the synchronization process.

[0057] In summary, data block partitioning improves transmission efficiency and fault tolerance, while orderly synchronization of data blocks and incremental data ensures consistency during the data synchronization process. This enhances the reliability, efficiency, and accuracy of the initial data synchronization phase and strengthens the practicality of backup methods in hyperconverged cross-regional scenarios.

[0058] In other embodiments, please refer to Figure 2 , Figure 2 A flowchart of a disaster recovery backup method according to another embodiment of the present invention includes the following steps: Step 201: Synchronize the initial data to the backup hyperconverged site, and acquire incremental data during the synchronization process of the initial data.

[0059] A communication connection is established between the local hyperconverged site and the backup hyperconverged site, and initial data is synchronously transmitted to the backup hyperconverged site. This allows for data backup transmission before a failure occurs, ensuring that at least one backup hyperconverged site can obtain a basic data sample and preventing data loss due to sudden failure of the local hyperconverged site. Understandably, during the initial data synchronization process, the local hyperconverged site continues to process business operations, thus generating some incremental data. This incremental data is then acquired during the data synchronization process.

[0060] In step 202, after the initial data synchronization is detected to be complete, the incremental data is synchronized to the backup hyperconverged site. Once all initial data has been synchronized to the corresponding hyperconverged site, the acquired incremental data is then transmitted. On the one hand, uploading incremental data while the initial data is being transmitted will consume the network channel for uploading the initial data, thus reducing the synchronization speed of the initial data. On the other hand, while the initial data and incremental data are being uploaded simultaneously, incremental data is still being generated continuously. In some embodiments, the incremental data may change. Therefore, synchronizing incremental data only after all the initial data has been synchronized can improve the utilization efficiency of network bandwidth and the effectiveness of the incremental data synchronization process, ensuring the consistency of disaster recovery backup.

[0061] In step 203, the data changes of multiple data blocks contained in the initial data are detected.

[0062] In embodiments of the present invention, multiple data blocks in the initial data are mapped to a hash graph. Each data block has a unique hash identifier as a key, and the corresponding value is the specific content of the data block. When a data block changes, its corresponding hash identifier also changes, thereby identifying the data changes of the data block. It is understood that in embodiments of the present invention, any change in the content of a data block is considered a change in the data of the data block, and its hash identifier also changes accordingly. It should be noted that a hash graph is a distributed ledger technology that uses a directed acyclic graph structure and a consensus algorithm to ensure data consistency and security.

[0063] Therefore, detecting changes in multiple data blocks can reduce the monitoring scope from massive amounts of initial data to the range of data blocks that have changed, and can be used to detect more subtle data changes, reducing the amount of data that needs to be processed later.

[0064] In step 204, the data changes of multiple data blocks are encapsulated into multiple events containing hash identifiers. The data blocks and hash identifiers correspond one-to-one, and the changes in the hash identifiers are used to characterize the data changes of the data blocks.

[0065] Encapsulating changes to data blocks into events containing one-to-one hash identifiers allows us to understand that each hash identifier corresponds to a unique and complete data block. The generation and comparison of hash identifiers is a deterministic computation process that can unambiguously identify whether a specific data block has undergone substantial changes, effectively avoiding misjudgments that may arise based on metadata such as modification time.

[0066] Changes to data blocks are encapsulated into corresponding events. Each event becomes a standardized data unit carrying change information (corresponding hash identifier), time sequence information (which may include timestamps), and association information (such as the hash of the preceding event).

[0067] In step 205, multiple events are synchronized to the backup hyperconverged site.

[0068] In this embodiment of the invention, multiple events are synchronized to a standby hyperconverged site. The synchronization object is transformed from the original data blocks into events that characterize their changes. This is generally lighter in terms of network transmission load. More importantly, after receiving the event-based data block structure, the standby hyperconverged site can not only know which data blocks have changed, but also verify the integrity of the received data blocks through the hash identifier carried by the events. Furthermore, the logical relationship between the events provides the possibility of constructing a consistent data state evolution history across nodes. This directly enhances the verifiability of the data synchronization process and the consistency guarantee of the final state.

[0069] In some embodiments, please refer to Figure 3 There are multiple backup hyperconverged sites. The virtual machine can propagate events to these sites via a gossip protocol in a hash graph. The gossip protocol is a decentralized communication mechanism used to efficiently propagate information between network nodes. It achieves efficient, decentralized information propagation and consensus through random communication and event logging, featuring high efficiency, decentralization, and high fault tolerance, making it suitable for various distributed systems. Due to the fast convergence of the gossip protocol, each event can be quickly propagated to every hyperconverged site in the hash graph. This invention achieves accurate and reliable detection of data block changes through hash identifiers, and also realizes the structuring and lightweighting of the data block synchronization process through event-based methods. Together, they provide a systematic method for achieving efficient and verifiable data block synchronization. This not only optimizes the utilization of network bandwidth and improves synchronization efficiency, but also enhances the ability of the entire backup method to guarantee data consistency in unreliable network environments by introducing deterministic data identifiers (i.e., hash identifiers), making the data synchronization steps more accurate, efficient, and reliable.

[0070] Specifically, step 205 includes the following steps: determining the order of the multiple events, and writing the multiple data blocks corresponding to the multiple events into the storage unit of the backup hyperconverged site in accordance with the order of the events.

[0071] In distributed systems, events are generated concurrently. However, network transmission delays and path differences may cause the physical order in which events are received to differ from their logical order of occurrence. By introducing a deterministic consensus mechanism, such as a virtual voting algorithm based on logical timestamps or hash graphs, the system can assign a sequential order to all events. This solves the problem of data state ambiguity that may be caused by out-of-order events. For example, applying multiple updates to the same data block in different orders can lead to drastically different final results, further improving the consistency of data transmitted across sites.

[0072] Multiple data blocks corresponding to multiple events are written to the storage unit of the standby hyperconverged site in the order of the events. When the standby hyperconverged site saves data blocks, it needs to follow the agreed event order to ensure that even if the data blocks arrive out of order due to network reasons, the standby hyperconverged site will cache them and store them in the correct order. This prevents data logic corruption or rollback caused by inconsistency between the writing order and the local hyperconverged site. This makes the storage image of the standby hyperconverged site not only a static copy of the data blocks of the local hyperconverged site at a certain moment, but also a dynamic copy of its data state evolving in the same historical order. This is crucial for applications that rely on the correct operation order.

[0073] The embodiments of the present invention enhance the robustness of the entire backup method in the face of problems such as network latency, out-of-order delivery, and partial data loss, ensuring that the data copy of the standby hyperconverged site is logically highly consistent with the local hyperconverged site.

[0074] In step 206, when a fault state is detected, a fault switching strategy is triggered. The fault switching strategy includes performing a cross-site migration operation, which is used to cooperate with the backup hyperconverged site to perform the migration strategy.

[0075] In step 206, if the local hyperconverged site experiences a failure, a failover strategy is triggered. It should be noted that the failure status of the local hyperconverged site includes various scenarios, and the corresponding failover strategies differ. Specifically, the failover strategy includes a cross-site migration operation, which is used to migrate all functions or data of the local converged site to a backup converged site. In this embodiment of the invention, the corresponding failover strategy is triggered and automatically executed based on the corresponding failure status, eliminating the need for manual judgment, thereby reducing delays and error risks, and effectively shortening service interruption time.

[0076] Specifically, the fault states include compute node faults, virtual machine faults, and network faults. That is, the fault states are clearly defined to include compute node faults, virtual machine faults, and network faults. Through multi-faceted monitoring and analysis, the problem of diverse fault sources and varying impact ranges in hyperconverged distributed environments can be solved.

[0077] First, compute node failures address scenarios where the physical server hosting the virtualization layer fails or completely crashes. These failures directly affect all virtual machine instances running on that node. Virtual machine failures correspond to software-level issues such as application crashes and service unresponsiveness within the virtual machine, and do not necessarily indicate a problem with the host compute node or underlying storage. Network failures address the complexity of network links in cross-regional disaster recovery scenarios. Network failures manifest in various forms, such as management network interruptions, storage network latency or packet loss, or business network partitions, with impacts ranging from disrupting data synchronization to misjudging site failures. By identifying these failure types, specific decision-making dimensions are provided for detecting failure states, enabling preliminary classification of failure root causes and laying the foundation for subsequently selecting and triggering the most suitable failover strategy.

[0078] Step 206 above also includes the following sub-steps: Sub-step 211: Obtain the status information of the computing node, and determine the fault status of the computing node based on the status information.

[0079] Acquiring status information and determining the fault status of compute nodes relies on agents or management components deployed on the nodes. Obtaining status information through methods such as heartbeat detection and hardware status monitoring provides real-time and reliable input for subsequent decision-making, enabling local hyperconverged sites to upgrade from simple fault diagnosis to assessment of where and to what extent the fault occurs.

[0080] Sub-step 212: When a failure of one of the compute nodes is detected, a virtual machine migration operation is performed to migrate the memory state of the virtual machine to another normally functioning compute node.

[0081] When a single compute node failure is detected, an operation is triggered to migrate the virtual machine's memory state to another normally functioning compute node. It's understood that in a hyperconverged architecture, due to the convergence of compute and storage, the virtual machine's disk data is typically already stored in a distributed storage pool across nodes, eliminating the need for migration. Therefore, the core of local migration is transferring the virtual machine's runtime state (memory). This operation is performed solely within the local hyperconverged site's internal network, resulting in low latency, ample bandwidth, and rapid service recovery, typically achieving a second-level RTO (Recovery Time Objective). Simultaneously, it avoids immediately activating a backup hyperconverged site due to a single compute node failure, thus saving cross-site bandwidth resources and maintaining the backup hyperconverged site's capacity for more severe failures, achieving optimized resource utilization.

[0082] Sub-step 213: When multiple computing node failures are detected, the cross-site migration operation is performed to migrate the virtual machine to the backup hyperconverged site.

[0083] When multiple compute node failures are detected, a cross-site migration operation is triggered to move the virtual machines to the backup hyperconverged site. It is understood that the simultaneous failure of multiple compute nodes may indicate an early sign of rack power outages, network partitioning, or site-wide disasters. In this situation, continuing to search for healthy compute nodes within the local hyperconverged site for migration may be infeasible or unstable. Immediately initiating a cross-site migration operation transfers the entire business load to a known, independent, and intact disaster recovery environment, thereby preventing a complete business interruption due to a potential further expansion of the failure scope.

[0084] By identifying the scale of the failure and matching the corresponding strategy, business continuity is ensured to the maximum extent, while the economy and resource utilization of the entire backup method are also significantly improved.

[0085] Sub-step 211 further includes the following steps: dividing the multiple computing nodes into multiple autonomous groups, and exchanging heartbeat data between the computing nodes in each autonomous group at regular intervals, wherein the heartbeat data is used to generate the status information of each computing node; determining the exchange status of the heartbeat data within a preset time period; if a computing node does not receive the heartbeat data of the corresponding computing node, then determining that the corresponding computing node is in a fault state and reporting the fault information.

[0086] It should be noted that traditional centralized monitoring architectures are susceptible to single-point-of-failure risks due to the monitoring server, and the heartbeats of all nodes must converge at a single point, with network overhead and latency potentially increasing significantly as the number of nodes grows. In this embodiment of the invention, by establishing autonomous groups, the global monitoring responsibility is decomposed into several autonomous groups. Each autonomous group can become an independent fault-aware unit, reducing the overall complexity of the system and its dependence on a central node. Even if an autonomous group temporarily loses connection to the global management platform due to network partitioning, it can still maintain basic mutual monitoring and fault detection capabilities within the group, thereby significantly enhancing the robustness and scalability of the entire state awareness system.

[0087] Within each autonomous group, compute nodes periodically exchange heartbeat data to generate status information. This heartbeat data serves as a periodic survival signal, and its exchange occurs only between a limited number of neighboring compute nodes within the group. This results in short communication paths and low overhead, enabling near real-time awareness of the compute node's survival status. Compared to methods that require continuous polling of complex performance metrics, the heartbeat mechanism identifies the compute node's status with minimal computational and network resource consumption, buying time for quickly triggering subsequent recovery processes.

[0088] Furthermore, using a preset time as a threshold provides a criterion for distinguishing between temporary network jitter and permanent faults, avoiding misjudgments. Once a fault is confirmed, the fault information is immediately reported, ensuring efficient linkage between the state awareness layer and the decision execution layer. The transformation from continuous monitoring to event triggering enables the entire system to proactively and promptly respond to fault events.

[0089] Sub-step 214 involves detecting the network connectivity status between each computing node, including the management network connectivity status, the service network connectivity status, and the storage network connectivity status.

[0090] This invention detects the network connectivity status between each compute node and specifically distinguishes between the management network, service network, and storage network. Since the normal operation of compute nodes in a hyperconverged appliance relies on multiple network planes: the management network controls signaling and heartbeats between compute nodes, the service network carries the service traffic of virtual machines, and the storage network is responsible for data synchronization and storage access between nodes. In some embodiments, only network connectivity may be detected, failing to distinguish faults in specific planes. This invention, by independently detecting the connectivity status of the management network, service network, and storage network, can accurately identify the specific network layer where the fault occurred, enabling precise fault location.

[0091] Sub-step 215: When the management network connectivity status is detected to be in a fault state, the management network connectivity status is forwarded to other computing nodes through the storage network; When a failure is detected in the management network connectivity status, this status is forwarded to other compute nodes via the storage network. A failure in the management network connectivity status means that the compute node may not be able to report its own status through the normal control path, which may be easily misjudged as a complete system crash by other compute nodes. Therefore, using the storage network, which is usually independent and used for large data transmission, as a backup communication channel to forward the status information of the management network failure ensures that even if the management network is interrupted, the critical fault metadata of the compute node can still be perceived. This avoids unnecessary virtual machine migration or site switching triggered by the interruption of a single network path, and significantly improves the accuracy of decision-making in some network failure scenarios.

[0092] In sub-step 216, if a fault is detected in the service network connectivity status or the storage network connectivity status, the corresponding compute node is determined to be in a fault state, and the virtual machine migration operation is executed.

[0093] When a connectivity failure is detected in the business network or storage network, the corresponding compute node is directly identified as faulty and a virtual machine migration operation is executed. A business network failure directly prevents virtual machines from providing services, while a storage network failure may jeopardize data consistency or cause data inaccessibility. Both are serious failures that directly affect business continuity. Therefore, immediately classifying the compute node as faulty and triggering migration is a strategy that prioritizes business availability. This ensures that when problems occur in the core service network or data network, the recovery process can be quickly initiated, and the business load can be transferred to healthy compute nodes or backup hyperconverged sites, thereby minimizing the impact on upper-layer applications.

[0094] By separating and detecting different network planes, accurate fault identification and management are achieved; by setting up backup reporting paths for management network faults, misjudgments are prevented and monitoring robustness is improved; and by quickly deciding on faults in the business network or storage network, the continuity of core services is ensured. This results in a better level of automated switching and service quality in complex network environments. Sub-step 217: When the virtual machine is in a fault state, control the virtual machine to restart; Sub-step 217: When the virtual machine is in a fault state, control the virtual machine to restart. Virtual machine fault states typically stem from software issues such as internal operating system crashes or critical application process termination. Compared to migrating the entire virtual machine instance, attempting a restart on the original compute node only involves the virtual machine's own startup process, avoiding the extensive data transfer and reconstruction work involved in cross-node memory state copying and storage mount switching. Therefore, this operation can be completed in a very short time (usually within seconds), and if successful, it can restore business operations with minimal overhead, making it the preferred path to achieve optimal recovery time. Simultaneously, this avoids unnecessary disturbance to the resources of other virtual machines.

[0095] Sub-step 218: If the virtual machine is still in a fault state after restarting, perform the virtual machine migration operation.

[0096] If the virtual machine remains in a faulty state after restarting, the virtual machine migration operation is executed. A failed restart indicates that the virtual machine fault may not be a simple temporary software error, but rather a deeper conflict with the current host environment (such as driver incompatibility or underlying resource anomalies) or damage that is difficult to heal itself. In this case, continuing recovery attempts on the original compute node may be ineffective, triggering the virtual machine migration operation. Essentially, this involves transferring the virtual machine's workload, along with its problem context, to a completely new, healthy hardware and software host environment. This is equivalent to providing a complete environment reset for the faulty virtual machine, effectively resolving faults caused by a specific node environment. This ensures that even when lightweight recovery methods fail, other means can be used to guarantee the eventual recoverability of the business.

[0097] Therefore, this strategy first attempts the fastest and most economical recovery method, and only initiates a more expensive recovery method when that method fails. This enables business to be restored in seconds in most common software failure scenarios, significantly improving average recovery efficiency and user experience. At the same time, it also provides a certain guarantee for a few complex failures, ensuring the eventual availability of the business and optimizing the overall resource utilization efficiency and recovery performance of the backup solution.

[0098] In scenarios with only two hyperconverged sites—one on-premises and one standby hyperconverged site—an arbitration node needs to be configured. The arbitration node is primarily used to deploy database services for compute management nodes and virtual machines, including highly available distributed key-value stores (ETCD) and MySQL-based cluster solutions (MySQL-Garela). The front end uses tools for high availability and load balancing (Keepalived) and high-performance, open-source load balancers and proxy software (HAproxy) to facilitate inter-component communication, ensuring high availability of the compute management components and the virtual machine's database. The latency between the on-premises and standby hyperconverged sites and the arbitration node must be less than 2.5ms. One management server is configured to communicate with both the on-premises and standby hyperconverged sites, with a management network bandwidth of at least 100Mbps. This multi-node arbitration mechanism prevents situations where a single arbitration node fails and failover is not possible.

[0099] The arbitration node needs to be configured with a custom arbitration IP address. The host hosting the arbitration IP needs to deploy a high-availability component on a virtual machine. This high-availability component is used for fault detection and triggering service switching. In the event that all other steps fail, the arbitration mechanism can determine which hyperconverged site can take over the resources, preventing a split-brain scenario.

[0100] In some embodiments, different monitoring technologies provide different switching standards for business applications. The monitoring technology can be selected as needed, such as service monitoring, process monitoring, memory monitoring, CPU monitoring, custom script monitoring, and disk monitoring. Multiple high-availability rules can be provided for protection during monitoring, and the monitored object can be a local hyperconverged site, a backup hyperconverged site, or a site monitored simultaneously.

[0101] When a business requires support from multiple applications, a high availability group approach is adopted. The applications are grouped and sorted according to certain rules. When the source end of a certain autonomous group fails, the corresponding primary and backup switch can be completed to ensure the normal operation of the application, thereby ensuring the continuity of business.

[0102] In some embodiments, please refer to Figure 4 The disaster recovery backup method in another embodiment of the present invention includes the following steps: In step 301, the initial data is synchronized to the backup hyperconverged site, and incremental data is acquired during the synchronization of the initial data.

[0103] A communication connection is established between the local hyperconverged site and the backup hyperconverged site, and initial data is synchronously transmitted to the backup hyperconverged site. This allows for data backup transmission before a failure occurs, ensuring that at least one backup hyperconverged site can obtain a basic data sample and preventing data loss due to sudden failure of the local hyperconverged site. Understandably, during the initial data synchronization process, the local hyperconverged site continues to process business operations, thus generating some incremental data. This incremental data is then acquired during the data synchronization process.

[0104] In step 302, after the initial data synchronization is detected to be complete, the incremental data is synchronized to the backup hyperconverged site. Once all initial data has been synchronized to the corresponding hyperconverged site, the acquired incremental data is then transmitted. On the one hand, uploading incremental data while the initial data is being transmitted will consume the network channel for uploading the initial data, thus reducing the synchronization speed of the initial data. On the other hand, while the initial data and incremental data are being uploaded simultaneously, incremental data is still being generated continuously. In some embodiments, the incremental data may change. Therefore, synchronizing incremental data only after all the initial data has been synchronized can improve the utilization efficiency of network bandwidth and the effectiveness of the incremental data synchronization process, ensuring the consistency of disaster recovery backup.

[0105] In step 303, when a fault state is detected, a fault switching strategy is triggered. The fault switching strategy includes performing a cross-site migration operation, which is used to cooperate with the backup hyperconverged site to perform the migration strategy.

[0106] In step 303, if the local hyperconverged site experiences a failure, a failover strategy is triggered. It should be noted that the failure status of the local hyperconverged site includes various scenarios, and the corresponding failover strategies differ. Specifically, the failover strategy includes a cross-site migration operation, which is used to migrate all functions or data of the local converged site to a backup converged site. In this embodiment of the invention, the corresponding failover strategy is triggered and automatically executed based on the corresponding failure status, eliminating the need for manual judgment, thereby reducing delays and error risks, and effectively shortening service interruption time.

[0107] Step 303 above also includes the following sub-steps: Sub-step 311: Obtain multiple local image data blocks.

[0108] By dividing the image into data blocks, it is possible to break away from the dependence on the entire image file and instead operate on local image data blocks that can be independently identified and processed. This enables the migration task to be processed in parallel and resumed from breakpoints, changing the traditional mode of full image transmission.

[0109] Sub-step 312: Compare multiple local mirror data blocks with multiple data blocks of the backup fusion site to determine the local mirror data block that needs to be transmitted.

[0110] By comparing the data storage status with that of the backup hyperconverged site in advance, redundant data blocks and truly missing differential data blocks that already exist in the target backup hyperconverged site can be accurately identified. This reduces the amount of cross-site data transmission from the full size of the local mirror data to the difference in data status between the two sites. On cross-regional links with high network latency and limited bandwidth resources, this can significantly reduce migration time and network resource consumption, which is key to improving the efficiency of cross-site migration operations and ensuring rapid business takeover.

[0111] Sub-step 313: Filter out the local mirror data blocks that are inconsistent with the data blocks of the backup fusion site, and mark these local mirror data blocks as to be transmitted; By filtering out data blocks that are inconsistent with the backup hyperconverged site and marking them as to be transmitted, the local mirror data blocks that must be physically transmitted are accurately located, making the target of subsequent transmission operations clear and avoiding the invalid bandwidth occupation that may be caused by misjudgment or the inclusion of redundant data in traditional solutions. This marking, as a kind of metadata, provides the transmission scheduling engine with direct and unambiguous instructions, ensuring that network resources can be used centrally and efficiently to transmit truly different content.

[0112] Sub-step 314: Select the portion of the local mirror data blocks that are consistent with the data blocks of the backup fusion site, and mark this portion of the local mirror data blocks as reusable.

[0113] By identifying and marking local image data blocks identical to those of the standby hyperconverged site as reusable, existing local resources at the standby hyperconverged site are utilized. The reusable marker means that for this local image data block, no cross-network copying is required. When reconstructing virtual machines at the standby hyperconverged site, an identical copy already existing in its local storage can be directly referenced. This not only reduces transmission requirements and significantly alleviates network load, but more importantly, it makes the migration time proportional to the overall difference in data between the two sites, rather than the absolute size of the virtual machine image data block. This advantage is particularly significant in scenarios where the standby hyperconverged site has deployed similar operating systems or application templates, enabling recovery time targets in the seconds or minutes.

[0114] Sub-step 315 involves combining the local mirror data blocks marked as to be transmitted and transmitting them to the backup fusion site.

[0115] By combining and transmitting selected local mirror data blocks with differences, rather than mixing in a large amount of duplicate data, the network bandwidth can be used entirely for the synchronization of effective data, maximizing the utilization of the transmission process. This not only shortens the migration window and reduces the risk of business interruption, but also reduces the probability of transmission failure due to prolonged occupation of high-latency links, thereby improving the overall success rate of cross-site migration tasks.

[0116] By sensing and comparing the content of local image data blocks, intelligently marking and classifying them, and focusing on the transmission of differing data, cross-site virtual machine migration is no longer constrained by the overall size of the image data, but depends on the actual degree of difference between the data between the two hyperconverged sites. This significantly reduces the dependence of cross-regional backup switching on network resources and the time cost of the switching process itself, thereby greatly enhancing the actual availability of the entire disaster recovery backup method in the face of major failures.

[0117] In addition, the migration of the following data is also included: user data migration: user data is migrated using a snapshot method to shorten the migration time; memory data migration: after the user data migration is completed, the memory data is migrated in a pre-copy manner; resource switching: when the iteration copy meets the end threshold, the virtual machine stops running on the hyperconverged appliance in the local hyperconverged site and starts running on the hyperconverged appliance in the target standby hyperconverged site, and releases the resources of the hyperconverged appliance in the hyperconverged site.

[0118] Secondly, Figure 5 This is a structural block diagram of a disaster recovery backup device according to an embodiment of the present invention. Please refer to it. Figure 5 This invention provides a disaster recovery backup device, comprising: Synchronization module 401 is used to synchronize initial data to the backup hyperconverged site and acquire incremental data during the synchronization process of the initial data; The first detection module 402 is used to synchronize the incremental data to the backup hyperconverged site after detecting that the initial data synchronization is complete. The execution module 403 is used to trigger a fault switching strategy when a fault state is detected. The fault switching strategy includes performing a cross-site migration operation. The cross-site migration operation is used to cooperate with the backup hyperconverged site to perform the migration strategy.

[0119] In this embodiment of the invention, the disaster recovery backup method is applied to a local hyperconverged site. The local hyperconverged site includes multiple hyperconverged appliances and virtual machines. The hyperconverged appliance includes multiple compute nodes. The local hyperconverged site is communicatively connected to at least one standby hyperconverged site. The disaster recovery backup method first performs an initial full data synchronization, continuously acquiring incremental data during this period. This ensures that the standby hyperconverged site obtains a complete basic data copy, while minimizing the data loss window caused by the long duration of full synchronization. Subsequently, after the initial data synchronization is completed, the acquired incremental data is synchronized, thereby avoiding congestion that may be caused by simultaneous transmission of incremental and full data over the network. This improves network bandwidth utilization efficiency and the reliability of the incremental data synchronization process, ensuring the consistency and accuracy of the disaster recovery backup. When a fault condition is detected, a failover strategy including cross-site migration operations is automatically triggered. The complex migration process and business takeover process are implemented through preset automated operations, significantly reducing the delay and error risks introduced by manual diagnosis, decision-making, and operation in traditional disaster recovery solutions, thereby effectively shortening business interruption time. This invention addresses the issue by breaking down the data synchronization process into steps of transmitting initial data and then incremental data, and by programming the fault response mechanism. This mitigates the negative impact of network latency on real-time data synchronization and overcomes the migration complexity challenges posed by the lack of centralized shared storage in hyperconverged architectures. Ultimately, it achieves highly reliable data backup and highly available business takeover in a distributed environment.

[0120] Optionally, the synchronization module includes; A partitioning submodule is used to divide the initial data into multiple data blocks; The first synchronization submodule is used to traverse multiple data blocks and synchronize the multiple data blocks to the backup hyperconverged site; The incremental data submodule is used to acquire the incremental data generated during the synchronization of multiple data blocks.

[0121] Optionally, the disaster recovery backup device also includes: The second detection module is used to detect the data changes of multiple data blocks contained in the initial data; An encapsulation module is used to encapsulate the data changes of multiple data blocks into multiple events containing hash identifiers, wherein the data blocks and the hash identifiers correspond one-to-one, and the changes in the hash identifiers are used to characterize the data changes of the data blocks; The second synchronization module is used to synchronize multiple events to the backup hyperconverged site.

[0122] Optionally, the second synchronization module includes: A sorting unit is used to determine the order of the multiple events and write the multiple data blocks corresponding to the multiple events into the storage unit of the backup hyperconverged site according to the order of the events.

[0123] Optionally, the fault status includes the compute node fault, the virtual machine fault, and the network fault.

[0124] Optionally, the execution module includes: The judgment submodule is used to obtain the status information of the computing node and judge the fault status of the computing node based on the status information; The first migration submodule is used to perform a virtual machine migration operation when a failure of one of the compute nodes is detected. The virtual machine migration operation is used to migrate the memory state of the virtual machine to other normally operating compute nodes. The second migration submodule is used to perform the cross-site migration operation when multiple compute node failures are detected. The cross-site migration operation is used to migrate the virtual machine to the backup hyperconverged site.

[0125] Optionally, the judgment submodule includes: A data exchange unit is used to divide the multiple computing nodes into multiple autonomous groups, and the computing nodes in each autonomous group exchange heartbeat data at regular intervals. The heartbeat data is used to generate the status information of each computing node. The detection unit is used to determine the exchange status of the heartbeat data within a preset time. If the computing node does not receive the heartbeat data corresponding to the computing node, it determines that the corresponding computing node is in a fault state and reports the fault information.

[0126] Optionally, the execution module also includes: The network detection submodule is used to detect the network connectivity status between each computing node, including the management network connectivity status, the service network connectivity status, and the storage network connectivity status. The forwarding submodule is used to forward the management network connectivity status to other computing nodes through the storage network when the management network connectivity status is detected to be in a fault state. The first execution submodule is used to determine that the corresponding compute node is in a fault state when the service network connectivity status or the storage network connectivity status is detected to be faulty, and to execute the virtual machine migration operation.

[0127] Optionally, the execution module also includes: The restart submodule is used to control the restart of the virtual machine when the virtual machine is in a fault state; The second execution submodule is used to perform the virtual machine migration operation when the virtual machine is still in a fault state after restarting.

[0128] Optionally, the disaster recovery backup device further includes a third detection module, which is used to acquire multiple local mirror data blocks, compare the multiple local mirror data blocks with the multiple data blocks of the backup converged site, and determine the local mirror data blocks that need to be transmitted.

[0129] Optionally, the third detection module is further configured to filter out the local mirror data blocks that are inconsistent with the data blocks of the backup fusion site, and mark these local mirror data blocks as to be transmitted; and to filter out the local mirror data blocks that are consistent with the data blocks of the backup fusion site, and mark these local mirror data blocks as reusable.

[0130] Optionally, the execution module is further configured to combine the portions of the local mirror data blocks marked as to be transmitted and transmit them to the backup fusion site.

[0131] As the apparatus embodiment is basically similar to the method embodiment, it is described in a relatively simple manner. For relevant details, please refer to the description of the method embodiment.

[0132] This invention also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the various processes of the above-described disaster recovery backup method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0133] This invention also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the disaster recovery backup method embodiments described above and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0134] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0135] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0136] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, electronic device, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0137] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.

Claims

1. A disaster recovery backup method, characterized in that, The disaster recovery backup method is applied to a local hyperconverged site, which includes multiple hyperconverged appliances and virtual machines. The hyperconverged appliance includes multiple compute nodes. The local hyperconverged site is communicatively connected to at least one backup hyperconverged site. The disaster recovery backup method includes: The initial data is synchronized to the backup hyperconverged site, and incremental data is acquired during the synchronization of the initial data; After the initial data synchronization is completed, the incremental data is synchronized to the backup hyperconverged site. When a fault condition is detected, a fault switching strategy is triggered. The fault switching strategy includes performing a cross-site migration operation, which is used to cooperate with the backup hyperconverged site to perform the migration strategy.

2. The disaster recovery backup method according to claim 1, characterized in that, The step of synchronizing initial data to the backup hyperconverged site and acquiring incremental data during the synchronization of the initial data includes: The initial data is divided into multiple data blocks; Traverse multiple data blocks and synchronize multiple data blocks to the backup hyperconverged site; Acquire the incremental data generated during the synchronization of the multiple data blocks.

3. The disaster recovery backup method according to claim 1, characterized in that, Also includes: Detect the data changes of multiple data blocks contained in the initial data; The data changes of multiple data blocks are encapsulated into multiple events containing hash identifiers. The data blocks and the hash identifiers correspond one-to-one, and the changes in the hash identifiers are used to characterize the data changes of the data blocks. The events are synchronized to the backup hyperconverged site.

4. The disaster recovery backup method according to claim 3, characterized in that, The process of synchronizing multiple events to the backup hyperconverged site includes: The order of the events is determined, and the data blocks corresponding to the events are written into the storage unit of the backup hyperconverged site in the order of the events.

5. The disaster recovery backup method according to claim 1, characterized in that, The fault states include compute node failure, virtual machine failure, and network failure.

6. The disaster recovery backup method according to claim 5, characterized in that, When a fault state is detected, a fault switching strategy is triggered, including: Obtain the status information of the computing node, and determine the fault status of the computing node based on the status information; When a failure of one of the compute nodes is detected, a virtual machine migration operation is performed to migrate the memory state of the virtual machine to another normally functioning compute node. When multiple compute node failures are detected, the cross-site migration operation is performed to migrate the virtual machine to the backup hyperconverged site.

7. The disaster recovery backup method according to claim 6, characterized in that, The step of obtaining the status information of the computing node and determining the fault status of the computing node based on the status information includes: The computing nodes are divided into multiple autonomous groups, and the computing nodes in each autonomous group exchange heartbeat data at regular intervals. The heartbeat data is used to generate the status information of each computing node. Within a preset time period, the exchange status of the heartbeat data is determined. If the computing node does not receive the heartbeat data corresponding to the computing node, the corresponding computing node is determined to be in a fault state, and the fault information is reported.

8. The disaster recovery backup method according to claim 6, characterized in that, The method of triggering a fault switching strategy when a fault state is detected also includes: Detect the network connectivity status between each computing node, including the management network connectivity status, the service network connectivity status, and the storage network connectivity status; When the management network connectivity status is detected to be faulty, the management network connectivity status is forwarded to other computing nodes through the storage network; If a failure is detected in the connectivity status of the service network or the connectivity status of the storage network, the corresponding compute node is determined to be in a fault state, and the virtual machine migration operation is executed.

9. The disaster recovery backup method according to claim 6, characterized in that, The method of triggering a fault switching strategy when a fault state is detected also includes: When the virtual machine is in a faulty state, control the virtual machine to restart; If the virtual machine remains in a faulty state after restarting, the virtual machine migration operation is performed.

10. The disaster recovery backup method according to claim 1, characterized in that, The fault switching triggering strategy also includes: Retrieve multiple local image data blocks; By comparing multiple local mirror data blocks with multiple data blocks from the backup fusion site, the local mirror data block that needs to be transmitted is determined.

11. The disaster recovery backup method according to claim 10, characterized in that, Also includes: Filter out the local mirror data blocks that are inconsistent with the data blocks of the backup fusion site, and mark these local mirror data blocks as to be transmitted; Select the portion of the local mirror data blocks that are consistent with the data blocks of the backup fusion site, and mark this portion of the local mirror data blocks as reusable.

12. The disaster recovery backup method according to claim 11, characterized in that, Also includes: The local mirror data blocks marked as to be transmitted are combined and transmitted to the backup fusion site.

13. A disaster recovery backup device, characterized in that, include: The synchronization module is used to synchronize initial data to the backup hyperconverged site and acquire incremental data during the synchronization process of the initial data; The first detection module is used to synchronize the incremental data to the backup hyperconverged site after detecting that the initial data synchronization is complete. An execution module is used to trigger a fault switching strategy when a fault state is detected, the fault switching strategy including performing a cross-site migration operation; The cross-site migration operation is used to coordinate with the backup hyperconverged site to execute the migration strategy.

14. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the disaster recovery backup method as described in claims 1-12.

15. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the disaster recovery backup method as described in claims 1-12.