A multi-cloud platform-based data recovery method and system
By employing an intelligent data acquisition layer, an adaptive transmission protocol stack, and a heterogeneous system recovery architecture, the problem of low data migration efficiency between multiple cloud platforms is solved, enabling efficient and secure data migration and system recovery, improving system reliability and stability, and reducing disaster recovery time.
Patent Information
- Application Number
- CN202510923810.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-07-04
AI Technical Summary
Existing technologies cannot efficiently and securely migrate systems and data in multi-cloud scenarios. In particular, when the number of files is large and the file system hierarchy is complex, the transmission efficiency is significantly reduced, which greatly limits the practicality and convenience of migration.
It adopts an intelligent data acquisition layer based on a multimodal data recognition engine, dynamically adapts to four data acquisition channels, combines an adaptive transmission protocol stack and a heterogeneous system recovery architecture, realizes cross-platform data capture and transmission through a virtual environment abstraction layer, uses a cloud-native storage architecture for data management, and automates the recovery process through an automated recovery orchestration engine.
It enables efficient and secure migration of systems and data across multiple cloud platforms, improves data reading efficiency, reduces the consumption of source machine resources, supports the reliability and stability of system recovery on heterogeneous platforms, and the automated agentless function saves the consumption of source machine resources. It supports whole-machine disaster recovery on heterogeneous platforms, avoids human error and compatibility issues, and reduces the RTO of disaster recovery.
Smart Images

Figure CN120429169B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data recovery technology, specifically to a data recovery method and system based on a multi-cloud platform. Background Technology
[0002] With the rapid development of cloud computing technology, more and more enterprises and individuals are choosing to migrate their application systems and data to the cloud to take advantage of the powerful services and flexibility of cloud computing. However, for enterprises that have already deployed systems locally, or users who have business needs to migrate systems and data between different cloud platforms, effectively migrating local or existing public cloud resources to a new cloud platform has become a new challenge.
[0003] Existing solutions primarily rely on file transfer technology between hosts and image import technology from hosts to cloud platforms. However, file transfer technology between hosts has significant limitations, such as the ability to transfer data only on bidirectional peer-to-peer networks, the inability to handle unidirectional transfers from local office networks to public clouds, low transfer efficiency, and complex manual configuration processes. Therefore, existing technologies cannot meet the needs for efficient and secure system and data migration in multi-cloud scenarios, especially when dealing with large numbers of files and complex file system hierarchies, where transfer efficiency drops significantly, severely limiting the practicality and convenience of migration. Summary of the Invention
[0004] Technical problems to be solved:
[0005] To address the shortcomings of existing technologies, this invention provides a solution that addresses the problem that existing technologies cannot meet the need for efficient and secure migration of systems and data in multi-cloud scenarios, especially when the number of files is large and the file system hierarchy is complex, where the transmission efficiency is significantly reduced, greatly limiting the practicality and convenience of migration.
[0006] (II) Technical Solution:
[0007] To achieve the above objectives, the present invention provides the following technical solution:
[0008] A data recovery method based on a multi-cloud platform includes the following steps:
[0009] Step S1, Intelligent Data Acquisition Layer Construction: Based on a multimodal data recognition engine, dynamically adapting to four data acquisition channels:
[0010] A lightweight client agent module that supports cross-platform data capture in heterogeneous ARM / x86 environments;
[0011] VMware's agentless channel achieves accurate identification of changed data through dual VADP / CBT interfaces;
[0012] Cloud-native agentless mode, using cloud platform API gateway to achieve instance-level data extraction;
[0013] The offline cold data synchronization module extracts physical machine data based on a UEFI-booted PE environment.
[0014] Step S2, Adaptive Transport Protocol Stack Construction: Based on a real-time network quality assessment model, dynamically select:
[0015] Bidirectional symmetric transmission mode: Employs zero-copy technology to achieve efficient transmission of block-level data;
[0016] One-way secure tunnel mode: An encrypted reverse channel is built based on the TLS 1.3 protocol;
[0017] Hybrid transmission optimization engine: Supports bandwidth-adaptive multi-channel aggregation transmission;
[0018] Step S3, Heterogeneous System Recovery Architecture: Implemented through a virtual environment abstraction layer:
[0019] Hardware abstraction layer construction based on ACPI table parsing;
[0020] Drives automatic identification and injection engine;
[0021] Network configuration intelligent migration module;
[0022] System startup path reconstruction algorithm;
[0023] Step S4, Spatiotemporal Dimension Data Management Model: Implemented based on cloud-native storage architecture:
[0024] Timeline snapshot indexing system;
[0025] Incremental change capture engine;
[0026] Multi-level caching optimization strategy;
[0027] Intelligent stratification of hot and cold data;
[0028] Step S5, Automated Recovery Orchestration Engine: Implemented through a workflow engine:
[0029] Restore scene intelligent recognition;
[0030] Hardware compatibility pre-detection;
[0031] The recovery path is generated automatically;
[0032] Post-scripts are executed automatically.
[0033] Preferably, the adaptive transmission protocol stack constructed in step S2 dynamically selects either a bidirectional symmetrical transmission mode or a unidirectional secure tunnel mode through a real-time network quality assessment model, and achieves adaptive optimization of transmission parameters based on a bandwidth prediction algorithm, specifically including:
[0034] The network quality probe module collects packet loss rate, latency, and bandwidth parameters in real time.
[0035] The transmission mode decision engine selects the optimal transmission mode based on reinforcement learning algorithms.
[0036] The dynamic congestion control module supports multiple congestion control algorithms, including BBR and CUBIC.
[0037] The transmission encryption module supports AES-256 and ChaCha20 encryption algorithms.
[0038] Preferably, the heterogeneous system recovery architecture in step S3 achieves compatibility adaptation between different operating systems and hardware platforms through a virtual environment abstraction layer, specifically including:
[0039] The ACPI table parsing module extracts hardware configuration information.
[0040] The driver repository management module maintains multiple versions of the driver library;
[0041] The driver injection engine automatically matches and injects the appropriate driver based on hardware information.
[0042] The system starts the configuration refactoring module and regenerates the boot configuration file;
[0043] The network configuration migration module enables automatic configuration of IP addresses and DNS network parameters.
[0044] Preferably, the spatiotemporal dimension data management model in step S4, based on a cloud-native storage architecture, achieves efficient management and rapid recovery of data replicas, specifically including:
[0045] The timeline index module creates a unique index for each snapshot at each point in time.
[0046] The incremental change capture module identifies data changes based on block-level comparison technology.
[0047] The multi-level caching module employs a hybrid LRU and LFU caching strategy to improve access efficiency.
[0048] The hot and cold data tiering module automatically migrates data storage tiers based on access frequency.
[0049] Preferably, the automated recovery orchestration engine in step S5, through a workflow engine, automates the execution and monitoring of the recovery process, specifically including:
[0050] The scene recognition module automatically matches a recovery template based on the features of the source system.
[0051] The hardware compatibility detection module pre-detects the compatibility between the target environment and the source system;
[0052] The recovery path generation module generates the optimal recovery path based on a graph algorithm.
[0053] The execution monitoring module tracks the recovery process in real time and provides visual feedback.
[0054] The post-processing module supports the automated execution of custom scripts.
[0055] Preferably, the multimodal data recognition engine constructed in the intelligent data acquisition layer in step S1 can automatically select the optimal data acquisition channel according to the source environment type, specifically including:
[0056] The environmental detection module identifies the source system's operating system, hardware architecture, and virtualization platform;
[0057] The channel evaluation module assesses the suitability of each acquisition channel based on performance indicators and environmental characteristics.
[0058] The intelligent switching module supports dynamic switching of acquisition channels during data acquisition;
[0059] The data quality verification module ensures the integrity and consistency of the collected data.
[0060] A data recovery system based on a multi-cloud platform includes:
[0061] The intelligent data acquisition layer includes four data acquisition modules: client agent software, VMware agentless, cloud instance agentless module, and source machine offline data synchronization.
[0062] An adaptive transport layer enables two transmission modes: bidirectional symmetrical transmission and unidirectional secure tunneling.
[0063] Heterogeneous recovery engine, supporting cross-platform and cross-architecture system recovery;
[0064] The spatiotemporal data management module enables the creation, management, and restoration of data snapshots;
[0065] The automated orchestration center provides visual configuration and execution of the recovery process.
[0066] Preferably, the client agent software adopts a microkernel architecture design, specifically including:
[0067] The data acquisition engine supports file system filter drivers and volume shadow copy services;
[0068] The data processing engine enables data compression, encryption, and segmentation.
[0069] The communication engine supports breakpoint resume and transmission optimization;
[0070] The resource management module enables intelligent scheduling of system resources.
[0071] Preferably, the agentless module of the cloud instance acquires data through the cloud platform API, specifically including:
[0072] The cloud platform adapter layer provides a unified interface to adapt to APIs of different cloud platforms.
[0073] The snapshot management module enables the creation and management of cloud disk snapshots;
[0074] The difference analysis engine identifies incremental data based on block-level comparison technology;
[0075] The data transmission optimization module supports multi-threaded parallel transmission and bandwidth control.
[0076] (III) Beneficial Effects:
[0077] Compared with existing technologies, the present invention provides a data recovery method and system based on a multi-cloud platform, which has the following beneficial effects:
[0078] 1. In this invention, block-level technology is used for source data acquisition, which improves data reading efficiency and reduces the consumption of source machine resources; the faster incremental synchronization interval improves the RPO index from one day to minutes; the source agent software is compatible with mainstream Windows versions and many Linux distributions, and the software is developed with a customized PE system that supports startup on multiple cloud platforms and integrates data processing programs to realize information collection, data writing, and system conversion functions. After the target cloud host switches to different states, the task information is always retained and will not be lost due to the restart of the PE system.
[0079] 2. This invention deeply integrates software workflows with cloud platforms, enabling automated cloud resource calls to automatically match the specifications of the source production and create disaster recovery resources on the target platform according to template planning. This saves on tedious supporting operations, eliminates the possibility of human error, saves labor costs, and improves system reliability and stability. The automated, agentless function eliminates the process of installing clients on each machine at the source, shortening the deployment cycle, avoiding client compatibility issues, and eliminating the need for end customers to provide account passwords. Automated conversion makes whole-machine disaster recovery on heterogeneous platforms possible, and heterogeneous platform disaster recovery can also avoid security risks caused by security vulnerabilities on the same platform, significantly reducing the RTO of disaster recovery.
[0080] 3. This invention supports both one-way and two-way data transmission technologies. It can not only meet the data transmission needs of networks such as VPNs, leased lines, and intranets, but also handle one-way transmission scenarios such as office network to Internet and firewall rules restricting one-way access. It also has a flexible network transmission architecture. In the scenario of two-way transmission network, the source machine can directly transmit to the target cloud host. In the scenario where the source machine and the target cloud host cannot communicate, data can be relayed through a data gateway.
[0081] 4. This invention can not only copy and restore file data, but also copy and restore the entire operating system. The target system and data are consistent with the source machine. After the system is restored, the user can quickly start the application and restore business operations without performing too many operations.
[0082] 5. This invention improves data reading efficiency through block-level data synchronization, enabling efficient input or output operations. It eliminates the need to maintain redundant information such as file attributes and directory structures, allows for random access optimization, avoids frequent head seeks caused by file fragmentation through block addressing, has batch processing capabilities, can process continuous blocks of data in a single input or output, reducing the number of operations, and reduces CPU load by optimizing memory usage and reducing the number of data copies between kernel and user modes. It detects changes in storage blocks and only transmits modified blocks instead of the entire file, resulting in lower data transfer volume. It supports high-frequency copy synchronization, thereby improving resource utilization. Attached Figure Description
[0083] Figure 1 This is a flowchart illustrating a data recovery method based on a multi-cloud platform proposed in this invention.
[0084] Figure 2 This is a schematic diagram of the process structure of the target resource allocation method for data replication that is combined with a cloud platform to achieve automation, as proposed in this invention. Detailed Implementation
[0085] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0086] Example 1:
[0087] like Figures 1-2 As shown, a data recovery method based on a multi-cloud platform includes the following steps:
[0088] Step S1: Construction of intelligent data acquisition layer: Based on a multimodal data recognition engine, dynamically adapt to four data acquisition channels;
[0089] A lightweight client agent module that supports cross-platform data capture in heterogeneous ARM / x86 environments;
[0090] VMware's agentless channel achieves accurate identification of changed data through dual VADP / CBT interfaces;
[0091] Cloud-native agentless mode, using cloud platform API gateway to achieve instance-level data extraction;
[0092] The offline cold data synchronization module extracts physical machine data based on a UEFI-booted PE environment.
[0093] The multimodal data recognition engine constructed in the intelligent data acquisition layer in step S1 can automatically select the optimal data acquisition channel according to the source environment type, specifically including:
[0094] The environmental detection module identifies the source system's operating system, hardware architecture, and virtualization platform;
[0095] Let the source environment feature vector be... .
[0096] The channel evaluation module assesses the suitability of each acquisition channel based on performance indicators and environmental characteristics.
[0097] Let the collection channel set be set. Evaluation function for:
[0098] ;
[0099] in, These are the weighting coefficients. Environmental characteristics With acquisition channel The degree of compatibility.
[0100] The intelligent switching module supports dynamic switching of acquisition channels during data acquisition;
[0101] Let the switching condition function be defined. When this condition is met, from the channel Switch to channel .
[0102] The data quality verification module ensures the integrity and consistency of the collected data;
[0103] Set data integrity indicators The calculation formula is as follows:
[0104]
[0105] Data integrity is calculated by comparing the original data size with the collected data size. If the difference is within the allowable error range, the data is considered complete. Simultaneously, to ensure data consistency, a hash verification algorithm is used to validate the data. The data consistency index is defined as the hash verification pass rate, calculated using the following formula:
[0106] The number of data that passed the verification / the total number of data × 100%.
[0107] : Represents the data integrity index, and the value range is usually between 0 and 1. The closer it is to 1, the better the data integrity, reflecting the proportion of valid data after collection / transmission.
[0108] : Indicates the size of valid data (such as the amount of complete and undamaged data after collection), which is the statistical data volume that meets quality requirements.
[0109] This refers to the total amount of data that should theoretically be collected / transmitted; it represents the full target scale of data operations.
[0110] Step S2, Adaptive Transport Protocol Stack Construction: Based on a real-time network quality assessment model, dynamically select:
[0111] Bidirectional symmetric transmission mode: Employs zero-copy technology to achieve efficient transmission of block-level data;
[0112] One-way secure tunnel mode: An encrypted reverse channel is built based on the TLS 1.3 protocol;
[0113] Hybrid transmission optimization engine: Supports bandwidth-adaptive multi-channel aggregation transmission;
[0114] The adaptive transmission protocol stack constructed in step S2 dynamically selects either a bidirectional symmetrical transmission mode or a unidirectional secure tunnel mode through a real-time network quality assessment model, and achieves adaptive optimization of transmission parameters based on a bandwidth prediction algorithm, specifically including:
[0115] The network quality probe module collects parameters such as packet loss rate, latency, and bandwidth in real time.
[0116] Define network quality evaluation function The calculation formula is as follows:
[0117] ;
[0118] in;
[0119] Let be the packet loss rate at time t;
[0120] The delay is at time t;
[0121] The amount paid at time t;
[0122] , , These are the weighting coefficients, and + + =1.
[0123] The transmission mode decision engine selects the optimal transmission mode based on reinforcement learning algorithms.
[0124] The dynamic congestion control module supports multiple congestion control algorithms such as BBR and CUBIC.
[0125] The transmission encryption module supports encryption algorithms such as AES-256 and ChaCha20.
[0126] Step S3, Heterogeneous System Recovery Architecture: Implemented through a virtual environment abstraction layer:
[0127] Hardware abstraction layer construction based on ACPI table parsing;
[0128] Drives automatic identification and injection engine;
[0129] Network configuration intelligent migration module;
[0130] System startup path reconstruction algorithm;
[0131] The heterogeneous system recovery architecture in step S3 achieves compatibility adaptation between different operating systems and hardware platforms through a virtual environment abstraction layer, specifically including:
[0132] The ACPI table parsing module extracts hardware configuration information.
[0133] Let the hardware configuration information set be... ,in express Hardware attributes.
[0134] The driver repository management module maintains multiple versions of the driver library;
[0135] Driver library ,in Indicates the hardware type. Indicates the driver version.
[0136] The driver injection engine automatically matches and injects the appropriate driver based on hardware information.
[0137] Matching function Defined as:
[0138]
[0139] in, For hardware With Driver The compatibility score.
[0140] The system starts the configuration refactoring module and regenerates the boot configuration file;
[0141] By configuring the generation function Generate a new boot configuration file, in which... This is operating system information.
[0142] The network configuration migration module enables automatic configuration of network parameters such as IP address and DNS.
[0143] ,in, This represents subnet information, and M represents the MAC address.
[0144] Step S4, Spatiotemporal Dimension Data Management Model: Implemented based on cloud-native storage architecture:
[0145] Timeline snapshot indexing system;
[0146] Incremental change capture engine;
[0147] Multi-level caching optimization strategy;
[0148] Intelligent stratification of hot and cold data;
[0149] Step S4, the spatiotemporal dimension data management model, based on a cloud-native storage architecture, achieves efficient management and rapid recovery of data replicas, specifically including:
[0150] The timeline index module creates a unique index for each snapshot at each point in time.
[0151] The incremental change capture module identifies data changes based on block-level comparison technology.
[0152] The multi-level caching module employs a hybrid LRU and LFU caching strategy to improve access efficiency.
[0153] The hot and cold data tiering module automatically migrates data storage tiers based on access frequency.
[0154] Step S5, Automated Recovery Orchestration Engine: Implemented through a workflow engine:
[0155] Restore scene intelligent recognition;
[0156] Hardware compatibility pre-detection;
[0157] The recovery path is generated automatically;
[0158] Post-script execution is automated;
[0159] Step S5, the automated recovery orchestration engine, uses a workflow engine to automate the execution and monitoring of the recovery process, specifically including:
[0160] The scene recognition module automatically matches a recovery template based on the features of the source system.
[0161] Let the source system feature vector Restore template collection Matching function ;
[0162] in, The similarity score is the score between the source system features and the recovered template.
[0163] The hardware compatibility detection module pre-detects the compatibility between the target environment and the source system;
[0164] Compatibility score The calculation formula is:
[0165] ;
[0166] in, These are the weighting coefficients. and The first and second parts of the source system and the target system, respectively. Hardware attributes.
[0167] The recovery path generation module generates the optimal recovery path based on a graph algorithm.
[0168] Set up a recovery task graph Where V is the set of nodes, E is the set of edges, and the optimal path P is determined by Dijkstra's algorithm. Generate, where s is the starting node and t is the target node.
[0169] The execution monitoring module tracks the recovery process in real time and provides visual feedback.
[0170] Set recovery progress The calculation formula is as follows:
[0171] ;
[0172] The post-processing module supports the automated execution of custom scripts;
[0173] Set a set of custom scripts Execute function via script Execute the script.
[0174] A data recovery system based on a multi-cloud platform includes:
[0175] The intelligent data acquisition layer includes four data acquisition modules: client agent software, VMware agentless, cloud instance agentless module, and source machine offline data synchronization.
[0176] The agentless module of the cloud instance acquires data through the cloud platform API, specifically including:
[0177] The cloud platform adapter layer provides a unified interface to adapt to APIs of different cloud platforms.
[0178] The snapshot management module enables the creation and management of cloud disk snapshots;
[0179] The difference analysis engine identifies incremental data based on block-level comparison technology;
[0180] The data transmission optimization module supports multi-threaded parallel transmission and bandwidth control.
[0181] An adaptive transport layer enables two transmission modes: bidirectional symmetrical transmission and unidirectional secure tunneling.
[0182] Heterogeneous recovery engine, supporting cross-platform and cross-architecture system recovery;
[0183] The spatiotemporal data management module enables the creation, management, and restoration of data snapshots;
[0184] The automated orchestration center provides visual configuration and execution of the recovery process.
[0185] Preferably, the client agent software adopts a microkernel architecture design, specifically including:
[0186] The data acquisition engine supports file system filter drivers and volume shadow copy services;
[0187] The data processing engine enables data compression, encryption, and segmentation.
[0188] The communication engine supports breakpoint resume and transmission optimization;
[0189] The resource management module enables intelligent scheduling of system resources.
[0190] Example 2:
[0191] A data recovery method based on a multi-cloud platform includes the following steps:
[0192] Step 1: Data Acquisition. Using block-level data read and write technology, four data acquisition methods were developed: client agent software, VMware agentless, cloud instance agentless module, and source machine offline data synchronization.
[0193] This technology directly manipulates the physical or logical blocks of storage devices, bypassing the file system hierarchy and reducing the following overheads: Metadata operations: No need to maintain redundant information such as file attributes and directory structures; Random access optimization: Block-based addressing avoids frequent head seeks caused by file fragmentation; Batch processing capability: A single input or output can process contiguous block data, reducing the number of operations. For example, when backing up a 1GB file, traditional methods require thousands of small file operations, while block-level operations only require dozens of batch transfers; Improved resource utilization and optimized memory usage: Block-level read / write can precisely control the buffer size, such as aligning to 4... K / 8K blocks avoid uncontrollable memory consumption due to file caching; reduced CPU load: fewer data copying operations between kernel and user modes; lower data transfer volume; synchronize only changed blocks: detect changes in storage blocks, such as 4KB-64MB, and only transfer modified blocks, not the entire file; millisecond-level incremental detection for rapid change location: quickly compare the state of source and target blocks using hashes and bitmaps; copy-on-write, recording data block modification logs; supports high-frequency copy synchronization; unaffected by frequent updates, even if the file is modified hundreds of times per hour, block-level synchronization can still be processed efficiently.
[0194] This invention improves data reading efficiency through block-level data synchronization, enabling efficient input or output operations. It eliminates the need to maintain redundant information such as file attributes and directory structures, optimizes random access, and avoids frequent disk head seeks caused by file fragmentation through block-based addressing. It features batch processing capabilities, handling consecutive blocks of data in a single input or output, reducing the number of operations. Memory usage optimization reduces the number of data copies between kernel and user modes, lowering CPU load. It detects changes in storage blocks, transmitting only modified blocks rather than the entire file, resulting in lower data transfer volumes. It supports high-frequency copy synchronization, thereby improving resource utilization. This software's block-level data processing technology, through underlying architectural innovation and dynamic block adjustment strategies, achieves comprehensive breakthroughs in performance, flexibility, and resource management. Whether for large-scale data migration, real-time backup, or high-performance computing scenarios, it provides efficiency and stability far exceeding traditional file synchronization solutions.
[0195] Step 2, Data Transmission: Supports two transmission connection modes: peer-to-peer and one-way connections. It supports various network scenarios including VPNs, leased lines, and the internet. Technically, it supports both bidirectional and one-way transmission. In bidirectional transmission, bidirectional network access is required between the two nodes. After the TCP connection is established, both communicating parties have independent send and receive buffers, allowing simultaneous data transmission and reception. Data separation is achieved by using port numbers, IP addresses, and sequence numbers to distinguish data streams in different directions. In one-way transmission, only one-way network access is required between the two nodes. The receiver does not need a reverse link; an HTTPS tunnel is used to establish a reverse transmission tunnel, enabling one-way transmission between the source and destination. Architecturally, it supports data replication in various scenarios.
[0196] Source → Network 1 → Data Gateway Node → Network 2 → Data Write Node;
[0197] Source → Network 1 → Data writing node.
[0198] Step 3: Data replication task establishment. The target host boots into PE mode using a customized image. Data processing programs run within the PE environment. During data replication, the target host maintains PE mode. During recovery, a system conversion process occurs, and the virtual machine restarts to access the target system. The data replication process requires automated target resource configuration via a cloud platform. The specific process is as follows: Synchronize task configuration, initiate synchronization, create a cloud host on the target cloud platform using the customized image, determine if creation was successful (report an error if it fails), stop the task, and if successful, read the cloud host information, register it to the management console, create a data disk, and mount it to the cloud host. The program within the cloud host identifies the disk and matches it with the cloud platform record, then matches it with the source machine to establish a complex data task, ultimately completing the process. This invention can not only replicate and restore file data but also replicate and restore the entire operating system. The restored target system and data remain consistent with the source machine. After system recovery, users can quickly start applications and resume business operations without extensive operations.
[0199] Step 4: Data Historical Copy Retention and Management. Utilizing the cloud platform's built-in disk snapshot function, a snapshot is created at the corresponding point in time after each synchronization, enabling the retention of historical copies. The main tasks of data copy management include: automatically creating a disk snapshot of the target cloud host after each synchronization according to the user-defined strategy; automatically maintaining the number of snapshots to retain on the target cloud host based on the user-defined configuration; rolling back the cloud host to the user-selected snapshot point in conjunction with the recovery function; and creating a cloud host using the user-selected snapshot in conjunction with the recovery function to complete the recovery process.
[0200] Step 5: Operating System and Data Recovery. This step addresses the issue of heterogeneous system startup, supports automated network card configuration after system recovery, and allows for user-defined execution scripts. Through process and architecture integration, it achieves diverse recovery capabilities, including: target cloud host state switching; system conversion; recovery mode 1: recovery using the target host; recovery mode 2: recovery using historical points; recovery mode 3: test recovery using snapshots. During different recovery processes, the target cloud host needs to switch states in conjunction with the current task. This switching process requires the integration of software processes and cloud platform functions. In replication and synchronization mode, the system within the target cloud host is a custom image system created by this software for receiving and processing data. In recovery mode, the system within the target cloud host undergoes system conversion, completing driver replacement, network card settings, and other related configurations, and boots into the user's operating system. This operating system and internal data are consistent with the user's source machine. By calling the cloud platform's snapshot rollback capability, the cloud host can be restored to the custom image system created by this software, allowing data replication tasks to continue or other recovery tasks to be re-executed.
[0201] It should be noted that after the user confirms the replication task configuration, the entire process will be automated by scheduling relevant resources from the source machine, cloud platform, and cloud host through the management console. The specific process includes:
[0202] a. The user determines the replication task configuration and initiates the task; the user needs to determine the target cloud platform, image, cloud platform network, security group, and storage-related resource configuration; the software automatically creates cloud hosts on the cloud platform according to these configurations; the parameters required will vary depending on the cloud platform.
[0203] b. Based on the user-selected configuration, the software automatically creates a cloud host by calling the cloud platform API;
[0204] c. Obtain cloud host status information, including start / stop status, network information, and disk information, through the cloud platform API;
[0205] d. Register the target cloud host to the management console and obtain internal information of the cloud host, including network information and disk information;
[0206] e. Match the information within the virtual machine with the information on the cloud platform;
[0207] f. Based on the matching information, create a replication task between the source machine and the target cloud host, and initiate a full synchronization. The full synchronization requires copying all data from the source machine to the target host.
[0208] G. After full synchronization is completed, incremental synchronization tasks are initiated periodically according to the user-configured strategy. This invention supports both one-way and two-way data transmission technologies. It can not only meet the data transmission needs of networks such as VPNs, dedicated lines, and intranets, but also handle one-way transmission scenarios such as office network to Internet and firewall rules restricting one-way access. It also has a flexible network transmission architecture. In the scenario of two-way transmission network, the source machine can directly transmit to the target cloud host. In the scenario where the source machine and the target cloud host cannot communicate, data can be relayed through a data gateway.
[0209] This invention also proposes a data recovery system based on a multi-cloud platform, including four methods: client agent software, VMware agentless, cloud instance agentless module, and source machine offline data synchronization. This invention deeply integrates software workflows with the cloud platform, automating cloud resource calls to automatically match the specifications of the source production and create disaster recovery resources on the target platform according to template planning. This saves on tedious setup operations, eliminates the possibility of human error, saves labor costs, and improves system reliability and stability. The automated agentless function eliminates the process of installing clients on each source machine, shortening the deployment cycle and avoiding client compatibility issues. The end customer needs to provide their account and password. Automated driver conversion makes whole-machine disaster recovery possible for heterogeneous platforms. Heterogeneous platform disaster recovery can also avoid security risks caused by security vulnerabilities on the same platform, significantly reducing the RTO of disaster recovery. The data acquisition method of the client agent software is to install the agent software on the source machine in scenarios such as physical machine environments or single-machine scenarios such as cloud platform virtual machines. By installing the agent software, data acquisition is completed directly within the operating system. The data is packaged and transmitted according to the software's unified data structure, and it supports compression and encryption of transmitted data; data transmission is carried out through the TCP / IP protocol; the client software supports running in ARM and x86 computing environments. The client software supports most mainstream Windows and Linux distributions. VMware's agentless data acquisition method utilizes the VADP and CBT interfaces provided by the VMware virtualization platform to acquire data from virtual machines, eliminating the need to install agent software within the virtual machine. The cloud instance agentless module acquires data by using a disk created from a snapshot of the source machine, mounted to the software's data reading server, and then reading and synchronizing the source machine's data, provided the cloud platform supports it. When incremental difference transmission is required, a full disk verification is performed on the read data disk; after verification, only the differences between the source and target data are transmitted. The offline data synchronization method involves using the source-end reading PE provided by this software to boot the source machine into PE state. The PE runs the data reading program of this software, and the source machine data is synchronized offline via TCP / IP. The automated data and recovery method based on a multi-cloud platform in this invention combines a multi-layered and novel system architecture of data acquisition, data transmission, and data writing. Unlike the traditional file transfer mode, the architecture innovatively combines the operation inside the cloud host with the operation of the external cloud platform. When performing tasks such as data copying and recovery, coordinated interaction can be achieved, and the entire process can be automated through the unified scheduling of the management console server.
[0210] This software supports full system recovery for any P2V or V2V platform, system conversion for heterogeneous platforms, and platform adaptation for multi-cloud disaster recovery, involving three aspects.
[0211] Drive conversion:
[0212] Physical machines typically use SAS drivers, while virtual machines use different drivers such as MPTSAS, virtio, Hyper-V, and Xen, depending on the type of Hypervisor. This software system can implant the corresponding driver for the target platform before the target machine starts up, ensuring that its bus, disk, network card, and other devices can be loaded and booted normally.
[0213] Boot type conversion; there are two types of operating system boot modes: Legacy and UEFI.
[0214] Legacy boot mode, as the name suggests, is the traditional boot mode, a common boot mode before Windows 8. Legacy boot mode has good system compatibility, running on both 32-bit and 64-bit systems. Legacy mode uses the MBR disk format, and its system can only be installed on MBR formatted disks; it only supports a maximum of 4 primary partitions; it does not support hard drives larger than 2TB; and a single MBR can only store the boot record of one operating system.
[0215] UEFI boot mode is the successor to Legacy and is currently the mainstream boot mode. Compared to Legacy, UEFI has better programmability, scalability, performance, and security. Windows systems have supported UEFI since WIN7, and WIN8 uses UEFI by default. UEFI mode uses the GPT disk format (GUID partition table). UEFI has no limitations on the number of partitions or disk size, provides secure boot functions, and prevents viruses from loading during boot. The software supports similar boot methods on the target machine and boot conversion.
[0216] Network adapter settings:
[0217] Depending on the hardware drivers and network architecture, the network card name and IP address of the target cloud host may differ from those of the source machine. Therefore, this software system will inject new network card information and IP address into the target system, supporting DHCP and Static depending on the platform, to meet the needs of user recovery scenarios.
[0218] Recovery Mode 1: Recover using the target host;
[0219] In this recovery mode, the target cloud host is operated directly. During this process, there are relatively heavy disk I / O operations such as invalid data rollback and disk creation. Therefore, no matter how large the data volume of the target cloud host is, the operating system and data can be restored within a few minutes, minimizing RTO and providing users with the highest switching efficiency.
[0220] Key steps of the specific process: The user defines a custom recovery task, setting parameters such as the size of the restored cloud host instance, network interface card configuration, post-processing scripts, and security groups; the replication task is interrupted, preventing subsequent replication tasks from being initiated to avoid errors caused by re-initiating synchronization after recovery; based on the user's task configuration, a decision is made on whether to create a rollback snapshot; system conversion is performed; the cloud host is shut down; the relevant parameters set by the user are applied to the cloud host configuration via the cloud platform API, including but not limited to instance specifications, network interface card configuration, security groups, and elastic IPs; the cloud host is started, and the cloud host enters the user's target system; after the cloud host starts in the target system, it performs one operation, sending a startup success message to the management console and executing the user-set post-processing scripts; to accommodate complex and varied real-world application scenarios and to enhance the integrity of the overall recovery mode, a re-execution process is designed. The user must have enabled "rollback snapshot" to run the following process: the user decides to re-execute the recovery task; the target cloud host is shut down; the cloud host is rolled back to its pre-recovery state using a snapshot; the target host is started; the management console confirms the target host's connectivity; the re-execution process is complete.
[0221] Recovery Mode 2: Restore using history points;
[0222] In disaster recovery scenarios, because the data on the target host is synchronized with the source host, file errors existing on the source host will also be synchronized to the target host, such as logical errors in the database or files tampered with after a ransomware attack. In this case, it is necessary to restore the user system and data through historical copies. Since different cloud platforms have different snapshot capabilities, there are two main process designs for recovery mode two: one is to roll back the target cloud host to a specific point in time, and the other is to create a new cloud host by taking a snapshot of the target cloud host.
[0223] Key steps of the specific process: The user defines a custom recovery task, setting parameters such as the size of the restored cloud host instance, network interface card configuration, post-processing scripts, and security groups; the replication task is interrupted, preventing subsequent replication tasks from being initiated to avoid errors caused by re-initiating replication synchronization after recovery; a rollback snapshot is created; the cloud host is shut down; based on the user's task configuration, the target cloud host is rolled back to the selected point in time; system conversion is performed; the relevant parameters set by the user are applied to the cloud host configuration via the cloud platform API, including but not limited to instance specifications, network interface card configuration, security groups, and elastic IPs; the cloud host is started, and the cloud host enters the user's target system; after the cloud host starts in the target system, it performs one operation, sending a startup success message to the management console and executing the user-set post-processing scripts; similarly, to handle complex and varied real-world application scenarios and enhance the integrity of the overall recovery mode, a re-execution process is designed: the user decides to re-execute the recovery task; the target cloud host is shut down; the cloud host snapshot is rolled back to its pre-recovery state; the target host is started; the management console confirms the target host's connectivity; the re-execution process is complete.
[0224] Recovery Mode 3: Use snapshots to perform test recovery;
[0225] In complex and ever-changing real-world application scenarios, users may want to test and verify target cloud servers or perform disaster recovery drills without interrupting periodic replication and synchronization. To address this need, the software is designed with a snapshot recovery testing mode. This mode differs from Mode 2 in that it does not interrupt the replication and synchronization task. The key steps are as follows: The user defines a recovery task, setting parameters such as the restored cloud server instance size, network interface card configuration, post-processing scripts, and security groups; A new cloud server is created using the selected snapshot point by utilizing the cloud platform's snapshot capabilities; The user-defined cloud server configuration is applied, including but not limited to instance size, network interface card configuration, security groups, and elastic IPs; The cloud server is started; System conversion is performed; The cloud server is restarted, and it enters the user's target system; After booting into the target system, the cloud server performs one operation: sending a successful startup message to the management console and executing the user-defined post-processing scripts.
[0226] To enhance the integrity of the overall recovery model, the process for deleting a recovery task is as follows: The user decides to delete the recovery task; a command to delete the corresponding cloud host is sent to the cloud platform; deletion is complete.
[0227] This invention improves data reading efficiency and reduces the consumption of source machine resources by using block-level technology for source data acquisition; faster incremental synchronization intervals reduce the RPO (Recovery Point Objective) from one day to minutes; the source agent software is compatible with mainstream Windows versions and numerous Linux distributions; the software is compatible with ARM and x86 platforms; the developed bidirectional and unidirectional transmission technologies, along with the designed and implemented flexible transmission architecture, meet the needs of diverse network scenarios; a customized PE system is developed to support startup on multiple cloud platforms and integrates data processing programs to achieve information collection, data writing, and system conversion functions; task information is always preserved after the target cloud host switches between different states, and task information is not lost due to PE system restarts; the system conversion program is compatible with mainstream Windows versions and numerous Linux distributions and matches drivers for different cloud platforms; the software process is combined with cloud platform APIs, and the operation process is automated through the matching of virtual machine internal resources with cloud platform layer resources; enabling the software to achieve diverse and rich data recovery capabilities.
[0228] In the description herein, it should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
Claims
1. A data recovery method based on a multi-cloud platform, characterized in that, Includes the following steps: Step S1: Construct an intelligent data acquisition layer: The multimodal data recognition engine is activated to identify the operating system, hardware architecture, and virtualization platform of the source system, generating source environment feature vectors. Based on performance indicators and environment feature vectors, the following four data acquisition channels are evaluated: client agent channel, VMware agentless channel, cloud instance agentless channel, and offline cold data synchronization agent channel. The intelligent switching module monitors the dynamic switching of acquisition channels during the data acquisition process, and the data quality verification module verifies the consistency and integrity of the data to ensure that the acquired data is qualified. Step S2: Construct an adaptive transport protocol stack: The network quality real-time assessment model is launched to collect packet loss rate, latency, and bandwidth parameters in real time to assess network quality. Then, the transmission mode decision engine selects the optimal transmission mode, including bidirectional symmetrical transmission mode or unidirectional secure tunnel mode, based on the assessed network quality. During data transmission, the dynamic congestion control module switches between multiple congestion control algorithms such as BBR and CUBIC to avoid data congestion, and the transmission encryption module encrypts the transmitted data using AES-256 and ChaCha20 encryption algorithms. Step S3: Build a heterogeneous system recovery architecture: First, the virtual environment abstraction layer calls the ACPI table to parse the table and generate hardware configuration information. Then, the driver repository management module calls the driver repository to match the driver that matches the hardware information. The driver injection engine injects the matched driver. Then, the network configuration migration module performs path reconstruction based on the network architecture, configured IP address, and DNS network parameters to generate the boot configuration file, ensuring that the target system can be used normally. Step S4: Construct a spatiotemporal dimension data management model: First, a timeline indexing system is built to create a unique index for system data collection. Then, the incremental change capture module identifies data changes based on block-level comparison technology, processes the changed data, and applies a multi-level caching module to optimize the data access path and dynamically allocate data cache space at multiple levels. Finally, the hot and cold data stratification module adjusts the access frequency of cached data priority and automatically migrates the data storage level. Step S5: Build an automated recovery orchestration engine: First, the workflow engine initiates the recovery scene recognition module to extract the similarity score between the source system features and the recovery template. Then, the hardware compatibility detection module... The compatibility score between the pre-detected target environment and the source system is calculated. The recovery path generation module then generates the optimal recovery path based on the compatibility score. During the path recovery process, the execution monitoring module provides real-time statistics on the path recovery process to achieve visual feedback. Finally, the post-processing module automatically executes the script based on a custom script.
2. The data recovery method based on a multi-cloud platform according to claim 1, characterized in that, In step S12, a set of acquisition channels is set. Evaluation function for: in, These are the weighting coefficients. Environmental characteristics With acquisition channel Adaptability; In step S13, a switching condition function is set. When this condition is met, from the channel Switch to channel ; In step S14, a data integrity index is set. The calculation formula is as follows: In the formula, : Represents a data integrity indicator; : Indicates the amount of valid data; : Indicates the total amount of data.
3. The data recovery method based on a multi-cloud platform according to claim 1, characterized in that, In step S21, the network quality assessment function The calculation formula is as follows: in, Let be the packet loss rate at time t; The delay is at time t; Let be the bandwidth at time t; , , These are the weighting coefficients, and + + =1.
4. The data recovery method based on a multi-cloud platform according to claim 1, characterized in that, In step S31, the hardware configuration information set is set. ,in Represents n hardware attributes; In step S32, the driver library ,in Indicates the hardware type. Indicates the driver version; Matching function Defined as: in, For hardware With Driver Compatibility score; The step S33, ,in, This refers to subnet information, where M is the MAC address; By configuring the generation function Generate a new boot configuration file, in which... This is operating system information.
5. The data recovery method based on a multi-cloud platform according to claim 1, characterized in that, Step S51, setting the source system feature vector Restore template collection , Matching function ; in, The similarity score between the source system features and the recovered template; Step S52, compatibility score The calculation formula is: in, These are the weighting coefficients. and The first and second parts of the source system and the target system, respectively. One hardware attribute; Step S53, Set up a recovery task graph Where V is the set of nodes, E is the set of edges, and the optimal path P is determined by Dijkstra's algorithm. Generate, where s is the starting node and t is the target node; Step S54: Set the recovery progress The calculation formula is as follows: ; Step S55: Set a custom script set Execute function via script Execute the script.
6. A data recovery system based on a multi-cloud platform, implementing the data recovery method based on a multi-cloud platform as described in claim 1, characterized in that, include: The intelligent data acquisition module includes four data acquisition modules: client agent module, VMware agentless module, cloud instance agentless module, and source machine offline data module synchronization. The adaptive transmission module enables both bidirectional symmetrical transmission and unidirectional secure tunnel transmission modes; the heterogeneous recovery engine supports cross-platform and cross-architecture system recovery. The spatiotemporal data management module enables the creation, management, and restoration of data snapshots; The automated orchestration center provides visual configuration and execution of the recovery process.
7. A data recovery system based on a multi-cloud platform according to claim 6, characterized in that, The client proxy module adopts a microkernel architecture design, specifically including: The data acquisition module supports file system filter drivers and volume shadow copy services; The data processing module performs data compression, encryption, and segmentation. The communication module supports breakpoint resume and transmission optimization; The resource management module enables intelligent scheduling of system resources.
8. A data recovery system based on a multi-cloud platform according to claim 7, characterized in that, The agentless module of the cloud instance acquires data through the cloud platform API, specifically including: The cloud platform adapter module provides a unified interface to adapt to different cloud platform APIs; The snapshot management module enables the creation and management of cloud disk snapshots; The difference analysis module identifies incremental data based on block-level comparison technology; The data transmission optimization module supports multi-threaded parallel transmission and bandwidth control.
Citation Information
Patent Citations
Data migration system and method
CN117111836A
Synchronous acquisition error detection method and system of multi-channel dynamic data acquisition system
CN120028043A