Data disaster recovery conversion method and system based on distributed storage, medium and product

By authenticating and chunking the source and target nodes of Ceph distributed storage system, data format conflicts between heterogeneous nodes are resolved, data disaster recovery conversion between heterogeneous nodes is realized, and data migration and backup are ensured.

CN120448195AActive Publication Date: 2025-08-08ANTUTE (BEIJING) TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510538267.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

In distributed storage systems, when data migration or disaster recovery switching between heterogeneous nodes is transmitted, it is difficult to achieve effective data disaster recovery conversion due to data format conflicts.

Method used

By authenticating and verifying the source distributed nodes and target distributed nodes, obtaining a mirror data list, determining the data block size based on the storage type, and sending the block data to the target node through the network transmission pipeline, format conversion and deletion pool filtering to realize data migration.

Benefits of technology

Ensure that the data block size and format meet the storage limits of the target nodes, realize data disaster recovery conversion between heterogeneous nodes, solve the problem of the 'protocol gap' between heterogeneous storage media, and support data migration and backup between heterogeneous nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448195A_ABST
    Figure CN120448195A_ABST
Patent Text Reader

Abstract

The invention discloses a data disaster recovery conversion method and system based on distributed storage, a medium and a product, and relates to the technical field of data disaster recovery conversion. Performing authentication verification on the source distributed node and the target distributed node; after the authentication verification is passed, acquiring a mirror data list of the to-be-migrated data from the source distributed node; determining the size of a data block based on the data storage type corresponding to the source distributed node and the data storage type corresponding to the target distributed node, and blocking each mirror data in the mirror data list based on the size of the data block; writing all the partitioned mirror image data into a network transmission pipeline, so that the network transmission pipeline sends all the partitioned mirror image data to a target distributed node; and performing format conversion, deduplication pool filtering and data writing on the received partitioned mirror image data through the target distributed node. According to the method, the system, the medium and the product disclosed by the invention, data disaster recovery conversion among heterogeneous nodes can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data disaster recovery conversion, and in particular relates to a data disaster recovery conversion method, system, medium and product based on distributed storage. Background Art

[0002] Data disaster recovery conversion refers to the process of migrating data from a primary storage system (or source node) to a backup storage system (or target node) to ensure data availability, consistency, and business continuity in a disaster scenario. In distributed storage systems, data disaster recovery is a core mechanism for ensuring service availability. Open-source distributed storage systems, such as Ceph, implement basic disaster recovery capabilities through technologies such as multi-copy replication and erasure coding.

[0003] However, in a distributed storage system, when data migration or disaster recovery switching is performed between heterogeneous nodes (such as different hardware architectures, storage media, Ceph versions or third-party storage systems), data format conflicts often make it difficult to achieve disaster recovery conversion of data between heterogeneous nodes.

[0004] Therefore, how to provide an effective solution to achieve data disaster recovery conversion between heterogeneous nodes in the Ceph distributed storage system has become a difficult problem that needs to be solved urgently in the existing technology. Summary of the Invention

[0005] The purpose of the present invention is to provide a data disaster recovery conversion method, system, medium and product based on distributed storage to solve the above-mentioned problems existing in the prior art.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a data disaster recovery conversion method based on distributed storage, which is used to implement data disaster recovery conversion between heterogeneous nodes in a Ceph distributed storage system, comprising:

[0008] Authenticate and verify the source distributed node and the target distributed node;

[0009] After authentication and verification are passed, the mirror data list of the data to be migrated is obtained from the source distributed node;

[0010] Determine a data block size based on a data storage type corresponding to a source distributed node and a data storage type corresponding to a target distributed node, and block each mirror data in the mirror data list based on the data block size;

[0011] Writing all the divided mirror data into a network transmission pipeline so that the network transmission pipeline sends all the divided mirror data to the target distributed node;

[0012] The target distributed node performs format conversion, deduplication pool filtering and data writing on the received divided mirror data to achieve data migration.

[0013] Based on the above disclosed content, the present invention authenticates the source distributed node and the target distributed node; after the authentication is passed, obtains the mirror data list of the data to be migrated from the source distributed node; determines the data block size based on the data storage type corresponding to the source distributed node and the data storage type corresponding to the target distributed node, and blocks each mirror data in the mirror data list based on the data block size; writes all the blocked mirror data to the network transmission pipeline so that the network transmission pipeline sends all the blocked mirror data to the target distributed node; and performs format conversion, deduplication pool filtering and data writing on the received blocked mirror data by the target distributed node to achieve data migration. In this way, it can be ensured that the size of the divided data block meets the storage limit of the target distributed node during data disaster recovery conversion, and the data format is the format of the storage type corresponding to the target distributed node, so that the data in the source distributed node of the Ceph distributed storage system can be migrated and backed up to the heterogeneous target distributed node, thereby achieving data disaster recovery conversion between heterogeneous nodes.

[0014] In one possible design, the authentication and verification of the source distributed node and the target distributed node includes:

[0015] Obtaining the first authentication information of the source distributed node and the second authentication information of the target distributed node entered by the user;

[0016] Performing authentication verification on the source distributed node and the target distributed node based on the first authentication information and the second authentication information;

[0017] The first authentication information includes the IP address, port PORT and access key of the source distributed node, and the second authentication information includes the IP address, port PORT and access key of the target distributed node.

[0018] In one possible design, after obtaining the mirror data list of the data to be migrated from the source distributed node, the method further includes:

[0019] Responding to user settings, determining the number of task threads and network transmission bandwidth for image migration;

[0020] A task list corresponding to the mirror data list is generated.

[0021] In a possible design, writing all the divided image data into a network transmission pipeline includes:

[0022] Starting at least one task in the task list according to the number of task threads;

[0023] Writing the divided mirror data corresponding to each task in the at least one task into a network transmission pipeline according to the network transmission bandwidth;

[0024] When any task is completed, check whether there is an unstarted task in the task list. If there is an unstarted task, start an unstarted task until there is no unstarted task in the task list.

[0025] In one possible design, after starting at least one task in the task list, the method further includes:

[0026] The metadata of the mirror data corresponding to the started task is obtained, and the metadata is recorded in the transaction log.

[0027] In one possible design, determining the data block size based on the data storage type corresponding to the source distributed node and the data storage type corresponding to the target distributed node includes:

[0028] If the data storage type corresponding to the source distributed node is the same as the data storage type corresponding to the target distributed node, the data block size is determined based on the size of the data block in the source distributed node;

[0029] If the data storage type corresponding to the source distributed node is different from the data storage type corresponding to the target distributed node, and the data storage type corresponding to the target distributed node is S3, the data block size is determined to be 4 MiB;

[0030] If the data storage type corresponding to the source distributed node is different from the data storage type corresponding to the target distributed node, and the data storage type corresponding to the target distributed node is local disk storage, the data block size is determined to be 16 MiB.

[0031] In one possible design, the target distributed node writes data by calling a storage write adapter.

[0032] In a second aspect, the present invention provides a data disaster recovery conversion system based on distributed storage, which is used to implement data disaster recovery conversion between heterogeneous nodes in a Ceph distributed storage system, including:

[0033] An authentication and verification unit, used to authenticate and verify the source distributed node and the target distributed node;

[0034] The acquisition unit is used to obtain the mirror data list of the data to be migrated from the source distributed node after the authentication verification is passed;

[0035] A determination unit, configured to determine a data block size based on a data storage type corresponding to a source distributed node and a data storage type corresponding to a target distributed node;

[0036] A block division unit, configured to divide each mirror data in the mirror data list into blocks based on the data block size;

[0037] a writing unit, configured to write all the divided mirror data into a network transmission pipeline, so that the network transmission pipeline sends all the divided mirror data to the target distributed node;

[0038] The data processing unit is used to perform format conversion, deduplication pool filtering and data writing on the received divided mirror data through the target distributed node to achieve data migration.

[0039] In a third aspect, the present invention provides a computer-readable storage medium having instructions stored thereon. When the instructions are executed on a computer, the data disaster recovery conversion method based on distributed storage as described in the first aspect or any possible design of the first aspect is executed.

[0040] In a fourth aspect, the present invention provides a computer program product comprising instructions, which, when executed on a computer, causes the computer to execute the data disaster recovery conversion method based on distributed storage as described in the first aspect or any possible design of the first aspect.

[0041] Beneficial effects:

[0042] The distributed storage-based data disaster recovery conversion method, system, medium and product provided by the present invention can ensure that the size of the divided data blocks meets the storage limit of the target distributed node during data disaster recovery conversion, and the data format is the format of the storage type corresponding to the target distributed node, so that the data in the source distributed node of the Ceph distributed storage system can be migrated and backed up to the heterogeneous target distributed node, realizing data disaster recovery conversion between heterogeneous nodes, which is convenient for practical application and promotion. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A flowchart of a data disaster recovery conversion method based on distributed storage provided in an embodiment of the present application;

[0044] Figure 2 This is a block diagram of a data disaster recovery conversion system based on distributed storage provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the present invention will be briefly introduced below in conjunction with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the structure of the drawings is only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.

[0046] It should be understood that although the terms "first," "second," etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element without departing from the scope of the exemplary embodiments of the present invention.

[0047] It should be understood that the term "and / or" that may appear in this document is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may indicate three situations: A exists alone, B exists alone, and A and B exist at the same time. The term " / and" that may appear in this document describes another type of association object relationship, indicating that two relationships may exist. For example, A / and B may indicate two situations: A exists alone, and A and B exist alone. In addition, the character " / " that may appear in this document generally indicates that the previous and subsequent associated objects are in an "or" relationship.

[0048] Example:

[0049] like Figure 1 As shown, the first aspect of this embodiment provides a data disaster recovery conversion method based on distributed storage, which can be applied to the Ceph distributed storage system to implement data disaster recovery conversion between heterogeneous nodes in the Ceph distributed storage system. Figure 1 As shown, the data disaster recovery conversion method based on distributed storage may include, but is not limited to, the following steps S101 to S105.

[0050] Step S101: Authenticate and verify the source distributed node and the target distributed node.

[0051] The source distributed node may be a node in the Ceph distributed storage system whose data needs to be migrated and backed up, and the target distributed node may be a node in the Ceph distributed storage system to which the data needs to be migrated.

[0052] When performing disaster recovery conversion on data between heterogeneous nodes in a Ceph distributed storage system, the user can enter the first authentication information of the source distributed node and the second authentication information of the target distributed node. At this time, the source distributed node and the target distributed node can be authenticated and verified based on the first authentication information and the second authentication information. The first authentication information may include, but is not limited to, the IP (Internet Protocol Address) address, port PORT (used to identify the logical channel of different services on the same device), and access key of the source distributed node, and the second authentication information may include, but is not limited to, the IP address, port PORT, and access key of the target distributed node.

[0053] Step S102: After the authentication verification is passed, a mirror data list of the data to be migrated is obtained from the source distributed node.

[0054] After authentication is successful, the distributed system can retrieve mirror data for all data from the source distributed node. At this point, the user can manually select the mirror data to be migrated and backed up. For ease of explanation, in this embodiment of the application, the selected mirror data to be migrated and backed up is referred to as the mirror data list.

[0055] It is understandable that if the user does not manually select the mirror data that needs to be migrated and backed up, the system may also default to using a portion of the mirror data or all of the mirror data as the mirror data that needs to be migrated and backed up.

[0056] In one or more embodiments, a user can manually set the number of task threads and network transmission bandwidth for image migration. In response to the user's setting operation, the number of task threads and network transmission bandwidth for image migration can be determined. Simultaneously, a task list corresponding to the image data list can be generated. Tasks in the task list can correspond one-to-one with the image data in the image data list, or a single task in the task list can correspond to multiple image data in the image data list. That is, each task corresponds to the migration and backup process of one or more image data.

[0057] Step S103: Determine the data block size based on the data storage type corresponding to the source distributed node and the data storage type corresponding to the target distributed node, and block each mirror data in the mirror data list based on the data block size.

[0058] Specifically, if the data storage type corresponding to the source distributed node is the same as the data storage type corresponding to the target distributed node, the data block size may be determined based on the size of the data block in the source distributed node.

[0059] If the data storage type corresponding to the source distributed node is different from the data storage type corresponding to the target distributed node, and the data storage type corresponding to the target distributed node is S3 (Amazon S3, Amazon Simple Storage Service) type, the data block size is determined to be 4 MiB.

[0060] If the data storage type corresponding to the source distributed node is different from the data storage type corresponding to the target distributed node, and the data storage type corresponding to the target distributed node is local disk storage, the data block size is determined to be 16 MiB.

[0061] In this way, the data block size can be dynamically adjusted to ensure that the size of the divided data block meets the storage limitations of the target distributed node.

[0062] Step S104: Write all the divided mirror data into a network transmission pipeline, so that the network transmission pipeline sends all the divided mirror data to the target distributed node.

[0063] Specifically, writing all the divided mirror data into the network transmission pipeline may include, but is not limited to, the following steps S1041-S1043.

[0064] Step S1041. Start at least one task in the task list according to the number of task threads.

[0065] For example, if the number of task threads is 16, then 16 tasks in the task list can be started.

[0066] Step S1042: Writing the divided mirror data corresponding to each task in the at least one task into a network transmission pipeline according to the network transmission bandwidth.

[0067] That is, the data transmission speed is controlled according to the network transmission bandwidth, and the blocked mirror data corresponding to each task in at least one task is written into the network transmission pipeline, so that the blocked mirror data corresponding to each task in at least one task can be transmitted to the target distributed node through the network transmission pipeline.

[0068] Step S1043: When any task is completed, check whether there is an unstarted task in the task list. If there is an unstarted task, start an unstarted task until there is no unstarted task in the task list.

[0069] When any task among all started tasks is completed, the task list may be checked to see if there is an unstarted task. If there is an unstarted task, an unstarted task is started until there is no unstarted task in the task list.

[0070] In one or more embodiments, after starting at least one task in the task list, the system may further obtain metadata of the mirror data corresponding to the started task, and record the metadata of the mirror data corresponding to the started task in a transaction log.

[0071] Step S105: The target distributed node performs format conversion, deduplication pool filtering, and data writing on the received divided mirror data to achieve data migration.

[0072] After sending the block-based mirror data to the target distributed node, the distributed system can perform format conversion on the received block-based mirror data through the target distributed node to convert it into a data format that is compatible with the data storage type corresponding to the target distributed node, and then perform dedupe pool ingress filtering on the format-converted mirror data (Dedupe Pool Ingress Filtering is a mechanism that filters or processes data according to predefined rules before the data enters a storage pool with dedupe / deduplicate functions. Its core purpose is to optimize storage efficiency, reduce unnecessary dedupe calculations, and ensure the security and compliance of key data). Finally, the target distributed node calls the storage write adapter to write the data obtained after dedupe pool filtration to achieve data migration. After the data is written, the transaction log can be submitted and the task list can be updated.

[0073] In one or more embodiments, before the target distributed node performs format conversion on the received segmented mirror data, verification data may be checked on the received segmented mirror data.

[0074] The present invention provides a data disaster recovery conversion method based on distributed storage. The method comprises the following steps: performing authentication and verification on a source distributed node and a target distributed node; obtaining a mirror data list of data to be migrated from the source distributed node after the authentication and verification are passed; determining a data block size based on the data storage type corresponding to the source distributed node and the data storage type corresponding to the target distributed node, and dividing each mirror data in the mirror data list based on the data block size; writing all divided mirror data into a network transmission pipeline so that the network transmission pipeline sends all divided mirror data to the target distributed node; and performing format conversion, deduplication pool filtering, and data writing on the received divided mirror data by the target distributed node to achieve data migration. In this way, during data disaster recovery conversion, the size of the divided data blocks can be ensured to meet the storage limit of the target distributed node, and the data format is the format of the storage type corresponding to the target distributed node, so that the data in the source distributed node of the Ceph distributed storage system can be migrated and backed up to the heterogeneous target distributed node, thereby solving the "protocol gap" problem between heterogeneous storage media, achieving data disaster recovery conversion between heterogeneous nodes, and facilitating practical application and promotion.

[0075] See also Figure 2 In a second aspect, an embodiment of the present application provides a data disaster recovery conversion system based on distributed storage, which is used to implement data disaster recovery conversion between heterogeneous nodes in a Ceph distributed storage system. The data disaster recovery conversion system based on distributed storage includes:

[0076] An authentication and verification unit, used to authenticate and verify the source distributed node and the target distributed node;

[0077] The acquisition unit is used to obtain the mirror data list of the data to be migrated from the source distributed node after the authentication verification is passed;

[0078] A determination unit, configured to determine a data block size based on a data storage type corresponding to a source distributed node and a data storage type corresponding to a target distributed node;

[0079] A block division unit, configured to divide each mirror data in the mirror data list into blocks based on the data block size;

[0080] a writing unit, configured to write all the divided mirror data into a network transmission pipeline, so that the network transmission pipeline sends all the divided mirror data to the target distributed node;

[0081] The data processing unit is used to perform format conversion, deduplication pool filtering and data writing on the received divided mirror data through the target distributed node to achieve data migration.

[0082] The working process, working details and technical effects of the distributed storage-based data disaster recovery conversion system provided in the second aspect of this embodiment can be found in the first aspect of the embodiment and will not be described in detail here.

[0083] A third aspect of this embodiment provides a computer-readable storage medium storing instructions for the data disaster recovery conversion method based on distributed storage as described in the first aspect of the embodiment. Specifically, the computer-readable storage medium stores instructions that, when executed on a computer, execute the data disaster recovery conversion method based on distributed storage as described in the first aspect. The computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash drive, and / or a memory stick. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable device.

[0084] The fourth aspect of this embodiment provides a computer program product containing instructions, which, when executed on a computer, enables the computer to execute the data disaster recovery conversion method based on distributed storage as described in the first aspect of the embodiment, wherein the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0085] It should be understood that certain details are provided in the following description to facilitate a thorough understanding of the example embodiments. However, one of ordinary skill in the art will appreciate that the example embodiments can be practiced without these specific details. For example, a system may be shown in block diagrams to avoid obscuring the example with unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail to avoid obscuring the example embodiments.

[0086] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.

Claims

1. A data disaster recovery conversion method based on distributed storage, used to implement data disaster recovery conversion between heterogeneous nodes in a Ceph distributed storage system, characterized in that: include: Authenticate and verify the source distributed node and the target distributed node; After authentication and verification are passed, the mirror data list of the data to be migrated is obtained from the source distributed node; Determine a data block size based on a data storage type corresponding to a source distributed node and a data storage type corresponding to a target distributed node, and block each mirror data in the mirror data list based on the data block size; Writing all the divided mirror data into a network transmission pipeline so that the network transmission pipeline sends all the divided mirror data to the target distributed node; The target distributed node performs format conversion, deduplication pool filtering and data writing on the received divided mirror data to achieve data migration.

2. The data disaster recovery conversion method based on distributed storage according to claim 1 is characterized in that: The authentication and verification of the source distributed node and the target distributed node includes: Obtaining the first authentication information of the source distributed node and the second authentication information of the target distributed node entered by the user; Performing authentication verification on the source distributed node and the target distributed node based on the first authentication information and the second authentication information; The first authentication information includes the IP address, port PORT and access key of the source distributed node, and the second authentication information includes the IP address, port PORT and access key of the target distributed node.

3. The data disaster recovery conversion method based on distributed storage according to claim 1, characterized in that: After obtaining the mirror data list of the data to be migrated from the source distributed node, the method further includes: Responding to user settings, determining the number of task threads and network transmission bandwidth for image migration; A task list corresponding to the mirror data list is generated.

4. The data disaster recovery conversion method based on distributed storage according to claim 3 is characterized in that: The step of writing all the divided image data into a network transmission pipeline includes: Starting at least one task in the task list according to the number of task threads; Writing the divided mirror data corresponding to each task in the at least one task into a network transmission pipeline according to the network transmission bandwidth; When any task is completed, the task list is checked to see if there is an unstarted task. If there is an unstarted task, an unstarted task is started until no unstarted task exists in the task list.

5. The data disaster recovery conversion method based on distributed storage according to claim 4 is characterized in that: After starting at least one task in the task list, the method further includes: The metadata of the mirror data corresponding to the started task is obtained, and the metadata is recorded in the transaction log.

6. The data disaster recovery conversion method based on distributed storage according to claim 1, characterized in that: The determining of the data block size based on the data storage type corresponding to the source distributed node and the data storage type corresponding to the target distributed node includes: If the data storage type corresponding to the source distributed node is the same as the data storage type corresponding to the target distributed node, the data block size is determined based on the size of the data block in the source distributed node; If the data storage type corresponding to the source distributed node is different from the data storage type corresponding to the target distributed node, and the data storage type corresponding to the target distributed node is S3, the data block size is determined to be 4 MiB; If the data storage type corresponding to the source distributed node is different from the data storage type corresponding to the target distributed node, and the data storage type corresponding to the target distributed node is local disk storage, the data block size is determined to be 16 MiB.

7. The data disaster recovery conversion method based on distributed storage according to claim 1 is characterized in that: The target distributed node writes data by calling a storage write adapter.

8. A data disaster recovery conversion system based on distributed storage, used to implement data disaster recovery conversion between heterogeneous nodes in a Ceph distributed storage system, characterized in that: include: An authentication and verification unit, used to authenticate and verify the source distributed node and the target distributed node; The acquisition unit is used to obtain the mirror data list of the data to be migrated from the source distributed node after the authentication verification is passed; A determination unit, configured to determine a data block size based on a data storage type corresponding to a source distributed node and a data storage type corresponding to a target distributed node; A block division unit, configured to divide each mirror data in the mirror data list into blocks based on the data block size; a writing unit, configured to write all the divided mirror data into a network transmission pipeline, so that the network transmission pipeline sends all the divided mirror data to the target distributed node; The data processing unit is used to perform format conversion, deduplication pool filtering and data writing on the received divided mirror data through the target distributed node to achieve data migration.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed on a computer, the data disaster recovery conversion method based on distributed storage according to any one of claims 1 to 7 is executed.

10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or the instruction is executed by a computer, the data disaster recovery conversion method based on distributed storage according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Heterogeneous data source standardized processing method and device and server

    CN108038239A

  • Data access method and device in distributed basic framework

    CN111694791A

  • Data processing method and data processing device

    CN115167773A

  • Distributed storage data migration method and system

    CN115167776A

  • Data migration method and device based on multi-type data sources

    CN116340295A