Data replication method and device, electronic equipment and readable storage medium

By receiving and processing hash value collection to determine the difference data and copying it directly into the snapshot, the problem that the slave volume cannot be used normally during remote replication of snapshots is solved, and the data volume is reduced and the normal use of the volume is achieved.

CN120075241APending Publication Date: 2025-05-30RUIJIE NETWORKS CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311621291.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

During snapshot remote replication, the volume in the slave electronic device cannot be used normally because the volume is not allowed to be opened when the master electronic device replicates and updates the data into the slave electronic device's volume.

Method used

By receiving the first set of hash values ​​and the first snapshot information sent by the second electronic device, the difference data from the first set of data in the second data set is determined based on these sets of hash values, and the difference data is directly copied into the snapshot in the second electronic device without first copying it to the volume of the slave electronic device.

Benefits of technology

This enables no need to copy the entire data set when replicating the snapshot remotely, reduces the total amount of data copied, and ensures that the volumes in the slave electronic devices can be used normally.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075241A_ABST
    Figure CN120075241A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data replication method and device, electronic equipment and a readable storage medium, and relates to the technical field of communication. The method is applied to first electronic equipment and comprises the steps that a first hash value set and first snapshot information sent by second electronic equipment are received, the first hash value set corresponds to a first data set, and the first data set is a set of duplicated data of data in a first data source of the first electronic equipment in the second electronic equipment; the first snapshot information is used for indicating a first snapshot of a first data set created by the second electronic equipment; on the basis of the first hash value set and the second hash value set, determining difference data in the second data set and the first data set, the second hash value set corresponding to the second data set, and the second data set being a set of data in the first data source; and copying the difference data into the first snapshot based on the first snapshot information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of communication technologies, and in particular, to a data replication method, apparatus, electronic device, and readable storage medium. Background Art

[0002] Currently, the master-end electronic device can copy the updated data in the snapshot of a data set in the master-end electronic device to the volume of the slave-end electronic device, and then the slave-end electronic device writes the updated data in the volume to the snapshot of the data set created in the slave-end electronic device.

[0003] However, according to the above method, during the process of copying the above updated data from the master-end electronic device to the volume of the slave-end electronic device, the volume is not allowed to be opened, that is, the volume cannot perform the transmission of service data. This causes the volume in the slave-end electronic device to be unable to be used normally during snapshot remote replication. Summary of the Invention

[0004] Embodiments of the present application provide a data replication method, an electronic device, and a readable storage medium, to solve the problem that the volume in the slave-end electronic device cannot be used normally during snapshot remote replication.

[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect of the embodiments of the present application, a data replication method is provided, which is applied to a first electronic device. The method includes:

[0007] Receiving a first hash value set and first snapshot information sent by a second electronic device, where the first hash value set corresponds to a first data set, and the first data set is a set of replication data of the data in a first data source of the first electronic device in the second electronic device, and the first snapshot information is used to indicate a first snapshot of the first data set created by the second electronic device;

[0008] Determining, based on the first hash value set and a second hash value set, the differential data in a second data set that is different from the first data set, where the second hash value set corresponds to the second data set, and the second data set is a set of data in the first data source;

[0009] Copying the differential data to the first snapshot based on the first snapshot information.

[0010] In the data replication method provided by the embodiments of the present application, since the first electronic device can determine the differential data in the second data set that is different from the first data set based on the second hash value set corresponding to the second data set and the first hash value set corresponding to the first data set generated by the second electronic device, and based on the first snapshot information sent by the second electronic device, copy the differential data to the first snapshot in the second electronic device. Therefore, during snapshot remote replication, on the one hand, it is not necessary to copy all the data in the second data set, thereby reducing the total amount of data to be replicated. On the other hand, the differential data can be directly copied to the snapshot in the second electronic device without first copying it to the volume in the second electronic device, thus ensuring that the volume in the second electronic device can be used normally.

[0011] In combination with the first aspect, in a possible implementation manner, each hash value in the first hash value set is used to indicate a data subset in the first data set, and each hash value in the second hash value set is used to indicate a data subset in the second data set.

[0012] In combination with the first aspect and the above possible implementation manner, in another possible implementation manner, both the first hash value set and the second hash value set include a first number of hash values arranged in sequence;

[0013] The determining of the differential data in the second data set that is different from the first data set based on the first hash value set and the second hash value set includes:

[0014] By sequentially comparing the hash values at the same arrangement positions in the first hash value set and the second hash value set, determine the target hash value in the second hash value set, where the target hash value is different from the hash value at the corresponding same arrangement position in the first hash value set;

[0015] Determine the data in the data subset indicated by the target hash value as the differential data.

[0016] In combination with the first aspect and the above possible implementation manner, in another possible implementation manner, there is a check code for every second number of hash values in both the first hash value set and the second hash value set, and the check code is used to detect hash duplicates.

[0017] In combination with the first aspect and the above possible implementation manner, in another possible implementation manner, the determining of the differential data in the second data set that is different from the first data set based on the first hash value set and the second hash value set includes:

[0018] In the case where the check codes with the same data position in the first hash value set and the second hash value set are all the same, based on the first hash value set and the second hash value set, determine the different data in the second data set from the first data set.

[0019] Combined with the first aspect and the above possible implementation manners, in another possible implementation manner, before receiving the first hash value set and the first snapshot information sent by the second electronic device, the method further includes:

[0020] Determine a first data length according to the data volume in the second data set;

[0021] Perform hashing processing on the data in the second data set according to the first data length to obtain the second hash value set;

[0022] Send a first message to the second electronic device, where the first message is used to instruct the second electronic device to generate the first hash value set according to the first data length.

[0023] Combined with the first aspect and the above possible implementation manners, in another possible implementation manner, after copying the different data to the first snapshot based on the first snapshot information, the method further includes:

[0024] Perform hashing processing on the data in the first data source to obtain a first target hash value;

[0025] Perform hashing processing on all the copied data of the data in the first data source in the second electronic device to obtain a second target hash value;

[0026] In the case where the first target hash value is different from the second target hash value, based on a second data length, determine the different data between the data in the first data source and all the copied data. In a second aspect of the embodiments of the present application, a data replication device is provided, and the data replication device includes a receiving unit, a determining unit, and a copying unit;

[0027] The receiving unit is configured to receive a first hash value set and a first snapshot information sent by a second electronic device, where the first hash value set corresponds to a first data set, and the first data set is a set of copied data of the data in a first data source of a first electronic device in the second electronic device, and the first snapshot information is used to indicate a first snapshot of the first data set created by the second electronic device;

[0028] The determining unit is configured to determine the differential data in the second data set that is different from the first data set based on the first hash value set and the second hash value set, where the second hash value set corresponds to the second data set, and the second data set is a set of data in the first data source;

[0029] The copying unit is configured to copy the differential data to the first snapshot based on the first snapshot information.

[0030] Combined with the second aspect, in a possible implementation manner, each hash value in the first hash value set is used to indicate a data subset in the first data set, and each hash value in the second hash value set is used to indicate a data subset in the second data set.

[0031] Combined with the second aspect and the above possible implementation manner, in another possible implementation manner, both the first hash value set and the second hash value set include a first number of hash values arranged in sequence;

[0032] The determining unit is specifically configured to determine the target hash value in the second hash value set by sequentially comparing the hash values at the same arranged positions in the first hash value set and the second hash value set, where the target hash value is different from the hash value at the corresponding same arranged position in the first hash value set; and determine the data in the data subset indicated by the target hash value as the differential data.

[0033] Combined with the second aspect and the above possible implementation manner, in another possible implementation manner, there is a check code for every second number of hash values in both the first hash value set and the second hash value set, and the check code is used to detect hash duplication.

[0034] Combined with the second aspect and the above possible implementation manner, in another possible implementation manner, the electronic device further includes a processing unit and a sending unit;

[0035] The determining unit is further configured to determine a first data length according to the data volume in the second data set before the receiving unit receives the first hash value set and the first snapshot information sent by the second electronic device;

[0036] The processing unit is configured to perform hash processing on the data in the second data set according to the first data length to obtain the second hash value set;

[0037] The sending unit is configured to send a first message to the second electronic device, where the first message is used to instruct the second electronic device to generate the first hash value set according to the first data length.

[0038] For the specific implementation method, reference can be made to the behavior and functions of the first electronic device in the data replication method provided in the first aspect or possible implementation manners of the first aspect.

[0039] In a third aspect of the embodiments of the present application, an electronic device is provided, including a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the data replication method as described in the first aspect and its possible implementation manners of the first aspect are implemented.

[0040] In a fourth aspect of the embodiments of the present application, a readable storage medium is provided. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps of the data replication method as described in the first aspect and its possible implementation manners of the first aspect are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 One of the flowcharts of a data replication method provided by an embodiment of the present application;

[0043] Figure 2 Another flowchart of a data replication method provided by an embodiment of the present application;

[0044] Figure 3 Another flowchart of a data replication method provided by an embodiment of the present application;

[0045] Figure 4 A schematic diagram of the composition of a data replication device provided by an embodiment of the present application;

[0046] Figure 5 A schematic diagram of the composition of another electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0048] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more.

[0049] In addition, the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after.

[0050] The terms "at least one (item)", "at least one of", etc. in the description and claims of this application refer to any one, any two or more combinations of the objects it contains. For example, at least one (item) of a, b, and c can represent: "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple. Similarly, "at least two (items)" means two or more, and its meaning is similar to that of "at least one (item)".

[0051] First, some nouns or terms involved in the description and claims of this application will be explained below.

[0052] Distributed storage is a data storage technology that uses the disk space on each machine in a local area network through a network and constructs these scattered storage resources into a virtual storage device, and the data is scattered and stored in this virtual storage device.

[0053] A volume is a logical volume based on distributed storage allocation processing, which is equivalent to a physical disk, except that the space it actually stores is in distributed storage.

[0054] A snapshot is a fully available copy of a specified data set, which includes the image of the corresponding data at a certain point in time (e.g., the time point when the copy starts). A snapshot can be a copy or a replica of the data. The main function of a snapshot is to enable online data backup and recovery. When an application failure or file corruption occurs in the storage device, quick data recovery can be performed to restore the data to the state at a certain available time point. Another function of a snapshot is to provide another data access channel for storage users. When the original data is being processed online, users can access the snapshot data and also use the snapshot for testing and other tasks. For all storage systems applied to online systems, snapshots become an essential function.

[0055] fingerprint_size (i.e., the fingerprint block size). The logical size of a volume is often dozens of GB or hundreds of GB or more. If fingerprint set data comparison is to be performed on the volume, if the entire volume is used as the fingerprint block size, in the scenario of remote replication, only whether the two volumes are the same can be compared, and it is impossible to more precisely compare the different data. Therefore, the fingerprint block size cannot be too large, otherwise it will cause additional data transmission. The fingerprint block size cannot be too small either, otherwise the fingerprint set will be too large, and it will be saved in the configuration file in the form of default parameters. Usually, the default fingerprint block size is 4MB. In addition, the fingerprint block size can also be calculated based on the actual quantity.

[0056] Remote replication relationship pair. For remote replication, it must be a combination of the source-side resources and the backup-side resources, and it contains necessary network transmission information, configuration information, etc.

[0057] The implementation manners of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0058] Currently, for snapshot remote replication, the mainstream practices in the industry include:

[0059] 1. Product limitations

[0060] Because the snapshot remote replication processing flow is relatively complex and there are many restrictions on the associated volumes during the snapshot replication process, the mainstream practice in the industry is to directly state in the product that snapshot remote replication is not supported.

[0061] 2. Design of the storage architecture

[0062] When implementing the cloud computing architecture, the virtual machine volume and snapshot are implemented in the form of file storage. In this design, the snapshot is also a file, and the storage design form is different, which is also convenient for snapshot remote replication.

[0063] 3. First, remotely replicate to the slave volume, and then create a snapshot

[0064] This conventional practice does not allow the opening of the slave volume during snapshot replication. That is to say, the slave volume cannot be used in business scenarios, and to a certain extent, it will affect the use of the master volume, with many restrictions on the product.

[0065] 4. The volume and the snapshot are in the same resource object or container and are remotely replicated as a whole.

[0066] This requires design at the storage level. The volume and the snapshot must be in the same resource object or container, and can be remotely replicated as a whole. The information at the management level needs to be synchronized to the slave end through an additional channel. However, this solution only supports full-copy and does not support incremental copy.

[0067] As can be seen from the above, in the current snapshot remote replication solution, the master electronic device first copies the updated data to the volume of the slave electronic device. However, during the process of copying the above updated data to the volume of the slave electronic device, the volume is not allowed to be opened, that is, the volume cannot perform the transmission of business data. This causes the volume in the slave electronic device to be unable to be used normally during snapshot remote replication.

[0068] To solve the above technical problems, the embodiments of the present application provide a data replication method, apparatus, data communication device, and readable storage medium. The data replication method provided by the embodiments of the present application may have an execution subject such as a data replication apparatus, an electronic device, or a functional module in an electronic device, etc. In the embodiments of the present application, the data replication method is taken as an example in which a first electronic device executes the data replication method to illustrate the data replication method provided by the embodiments of the present application.

[0069] Figure 1 The flowchart of a data replication method provided by the embodiments of the present application is shown. The data replication method provided by the embodiments of the present application is applied to a first electronic device; as Figure 1 shown, the data replication method provided by the embodiments of the present application may include the following steps 101 to 103.

[0070] Step 101: The first electronic device receives a first hash value set and first snapshot information sent by the second electronic device.

[0071] Wherein, the above first hash value set corresponds to a first data set, and the first data set is a set of replicated data of the data in the first data source of the first electronic device in the second electronic device. The first snapshot information is used to indicate a first snapshot of the first data set created by the second electronic device.

[0072] In a possible implementation manner, the above first data source may be any data source in the first electronic device.

[0073] It can be understood that the above first data set is the data set in the second electronic device, and the data in the first data set is the data obtained by copying the data in the above first data source to the second electronic device.

[0074] In a possible implementation, the above first snapshot information may include the identity document (ID) of the above first snapshot, and the ID of the first snapshot may indicate the storage location of the first snapshot in the second electronic device.

[0075] In a possible implementation, the above first hash value set, which can also be called the first fingerprint set, includes multiple hash values, and each hash value includes a preset number of characters.

[0076] For example, taking the case where each of the above hash values is calculated by the Message-Digest Algorithm (MD) 5 hash, each of the hash values includes 32 characters.

[0077] In a possible implementation, the above first hash value set corresponds to the above first data set, which can be understood as: a part of the data in the first data set can be indicated by one of the hash values in the first hash value set.

[0078] In a possible implementation, there is a remote replication relationship pair between the first electronic device and the second electronic device.

[0079] In a possible implementation, the first electronic device may first send a first request message to the second electronic device to request the fingerprint information (i.e., the hash value set-related information) of the second electronic device. After receiving the first request message, the second electronic device may generate the above first hash value set and create the above first snapshot, and then send the first hash value set and the above first snapshot information to the first electronic device; where the first electronic device may be the master electronic device and the second electronic device may be the slave electronic device.

[0080] Step 102: The first electronic device determines the differential data between the second data set and the first data set based on the first hash value set and the second hash value set.

[0081] Wherein, the above second hash value set corresponds to the above second data set, and the second data set is the set of data in the above first data source.

[0082] It can be understood that the above second data set is the data set in the first electronic device.

[0083] In a possible implementation, the above-mentioned second hash value set, which can also be referred to as the second fingerprint set, includes multiple hash values, and each hash value includes a preset number of characters.

[0084] In a possible implementation, the number of hash values included in the above-mentioned first hash value set is the same as the number of hash values included in the above-mentioned second hash value set.

[0085] In a possible implementation, the above-mentioned difference data is the data updated in the above-mentioned first data source.

[0086] The specific method for the first electronic device to generate the above-mentioned second hash value set will be described in detail in the following embodiments. To avoid repetition, it will not be elaborated here.

[0087] In a possible implementation, there is a check code for every second number of hash values in the above-mentioned first hash value set and the above-mentioned second hash value set, and this check code is used to detect hash duplication.

[0088] In a possible implementation, the above-mentioned second number can be determined according to a check ratio, and this check ratio can be preset by the system or can be set according to actual usage requirements.

[0089] For example, taking the case where the above-mentioned check ratio is set according to actual usage requirements as an example, assuming that the fingerprint block size of a hash value set is 4MB, then the check ratio can be set to 10, so that 10 * 4MB can be determined as the above-mentioned second number. In this way, a check code is added every 10 * 4MB hash values in this hash value set to be used to detect hash duplication.

[0090] In the embodiments of the present application, since there is a check code for detecting hash duplication for every second number of hash values in the above-mentioned first hash value set and the above-mentioned second hash value set, the problem of hash duplication can be avoided, and the accuracy of the data can be improved.

[0091] In a possible implementation, the above-mentioned step 102 can be specifically implemented by the following step A.

[0092] Step A: When the check codes at the same data positions in the first hash value set and the second hash value set are the same, the electronic device determines the difference data between the second data set and the first data set based on the first hash value set and the second hash value set.

[0093] In a possible implementation, when the check codes with the same data positions in the above first hash value set and the above second hash value set are all the same, it can be considered that there is no hash duplication problem in the first hash value set and the second hash value set. In a possible implementation, if the above first hash value set and the above second hash value set are identical, but a pair of check codes with the same position therein are inconsistent, or the first hash value set and the second hash value set are inconsistent, it is necessary to record the inconsistent data addresses and adjust the above fingerprint block size for secondary hash comparison. When performing the secondary hash comparison, the above check ratio can be reduced (for example, the check ratio is 2) to reduce the data check granularity.

[0094] In the embodiments of the present application, since the electronic device will determine the different data in the second data set from the first data set based on the first hash value set and the second hash value set only when the check codes with the same data positions in the first hash value set and the second hash value set are all the same, the influence of hash duplication can be avoided, and the accuracy of the determined different data can be improved.

[0095] In a possible implementation, each hash value in the above first hash value set can be used to indicate a data subset in the above first data set, and each hash value in the above second hash value set can be used to indicate a data subset in the above second data set.

[0096] In a possible implementation, each of the above hash values can be calculated according to a corresponding data subset.

[0097] In a possible implementation, the division of the data subsets in a hash value set can be determined according to the above fingerprint block size.

[0098] In the embodiments of the present application, since each hash value in the above first hash value set can be used to indicate a data subset in the above first data set, and each hash value in the above second hash value set can be used to indicate a data subset in the above second data set, a one-to-one correspondence between the hash values and the data subsets can be established, which is convenient for data comparison and query.

[0099] In a possible implementation, both the above first hash value set and the above second hash value set include a first number of hash values arranged in sequence. Exemplarily, in combination Figure 1 , as Figure 2 shown, the above step 102 can be specifically implemented by the following step 102a and step 102b.

[0100] Step 102a: The first electronic device determines the target hash value in the second hash value set by sequentially comparing the hash values at the same arrangement positions in the first hash value set and the second hash value set.

[0101] Wherein, the above target hash value is different from the hash value at the corresponding same arrangement position in the above first hash value set.

[0102] In a possible implementation manner, the above first quantity is the quantity of data subsets included in the above first data set and the above second data set respectively.

[0103] It should be noted that if the data in the above first data set and the above second data set have not been updated, the hash values at the same arrangement positions in the above first hash value set and the above second hash value set are the same.

[0104] In a possible implementation manner, the above target hash value can be indicated by a difference serial number list; for example, if the above fingerprint block size is 4MB, the input above second hash value set is ["18c982a8a32", "816ce2b7280", "01b5e8d42bc"], and the above first hash value set is ["18c982a8a32", "81677557280", "01b5e8d42bc"]. By comparing in order, it is found that the second hash value is different, then this second hash value can be determined as the above target hash value, and then the difference data serial number list diff_list = [2] is output to indicate this second hash value.

[0105] Step 102b: The first electronic device determines the data in the data subset indicated by the target hash value as the difference data.

[0106] In a possible implementation manner, since the above target hash value is different from the hash value at the corresponding same arrangement position in the above first hash value set, it can be considered that the data in the data subset indicated by this target hash value has been updated, and thus the data in the data subset indicated by this target hash value can be determined as the difference data.

[0107] In the embodiment of the present application, since the first electronic device can determine the difference data by sequentially comparing the hash values at the same arrangement positions in the first hash value set and the second hash value set, and determining the data in the data subset indicated by the above target hash value in the above second hash value set as the difference data, that is, only by comparing the hash values can the difference data be determined, so the process of determining the difference data can be simplified.

[0108] Step 103: The first electronic device copies the difference data to the first snapshot based on the first snapshot information.

[0109] In a possible implementation, the first electronic device may determine the first snapshot in the second electronic device according to the above-mentioned first snapshot information, and then copy the difference data into the first snapshot.

[0110] In a possible implementation, after the first electronic device copies the difference data into the first snapshot, the second electronic device may compare the difference data in the first snapshot with the relevant data stored in the volume of the second electronic device, further screen out the accurate difference data, and write the screened difference data into the data pointer table of the first snapshot to complete the entire snapshot copy.

[0111] In the data replication method provided in the embodiments of the present application, since the first electronic device can determine the difference data between the second data set and the first data set based on the second hash value set corresponding to the second data set and the first hash value set generated by the second electronic device corresponding to the first data set, and copy the difference data into the first snapshot in the second electronic device based on the first snapshot information sent by the second electronic device, therefore, during snapshot remote replication, on the one hand, it is not necessary to copy all the data in the second data set to reduce the total amount of data to be copied, and on the other hand, the difference data can be directly copied into the snapshot in the second electronic device without first copying it into the volume of the second electronic device, so as to ensure that the volume in the second electronic device can be used normally.

[0112] In a possible implementation, in combination Figure 1 , as Figure 3 shown, before step 101, the data replication method provided in the embodiments of the present application may further include the following steps 104 to 106.

[0113] Step 104: The first electronic device determines the first data length according to the data volume in the second data set.

[0114] In a possible implementation, the above-mentioned first data length is the above-mentioned fingerprint block size.

[0115] In a possible implementation, the first electronic device may calculate a reasonable above-mentioned first data length according to the transmission link bandwidth with the second electronic device and the above-mentioned data volume.

[0116] In a possible implementation, the above-mentioned first data length may also be user-defined.

[0117] Step 105: The first electronic device performs hash processing on the data in the second data set according to the first data length to obtain a second hash value set.

[0118] In a possible implementation, the first data length must not exceed 50% of the transmission link bandwidth, and this value can be adjusted through a configuration file.

[0119] Exemplarily, the first electronic device may obtain the second hash value set through the following steps 1 to 6:

[0120] Step 1: Obtain the first data source, fingerprint block size fingerprint_size (i.e., the first data length) and secondary fingerprint calculation size secondary_fingerprint_size;

[0121] Step 2: Based on fingerprint_size, perform hash processing on the data in the first data source to obtain a primary fingerprint table (i.e., a primary hash value set);

[0122] Step 3: Based on check_fingerprint_size, hash the data in the first data source to obtain a peer verification fingerprint table;

[0123] Step 4: Determine whether the length of the primary fingerprint table is greater than secondary_fingerprint_size. If so, perform secondary hashing, that is, perform secondary hashing for each next_fingerprint_size primary fingerprint table unit to obtain the secondary fingerprint table (i.e., the secondary hash value set);

[0124] Step 5: Recursively determine whether the secondary hash is greater than next_fingerprint_size, and loop through step 4;

[0125] Step 6: Output the second hash value set.

[0126] In a possible implementation, the format of the second hash value set is as follows:

[0127]

[0128]

[0129] Among them, fingerprint_list represents the fingerprint table of the current level data, fingerprint_size represents the data hash unit in bytes, secondary_fingerprint_size represents the hash size of the secondary fingerprint table, and next_fingerprint_list represents the content in the next level hash table. The fingerprint table can be nested multiple times.

[0130] In a possible implementation, fingerprint_size in the above first-level list represents a hash unit, where the unit is a byte, and 4194304 means that one hash unit is 4MB. In the fingerprint_list in the first-level list, ["18c982a8a32", "816ce2b7280",

[0131] "01b5e8d42bc", "9389cd21c96b"] means:

[0132] "18c982a8a32" represents the digital fingerprint of the first 4MB data (the size of one hash unit) of the data source;

[0133] "816ce2b7280" represents the digital fingerprint of the second 4MB data of the data source;

[0134] "01b5e8d42bc" represents the digital fingerprint of the third 4MB data of the data source;

[0135] "9389cd21c96b" represents the digital fingerprint of the fourth 4MB data of the data source.

[0136] check_fingerprint_size represents the verification fingerprint size, which is calculated based on

[0137] fingerprint_size * check_rate (the default value is 10), and is used for multi-level data verification to prevent data anomalies caused by hash duplication;

[0138] check_fingerprint_list represents the verification hash values of the source data done in sequence based on check_fingerprint_size.

[0139] secondary_fingerprint_size in the first-level list means that the numbers in the secondary fingerprint table are hash values of the specified number of the upper level. Among them, in the "fingerprint_list" in the second-level list:

[0140] ["4a4d7c1870d", "6b9fc587e16"]:

[0141] "4a4d7c1870d" represents the secondary hash digital fingerprint of "18c982a8a32",

[0142] "816ce2b7280", "01b5e8d42bc" in the first-level list.

[0143] "6b9fc587e16" represents the secondary hash digital fingerprint of "9389cd21c96b" in the first-level list.

[0144] The next_fingerprint_list in the first-level list indicates that the number of hash values in this level list exceeds the secondary_fingerprint_size, so secondary hashing is required.

[0145] It should be noted that the advantage of multiple hashing is that it can compare the source and destination data from the innermost fingerprint table. If they are the same, the numbers in its lower-level fingerprint table can be omitted, which can greatly improve the hashing comparison efficiency.

[0146] Step 106, the first electronic device sends a first message to the second electronic device.

[0147] Among them, the above first message is used to instruct the second electronic device to generate the above first hash value set according to the above first data length.

[0148] In the embodiment of the present application, since the data in the second data set can be hashed according to the determined first data length to obtain the second hash value set, and the above first message is sent to the second electronic device so that the second electronic device generates the above first hash value set according to the first data length, it is possible to ensure that the number of hash values in the first hash value set and the second hash value set is the same, which is convenient for comparing the differential data.

[0149] In a possible implementation manner, before the above step 101, the data replication method provided by the embodiment of the present application may further include the following steps 107 to 109.

[0150] Step 107, the first electronic device hashes the data in the first data source to obtain a first target hash value.

[0151] Step 108, the first electronic device hashes all the replicated data of the data in the first data source in the second electronic device to obtain a second target hash value.

[0152] In a possible implementation manner, the number of the above first target hash value and the above second target hash value is one.

[0153] Step 109, in the case where the first target hash value is different from the second target hash value, the first electronic device determines the differential data between the data in the first data source and all the replicated data based on the second data length.

[0154] In a possible implementation, when the first target hash value is different from the second target hash value, it can be considered that there are still differences between the data in the first data source and all the replicated data; at this time, the fingerprint block size (i.e., the second data length) can be changed to re-determine the differential data between the two.

[0155] Exemplarily, after the first electronic device and the second electronic device complete data transmission, total_fingerprint can be verified to ensure that the data fingerprints of the snapshots on both sides are consistent. If total_fingerprint is inconsistent when the snapshot remote replication is completed, then a comparison verification is performed on the check_fingerprint_list in the fingerprint set to find the differences and check whether there are differences in the fingerprint_list. If the corresponding values of the check_fingerprint_list are consistent while the corresponding fingerprint_list is inconsistent, it indicates that there is a hash conflict. Record these addresses, reduce the fingerprint_size, and change the check_rate to 2. Calculate the hash values of these specific addresses, re-compare the specified fingerprint set after verification and then transmit the data, and finally perform the verification of total_fingerprint to ensure the accuracy of the data.

[0156] In the embodiments of the present application, after the differential data is replicated into the first snapshot, the first electronic device can, based on the comparison of the target hash values, detect again whether there are differences between the source data and the replicated data. In the case of differences, the fingerprint block size is changed to continue to determine the differential data between the two, and then replicated again. Therefore, the accuracy of data replication can be improved.

[0157] The embodiments of the present application can divide the above-mentioned electronic device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0158] In the case of dividing each functional module corresponding to each function, Figure 4 shows a possible composition schematic diagram of a data replication device; as Figure 4 shown, the data replication device 40 may include: a receiving unit 41, a determining unit 42, and a replicating unit 43.

[0159] Among them, the receiving unit 41 can be used to receive the first hash value set and the first snapshot information sent by the second electronic device. The first hash value set corresponds to the first data set, and the first data set is a set of replicated data in the second electronic device of the data in the first data source of the first electronic device. The first snapshot information is used to indicate the first snapshot of the first data set created by the second electronic device. The determining unit 42 can be used to determine the differential data in the second data set that is different from the first data set based on the first hash value set and the second hash value set. The second hash value set corresponds to the second data set, and the second data set is a set of data in the first data source. The copying unit 43 can be used to copy the differential data into the first snapshot based on the first snapshot information.

[0160] In a possible implementation manner, each hash value in the above-mentioned first hash value set can be used to indicate a data subset in the above-mentioned first data set, and each hash value in the above-mentioned second hash value set can be used to indicate a data subset in the above-mentioned second data set.

[0161] In a possible implementation manner, both the above-mentioned first hash value set and the above-mentioned second hash value set include a first number of hash values arranged in sequence. The determining unit 42 can specifically be used to determine the target hash value in the second hash value set by sequentially comparing the hash values with the same arrangement position in the first hash value set and the second hash value set. The target hash value is different from the hash value at the corresponding same arrangement position in the first hash value set; and determine the data in the data subset indicated by the target hash value as the above-mentioned differential data.

[0162] In a possible implementation manner, there is a check code for every second number of hash values in both the above-mentioned first hash value set and the above-mentioned second hash value set, and the check code is used to detect hash duplication.

[0163] In a possible implementation manner, the determining unit 42 can specifically be used to determine the differential data in the second data set that is different from the first data set based on the first hash value set and the second hash value set when the check codes with the same data position in the first hash value set and the second hash value set are the same.

[0164] In a possible implementation, the data replication device 40 may further include a first processing unit and a sending unit. The determining unit 42 may further be configured to determine a first data length according to the data volume in the second data set before the receiving unit 41 receives the first hash value set and the first snapshot information sent by the second electronic device. The first processing unit may be configured to perform a hash process on the data in the second data set according to the first data length to obtain the second hash value set. The sending unit may be configured to send a first message to the second electronic device, where the first message is used to instruct the second electronic device to generate the first hash value set according to the first data length.

[0165] In a possible implementation, the data replication device 40 may further include a second processing unit. The first processing unit may be configured to perform a hash process on the data in the first data source to obtain a first target hash value; and perform a hash process on all the replicated data of the data in the first data source in the second electronic device to obtain a second target hash value. The determining unit 42 may further be configured to, when the first target hash value is different from the second target hash value, determine the differential data between the data in the first data source and all the replicated data based on the second data length.

[0166] It should be noted that all relevant content of each step involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules, and will not be elaborated here.

[0167] It should be noted that the specific working processes of the functional modules in the data replication device provided in the embodiments of the present application can refer to the specific descriptions of the corresponding processes in the method embodiments, and will not be elaborated in detail here. The data replication device provided in the embodiments of the present application is used to execute the above data replication method, and thus can achieve the same effect as the above data replication method.

[0168] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0169] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.

[0170] The units described as separate components may or may not be physically separated. The components displayed as units can be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0171] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0172] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks or optical discs and other various media that can store program codes.

[0173] As Figure 5 shown, the embodiments of the present application further provide an electronic device 50, including a processor 51 and a memory 52. A program or instruction that can run on the processor 51 is stored on the memory 52. When the program or instruction is executed by the processor 51, it implements each step of the data copying method embodiment as described above and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0174] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above data replication method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0175] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A data replication method, applied to a first electronic device, Characterized in that, The method includes: Receiving a first hash value set and first snapshot information sent by a second electronic device, where the first hash value set corresponds to a first data set, and the first data set is a set of replicated data of the data in the first data source of the first electronic device in the second electronic device, and the first snapshot information is used to indicate a first snapshot of the first data set created by the second electronic device; Based on the first hash value set and a second hash value set, determining the differential data in the second data set that is different from the first data set, where the second hash value set corresponds to the second data set, and the second data set is a set of data in the first data source; Based on the first snapshot information, copying the differential data to the first snapshot.

2. The method according to claim 1, Characterized in that, Each hash value in the first hash value set is used to indicate a data subset in the first data set, and each hash value in the second hash value set is used to indicate a data subset in the second data set.

3. The method according to claim 2, Characterized in that, Both the first hash value set and the second hash value set include a first number of hash values arranged in sequence; The determining, based on the first hash value set and the second hash value set, the differential data in the second data set that is different from the first data set includes: By sequentially comparing the hash values at the same arrangement positions in the first hash value set and the second hash value set, determining a target hash value in the second hash value set, where the target hash value is different from the hash value at the corresponding same arrangement position in the first hash value set; Determining the data in the data subset indicated by the target hash value as the differential data.

4. The method according to claim 1, Characterized in that, There is a check code for every second number of hash values in both the first hash value set and the second hash value set, and the check code is used to detect hash duplication.

5. The method according to claim 4, Characterized in that, The determining, based on the first hash value set and the second hash value set, the differential data in the second data set that is different from the first data set includes: In the case where the check codes at the same data positions in the first hash value set and the second hash value set are the same, based on the first hash value set and the second hash value set, determining the differential data in the second data set that is different from the first data set.

6. The method according to any one of claims 1 to 5, Characterized in that, Before receiving the first hash value set and the first snapshot information sent by the second electronic device, the method further includes: Determining a first data length according to the data volume in the second data set; Performing hash processing on the data in the second data set according to the first data length to obtain the second hash value set; Send a first piece of information to the second electronic device, where the first piece of information is used to instruct the second electronic device to generate the first hash value set according to the first data length.

7. The method according to any one of claims 1 to 5, wherein, after copying the differential data to the first snapshot based on the first snapshot information, the method further includes: Performing a hash process on the data in the first data source to obtain a first target hash value; Performing a hash process on all the copied data of the data in the first data source in the second electronic device to obtain a second target hash value; In the case where the first target hash value is different from the second target hash value, determining the differential data between the data in the first data source and all the copied data based on the second data length.

8. A data replication device, characterized in that , the data replication device includes a receiving unit, a determining unit, and a copying unit; the receiving unit is configured to receive a first hash value set and first snapshot information sent by a second electronic device, the first hash value set corresponds to a first data set, the first data set is a set of copied data of the data in the first data source of the first electronic device in the second electronic device, and the first snapshot information is used to indicate a first snapshot of the first data set created by the second electronic device; the determining unit is configured to determine, based on the first hash value set and a second hash value set, the differential data between the second data set and the first data set, the second hash value set corresponds to the second data set, and the second data set is the set of data in the first data source; the copying unit is configured to copy the differential data to the first snapshot based on the first snapshot information.

9. An electronic device, wherein, it includes a processor and a memory, the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the data replication method according to any one of claims 1 to 7 are implemented.

10. A readable storage medium, wherein, a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the data replication method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Service flow copying method, device and system based on switch chip

    CN121037330A