Snapshot rollback verification method and device and distributed file system
By using a preset verification digest generation algorithm during snapshot rollback and generating a unique verification identifier using the SHA256 hash algorithm, the consistency problem in snapshot volume data replication is solved, thereby improving the reliability of data rollback and enhancing the user experience.
Patent Information
- Application Number
- CN202211259569.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-10-14
AI Technical Summary
In existing technologies, snapshot volumes are prone to abnormal interruptions during data replication, leading to data loss or tampering, which affects the reliability of data rollback.
A pre-designed checksum generation algorithm is designed to generate a unique check identifier using the SHA256 hash algorithm, which determines the consistency between the copy data in the snapshot volume and the source data, ensuring the integrity and consistency of the data replication process.
It improves the reliability of the snapshot rollback process, ensures the accuracy of data rollback and user experience, and enhances the reliability of the distributed file system.
Smart Images

Figure CN115563056B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of file system management technology, and in particular to a snapshot rollback verification method, apparatus, and distributed file system. Background Technology
[0002] A snapshot is a fully usable copy of a specified set of data. This copy includes an image of the corresponding data at a certain point in time (i.e., the time when the copy started). Its main function is to perform online data backup and recovery, roll back the data to a usable state at a certain point in time, and provide another data access channel.
[0003] In existing technologies, common snapshot techniques include the COW (Copy on First Write) method. Data blocks consist of metadata and a corresponding data ontology. In this method, the metadata in the source volume is first copied to create the snapshot volume. Then, if a data ontology in the source volume changes, according to the COW design principle, the data ontology to be changed is copied to the corresponding area allocated on the snapshot volume. Finally, the data ontology in the source volume is rewritten. This entire writing process involves two operations. As write operations on the data ontology in the source volume continue, a complete copy of the data ontology is eventually obtained. However, during the copying of the data ontology from the source volume to the snapshot volume, abnormal interruptions may occur, leading to data loss or even data tampering. This can result in incorrect information stored in the snapshot volume, making data rollback impossible and impacting user experience.
[0004] Therefore, finding an effective way to verify the consistency of data replicated to snapshot volumes is a problem that urgently needs to be solved. Summary of the Invention
[0005] The purpose of this invention is to provide a snapshot rollback verification method, apparatus, and distributed file system. A preset verification digest generation algorithm is designed, and the same data is set to produce the same result under the algorithm, which serves as the basis for consistency verification. This facilitates subsequent data rollback, ensures user experience, and further improves the reliability of the distributed file system.
[0006] To address the aforementioned technical problems, this invention provides a snapshot rollback verification method, comprising:
[0007] When a control instruction is received at the current moment to create a snapshot copy of the file to be modified in the source volume, the data body of the file to be modified is copied to the corresponding target storage area in the snapshot volume;
[0008] Based on the replica data in the target storage area and the preset verification digest generation algorithm, a first unique verification identifier corresponding to the replica data is determined;
[0009] Based on the identity identifier of the file to be rewritten and the preset identity-verification identifier correspondence, a second unique verification identifier corresponding to the data body is determined. The second unique verification identifier is predetermined based on the data body and the preset verification digest generation algorithm.
[0010] Determine whether the first unique check identifier and the second unique check identifier are the same;
[0011] If so, confirm that the replica data passes the consistency check.
[0012] Preferably, the step of establishing the preset identity-verification identifier correspondence includes:
[0013] Identify all files to be snapshotted in the source volume, wherein the files to be snapshotted comprise N layers of sub-files, where N is an integer not less than 1;
[0014] The sub-files in each layer of the file to be snapshotted are traversed according to a preset search algorithm to determine the sub-files as target snapshot files in turn, and the file to be rewritten is one of the sub-files;
[0015] Based on the preset verification digest generation algorithm and the data ontology in the target snapshot file, a second unique verification identifier of the target snapshot file is determined to obtain the preset identity-verification identifier correspondence.
[0016] Preferably, before copying the data body of the file to be rewritten to the corresponding target storage area in the snapshot volume, the method further includes:
[0017] Set the file read / write attribute of the file to be modified to read-only mode;
[0018] Correspondingly, after determining that the replica data has passed the consistency check, the process also includes:
[0019] Set the file read / write attributes of the file to be modified to read / write mode so that the data in the file can be modified.
[0020] Preferably, when determining that the first unique check identifier and the second unique check identifier are different, the following steps are included:
[0021] Upon receiving a data modification instruction for the file to be modified, the execution of the action corresponding to the data modification instruction is stopped;
[0022] The copy data in the target storage area is deleted and the control prompt module issues an alarm prompt so that the data body of the file to be rewritten can be copied again and a new consistency check can be performed.
[0023] Preferably, after determining the second unique verification identifier of the target file to be snapshotted based on the preset verification digest generation algorithm and the data ontology of the target file to be snapshotted, the method further includes:
[0024] Determine the name of the second unique verification identifier;
[0025] The file extended attribute function component is invoked to store the name and value of the second unique verification identifier as metadata information of the target snapshot file in the form of key-value pairs.
[0026] Preferably, the preset verification digest generation algorithm is a hash encryption algorithm.
[0027] Preferably, the total number of files to be rewritten is multiple, and each file to be rewritten belongs to the same upper-level management file;
[0028] The snapshot rollback verification method further includes:
[0029] When it is determined that the consistency check of the copy data corresponding to all the files to be rewritten has passed, the first total check identifier of the copy data corresponding to the upper-level management file is determined based on the first unique check identifier of each copy data and the preset total check strategy.
[0030] Based on the identity identifier of the upper-level management file and the preset identity-verification identifier correspondence, a second total verification identifier corresponding to the total data of the data body of the upper-level management file is determined. The second total verification identifier is predetermined based on the second unique verification identifier of each data body and the preset total verification strategy.
[0031] Determine whether the first total checksum identifier and the second total checksum identifier are the same;
[0032] If so, confirm that the upper-level management file has passed the consistency check.
[0033] Preferably, based on the first unique verification identifier of each copy data and a preset overall verification strategy, the first overall verification identifier corresponding to the total copy data of the upper-level management file is determined, including:
[0034] The first unique verification identifier of each copy of the data is incremented;
[0035] Based on the first accumulation result after accumulation and the preset check digest generation algorithm, the first check total identifier corresponding to the total copy data of the upper-level management file is determined;
[0036] The second step of pre-determining the total verification identifier includes:
[0037] The second unique verification identifier of each of the data bodies is accumulated;
[0038] Based on the second accumulation result after accumulation and the preset verification digest generation algorithm, the second verification total identifier corresponding to the total data of the data ontology of the upper-level management file is determined.
[0039] To address the aforementioned technical problems, the present invention also provides a snapshot rollback verification device, comprising:
[0040] Memory, used to store computer programs;
[0041] A processor is configured to implement the steps of the snapshot rollback verification method as described above when executing the computer program.
[0042] To address the aforementioned technical problems, the present invention also provides a distributed file system, comprising:
[0043] The copying unit is used to copy the data body of the file to be modified to the corresponding target storage area in the snapshot volume when it receives a control instruction representing the creation of a snapshot copy of the file to be modified in the source volume at the current moment.
[0044] The first determining unit is configured to determine the first unique verification identifier corresponding to the replica data based on the replica data in the target storage area and a preset verification digest generation algorithm.
[0045] The second determining unit is used to determine the second unique verification identifier corresponding to the data body based on the identity identifier of the file to be rewritten and the preset identity-verification identifier correspondence. The second unique verification identifier is predetermined based on the data body and the preset verification digest generation algorithm.
[0046] The first judgment unit is used to determine whether the first unique verification identifier and the second unique verification identifier are the same; if so, the verification pass unit is triggered.
[0047] The verification unit is used to determine that the replica data has passed the consistency check.
[0048] This application provides a snapshot rollback verification method, apparatus, and distributed file system. When a control instruction representing the creation of a snapshot copy of a file to be modified in the source volume is received at the current moment, the data body of the file to be modified is copied to the corresponding target storage area in the snapshot volume. Considering that failures may occur during the copying process, leading to errors in the copying result, a preset verification digest generation algorithm is designed, and the same data is set to produce the same result under the algorithm, which serves as the basis for consistency verification. Specifically, based on the copy data in the target storage area and the preset verification digest generation algorithm, a first unique verification identifier corresponding to the copy data is determined. Based on the identity identifier of the file to be modified and the preset identity-verification identifier correspondence, a second unique verification identifier corresponding to the data body is determined. The second unique verification identifier is pre-determined based on the data body and the preset verification digest generation algorithm. Subsequently, it is determined whether the first unique verification identifier and the second unique verification identifier are the same. If they are the same, it indicates that the data body and the copy data are consistent, and the copy data passes the consistency verification, indicating successful data copying. This facilitates subsequent data rollback, ensures user experience, and further improves the reliability of the distributed file system. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the prior art and embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 A flowchart of a snapshot rollback verification method provided by the present invention;
[0051] Figure 2 This is a schematic diagram of the structure of a snapshot rollback verification device provided by the present invention;
[0052] Figure 3 This is a schematic diagram of the structure of a distributed file system provided by the present invention. Detailed Implementation
[0053] The core of this invention is to provide a snapshot rollback verification method, device, and distributed file system. It designs a preset verification digest generation algorithm and sets that the same data will produce the same result under the algorithm, which serves as the basis for consistency verification. This facilitates subsequent data rollback, ensures user experience, and further improves the reliability of the distributed file system.
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] Please refer to Figure 1 , Figure 1 The flowchart illustrates a snapshot rollback verification method provided by this invention.
[0056] In this embodiment, considering that in the existing COW method, after creating a snapshot volume, if there is a data ontology modification operation, the data ontology to be modified needs to be copied to the corresponding area allocated on the snapshot volume before the data ontology in the source volume is rewritten. However, during the process of copying the data ontology from the source volume to the snapshot volume, abnormal interruptions may occur, leading to data loss, or even data tampering, resulting in incorrect information stored in the snapshot volume, making data rollback impossible and affecting user experience. To solve the above technical problems, this application provides a snapshot rollback verification method, which realizes consistency verification between the data ontology and the copied replica data, facilitating subsequent data rollback.
[0057] The snapshot rollback verification method includes:
[0058] S11: When a control instruction representing the creation of a snapshot copy of the file to be modified in the source volume is received at the current moment, the data body of the file to be modified is copied to the corresponding target storage area in the snapshot volume;
[0059] Specifically, this method includes, but is not limited to, applications in distributed file systems. The source volume of this file system includes multiple files. It can be understood that all files in the multiple files that require snapshotting are pre-determined, and corresponding snapshot volumes are pre-established based on the metadata information of the files to be snapshotted. The files to be snapshotted may specifically include multiple layers of sub-files, and each sub-file has a corresponding storage area in the snapshot volume. The file to be rewritten can be one of the multiple sub-files mentioned above, and the storage area corresponding to the file to be rewritten is the target storage area.
[0060] Furthermore, the development environment for this method includes, but is not limited to, C++ or the Linux operating system; no special restrictions are imposed here.
[0061] S12: Determine the first unique verification identifier corresponding to the replica data based on the replica data in the target storage area and the preset verification digest generation algorithm;
[0062] Specifically, a pre-designed checksum generation algorithm is provided. This algorithm includes, but is not limited to, the SHA256 hash algorithm (an algorithm that generates a checksum based on the data to be encrypted, with a hash value length of 256, or 32 bytes). Furthermore, when this algorithm is used to process the same data, the resulting checksum identifiers are always identical. Therefore, this pre-designed checksum generation algorithm is applied to generate the first unique checksum identifier for the copy data. When the algorithm is the SHA256 hash algorithm, the first unique checksum identifier is the first hash value.
[0063] S13: Based on the identity identifier of the file to be rewritten and the preset identity-verification identifier correspondence, determine the second unique verification identifier corresponding to the data body. The second unique verification identifier is predetermined based on the data body and the preset verification digest generation algorithm.
[0064] Specifically, a pre-stored identity-verification identifier mapping relationship is provided. This mapping relationship includes the identifiers of all files to be snapshotted and a corresponding second unique verification identifier generated according to a pre-defined verification digest generation algorithm. The specific establishment steps are described in the following embodiments and will not be repeated here. It is understood that when the algorithm is the SHA256 hash algorithm described above, the second unique verification identifier is the second hash value.
[0065] S14: Determine whether the first unique check identifier and the second unique check identifier are the same; if so, proceed to S15;
[0066] S15: Determine that the replica data has passed the consistency check.
[0067] It is understandable that when the first unique verification identifier is the same as the second unique verification identifier, it means that the replica data is the same as the original data body and no abnormal situation occurred during the data replication process. Therefore, it is determined that the replica data passes the consistency verification. Then, when a specific data change instruction including the rewrite information of the file to be rewritten is received in the future, it can respond to it to execute the corresponding rewrite action for the file to be rewritten.
[0068] Furthermore, it should be noted that the preset verification digest generation algorithm in this application does not directly encrypt the data ontology itself (algorithms that encrypt the data ontology itself need to be decrypted in the snapshot volume later). Thus, the method provided by this application not only verifies whether the data ontology has been tampered with during the copying process, but also verifies whether there is a transmission interruption during the copying process that causes data transmission errors. Moreover, it eliminates the need for decryption steps and avoids increasing the computational burden on the distributed file system, thereby better ensuring the efficiency of subsequent snapshot rollback.
[0069] In summary, this application provides a snapshot rollback verification method. Considering that failures may occur during the replication process, leading to errors in the replication results, a preset verification digest generation algorithm is designed, and the same data is set to produce the same result under the algorithm, which serves as the basis for consistency verification. When the first unique verification identifier and the second unique verification identifier are the same, it indicates that the data body and the replica data are consistent, and the replica data is confirmed to have passed the consistency verification. The data replication is successful, which facilitates subsequent data rollback, ensures the user experience, and further improves the reliability of the distributed file system.
[0070] Based on the above embodiments:
[0071] As a preferred embodiment, the step of establishing the pre-defined identity-verification identifier correspondence includes:
[0072] Identify all files to be snapshotted in the source volume. The files to be snapshotted consist of N sub-files, where N is an integer not less than 1.
[0073] According to the preset search algorithm, traverse the sub-files in each layer of each snapshot file to determine the sub-file as the target snapshot file in turn, and the file to be rewritten is one of the sub-files;
[0074] Based on the preset verification digest generation algorithm and the data ontology in the target snapshot file, a second unique verification identifier of the target snapshot file is determined to obtain the preset identity-verification identifier correspondence.
[0075] In this embodiment, the steps for establishing the correspondence between the preset identity and the verification identifier are given. Specifically, the source volume of the distributed file system includes multiple files. When a snapshot instruction is received at the current moment, all the files in the source volume that need to be snapshotted can be determined according to the instruction. Considering that the files to be snapshotted include files and directories, and there may be subdirectories under the directory, the files to be snapshotted can include N layers of sub-files.
[0076] The sub-files in each layer of each snapshot file are traversed according to a preset search algorithm, which includes, but is not limited to, a breadth-first search algorithm. Then, according to the recursive method of the breadth-first search algorithm, each of the sub-files is taken as a target snapshot file. Based on a preset verification digest generation algorithm and the data ontology in the target snapshot file, a second unique verification identifier of the target snapshot file is determined to obtain a preset identity-verification identifier correspondence. The actual storage form of this correspondence includes, but is not limited to, storing it in the form of a table.
[0077] As can be seen, the above method can simply and reliably determine the correspondence between the preset identity and the verification identifier, laying the foundation for subsequent data integrity and consistency verification.
[0078] As a preferred embodiment, before copying the data body of the file to be rewritten to the corresponding target storage area in the snapshot volume, the method further includes:
[0079] Set the file read / write attribute of the file to be modified to read-only mode;
[0080] Correspondingly, after confirming that the replica data has passed the consistency check, the following steps are also included:
[0081] Set the file read / write attributes of the file to be modified to read / write mode so that the data in the file can be modified.
[0082] In this embodiment, to ensure that the consistency verification of the replica data used for snapshot rollback is passed, user rewriting operations are prohibited to prevent subsequent data from being unrecoverable. Therefore, before copying the data body of the file to be rewritten to the corresponding target storage area in the snapshot volume, the file read / write attribute of the file to be rewritten can be set to read-only mode. After confirming that the replica data has passed the consistency verification, the user's rewriting operation is allowed, that is, the file read / write attribute of the file to be rewritten is set to read / write mode so that the data body in the file to be rewritten can be modified, thus ensuring the reliable implementation of snapshot rollback.
[0083] As a preferred embodiment, when determining that the first unique verification identifier and the second unique verification identifier are different, the following steps are taken:
[0084] Upon receiving a data modification instruction for a file to be modified, stop executing the corresponding action.
[0085] Delete the copy data in the target storage area and control the prompting module to issue an alarm so that the data body in the file to be rewritten can be copied again and a new consistency check can be performed.
[0086] In this embodiment, when the first unique verification identifier and the second unique verification identifier are different, the copy data in the target storage area can be deleted (or the target storage area can be marked for later log viewing and cause analysis). The alert module is then directly controlled to issue an alarm, allowing the user to promptly understand the situation and pause subsequent rewrite operations to prevent data loss and irreversibility. The alert module includes, but is not limited to, a display module, such as a human-computer interaction display interface. After understanding the situation, the user can resend the corresponding control command to recopy the data in the file to be rewritten and perform a new consistency check. Furthermore, considering that after sending the control command to create a snapshot copy, the user will send a data change command indicating that the file to be rewritten is to be modified, if the consistency check fails, the rewrite action corresponding to the data change command can be stopped to avoid data loss and irreversibility.
[0087] It is understood that there is no particular limitation on the order of execution of the two steps in this embodiment. Furthermore, the final snapshot backup of all files to be snapshotted in the source volume may fail the consistency check; in such cases, a certain number of repeated copying steps are required to obtain a final, fully usable backup.
[0088] As a preferred embodiment, after determining the second unique verification identifier of the target snapshot file based on the preset verification digest generation algorithm and the data ontology in the target snapshot file, the method further includes:
[0089] Determine the name of the second unique check identifier;
[0090] Call the file extended attributes function component to store the name and value of the second unique check identifier as key-value pairs as metadata information of the target snapshot file.
[0091] In this embodiment, the second unique checksum determined at the current moment can also be stored as metadata information. Therefore, after determining the second unique checksum of the target file to be snapshotted, the name of the second unique checksum can also be determined. Then, the file extended attribute function component is invoked to store the name and value of the second unique checksum as key-value pairs as metadata information of the target file to be snapshotted. As an example, this method is applied to the Linux operating system. The Linux file extended attribute function (i.e., EA function) can be used to associate metadata information with the target file node to be snapshotted. The user's file extended attribute namespace is accessed, and the lsetattr or setattr command is used to implement a system call. Finally, the name and value of the second unique checksum are stored as key-value pairs as metadata information of the target file to be snapshotted, facilitating subsequent viewing.
[0092] In a preferred embodiment, the preset verification digest generation algorithm is a hash encryption algorithm.
[0093] In this embodiment, the preset checksum generation algorithm is provided, which can be a hash encryption algorithm, more specifically, the SHA256 hash algorithm described above. It can generate a unique identification code for the processed data in a relatively short time and has high security, thus better fulfilling the design requirements of the preset checksum generation algorithm in this application.
[0094] In one preferred embodiment, the total number of files to be rewritten is multiple, and each file to be rewritten belongs to the same upper-level management file;
[0095] The snapshot rollback verification method also includes:
[0096] Once it is determined that the consistency check of the copy data corresponding to all files to be rewritten has passed, the first unique check identifier of the copy data corresponding to the upper-level management file is determined based on the first unique check identifier of each copy data and the preset total check strategy.
[0097] Based on the identity identifier of the upper-level management file and the preset identity-verification identifier correspondence, a second total verification identifier corresponding to the total data of the data body of the upper-level management file is determined. The second total verification identifier is predetermined based on the second unique verification identifier of each data body and the preset total verification strategy.
[0098] Determine whether the first checksum identifier and the second checksum identifier are the same;
[0099] If so, confirm that the upper-level management files have passed the consistency check.
[0100] In this embodiment, the inventors further considered that multiple files to be rewritten in practical applications may belong to the same upper-level management file. Therefore, redundant checks can be further set to maximize the effectiveness and reliability of the snapshot rollback verification method provided in this application. Specific steps are as described above and will not be repeated here. It should be noted that the same data will have the same total verification identifier under the preset total verification strategy. Therefore, when the first total verification identifier and the second total verification identifier are the same, it can be determined that the upper-level management file has passed the consistency check, and then the user can be allowed to rewrite the files to be rewritten in that upper-level management file.
[0101] As a preferred embodiment, based on the first unique verification identifier of each copy of the data and a preset overall verification strategy, a first overall verification identifier corresponding to the total copy data of the upper-level management file is determined, including:
[0102] The first unique verification identifier of each copy of the data is incremented;
[0103] Based on the first accumulation result after accumulation and the preset check digest generation algorithm, determine the first check total identifier corresponding to the total copy data of the upper-level management file;
[0104] The second step of pre-determining the overall identifier includes:
[0105] The second unique verification identifier of each data body is accumulated;
[0106] Based on the second accumulation result after accumulation and the preset checksum generation algorithm, the second checksum identifier corresponding to the total data of the data ontology of the upper-level management file is determined.
[0107] In this embodiment, the steps for determining the first total verification identifier and the second total verification identifier are given. Specifically, taking the SHA256 hash algorithm as the preset verification digest generation algorithm as an example, the determination of the first total verification identifier is carried out by accumulating the first unique verification identifier (i.e., the first unique verification identifier of 32 bytes) of each copy data. Based on the first accumulation result and the preset verification digest generation algorithm, the first total verification identifier (i.e., the first total verification identifier of 32 bytes) of the copy total data corresponding to the upper-level management file can be determined. The determination of the second total verification identifier is similar and will not be repeated here.
[0108] As can be seen, compared to the method of calculating the corresponding verification identifier by relying on a preset verification digest generation algorithm for the entire upper-level management file, the above method is simpler and takes less time when the data volume of the upper-level management file is large, making it easier to apply in practice and ensuring the efficiency of consistency verification.
[0109] Please refer to Figure 2 , Figure 2 This is a schematic diagram of the structure of a snapshot rollback verification device provided by the present invention.
[0110] The snapshot rollback verification apparatus includes:
[0111] Memory 21 is used to store computer programs;
[0112] The processor 22 is configured to implement the steps of the snapshot rollback verification method as described above when executing the computer program.
[0113] For a description of the snapshot rollback verification device provided in this invention, please refer to the above-described embodiments of the snapshot rollback verification method; further details will not be repeated here.
[0114] Please refer to Figure 3 , Figure 3 This is a schematic diagram of the structure of a distributed file system provided by the present invention.
[0115] This distributed file system includes:
[0116] The copying unit 31 is used to copy the data body of the file to be modified to the corresponding target storage area in the snapshot volume when it receives a control instruction representing the creation of a snapshot copy of the file to be modified in the source volume at the current moment.
[0117] The first determining unit 32 is used to determine the first unique verification identifier corresponding to the copy data based on the copy data in the target storage area and the preset verification digest generation algorithm.
[0118] The second determining unit 33 is used to determine the second unique verification identifier corresponding to the data body based on the identity identifier of the file to be rewritten and the preset identity-verification identifier correspondence relationship. The second unique verification identifier is predetermined based on the data body and the preset verification digest generation algorithm.
[0119] The first judgment unit 34 is used to determine whether the first unique verification identifier and the second unique verification identifier are the same; if so, the verification pass unit 35 is triggered.
[0120] Verification unit 35 is used to determine that the replica data has passed the consistency check.
[0121] For an introduction to the distributed file system provided in this invention, please refer to the above-described embodiment of the snapshot rollback verification method; it will not be repeated here.
[0122] In a preferred embodiment, the distributed file system further includes a pre-establishment unit, which is used to establish the preset identity-verification identifier correspondence, including:
[0123] The third determining unit is used to determine all the files to be snapshotted in the source volume, wherein the files to be snapshotted include N layers of sub-files, where N is an integer not less than 1;
[0124] The traversal unit is used to traverse the sub-files in each layer of the file to be snapshotted according to a preset search algorithm, so as to determine the sub-files as target snapshot files in turn, and the file to be rewritten is one of the sub-files;
[0125] The fourth determining unit is used to determine the second unique verification identifier of the target snapshot file based on the preset verification digest generation algorithm and the data ontology of the target snapshot file, so as to obtain the preset identity-verification identifier correspondence.
[0126] As a preferred embodiment, the distributed file system further includes:
[0127] The first setting unit is used to set the file read / write attribute of the file to be modified to read-only mode;
[0128] Correspondingly, the distributed file system also includes:
[0129] The second setting unit is used to set the file read / write attribute of the file to be modified to read / write mode after the verification unit 35 passes, so as to modify the data body of the file to be modified.
[0130] In a preferred embodiment, the first determination unit 34 is used to determine whether the first unique verification identifier and the second unique verification identifier are the same; if not, the stop response unit and the feedback unit are triggered.
[0131] The stop response unit is used to stop executing the action corresponding to the data change instruction when it receives a data change instruction for the file to be modified.
[0132] The feedback unit is used to delete the copy data in the target storage area and control the prompting module to issue an alarm prompt, so as to re-copy the data body of the file to be rewritten and perform a new consistency check.
[0133] As a preferred embodiment, the distributed file system further includes:
[0134] The fifth determining unit is used, after the fourth determining unit, to determine the name of the second unique verification identifier;
[0135] The component invocation unit is used to invoke the file extended attribute function component to store the name and value of the second unique verification identifier as metadata information of the target snapshot file in the form of key-value pairs.
[0136] In a preferred embodiment, the total number of files to be rewritten is multiple, and each file to be rewritten belongs to the same upper-level management file;
[0137] The distributed file system also includes:
[0138] The sixth determining unit is used to determine the first total verification identifier of the total copy data corresponding to the upper-level management file based on the first unique verification identifier of each copy data and the preset total verification strategy when it is determined that the consistency verification of the copy data corresponding to all the files to be rewritten has passed.
[0139] The seventh determining unit is used to determine the first total verification identifier of the total copy data corresponding to the upper-level management file based on the first unique verification identifier of each copy data and the preset total verification strategy when it is determined that the consistency verification of the copy data corresponding to all the files to be rewritten has passed.
[0140] The second judgment unit is used to determine whether the first total verification identifier and the second total verification identifier are the same; if so, the eighth determination unit is triggered.
[0141] The eighth determining unit is used to determine that the upper-level management file has passed the consistency check.
[0142] In a preferred embodiment, the sixth determining unit specifically includes:
[0143] The first accumulation unit is used to accumulate the first unique verification identifier of each of the replica data;
[0144] The ninth determining unit is used to determine the first total verification identifier corresponding to the total copy data of the upper-level management file based on the first accumulated result after accumulation and the preset verification digest generation algorithm;
[0145] The distributed file system further includes a second total verification identifier pre-determination unit, used for pre-determining the second total verification identifier;
[0146] The second verification total identifier pre-determination unit includes:
[0147] The second accumulation unit is used to accumulate the second unique verification identifier of each of the data bodies;
[0148] The tenth determining unit is used to determine the second total verification identifier corresponding to the total data of the data ontology of the upper-level management file based on the second accumulated result after accumulation and the preset verification digest generation algorithm.
[0149] The various embodiments described in this specification are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0150] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can implement the described functions using different methods for each specific application, but such implementation should not be considered beyond the scope of the invention. The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A snapshot rollback verification method, characterized in that, include: When a control instruction is received at the current moment to create a snapshot copy of the file to be modified in the source volume, the data body of the file to be modified is copied to the corresponding target storage area in the snapshot volume; Based on the replica data in the target storage area and the preset verification digest generation algorithm, determine the first unique verification identifier corresponding to the replica data; Based on the identity identifier of the file to be rewritten and the preset identity-verification identifier correspondence, a second unique verification identifier corresponding to the data body is determined. The second unique verification identifier is predetermined based on the data body and the preset verification digest generation algorithm. Determine whether the first unique check identifier and the second unique check identifier are the same; If so, confirm that the replica data passes the consistency check; The total number of files to be rewritten is multiple, and each file to be rewritten belongs to the same upper-level management file; The snapshot rollback verification method further includes: When it is determined that the consistency check of the copy data corresponding to all the files to be rewritten has passed, the first total check identifier of the copy data corresponding to the upper-level management file is determined based on the first unique check identifier of each copy data and the preset total check strategy. Based on the identity identifier of the upper-level management file and the preset identity-verification identifier correspondence, a second total verification identifier corresponding to the total data of the data body of the upper-level management file is determined. The second total verification identifier is predetermined based on the second unique verification identifier of each data body and the preset total verification strategy. Determine whether the first total checksum identifier and the second total checksum identifier are the same; If so, confirm that the upper-level management file has passed the consistency check.
2. The snapshot rollback verification method as described in claim 1, characterized in that, The steps for establishing the preset identity-verification identifier correspondence include: Identify all files to be snapshotted in the source volume, wherein the files to be snapshotted comprise N layers of sub-files, where N is an integer not less than 1; The sub-files in each layer of the file to be snapshotted are traversed according to a preset search algorithm to determine the sub-files as target files to be snapshotted in sequence, and the file to be rewritten is one of the sub-files; Based on the preset verification digest generation algorithm and the data ontology in the target snapshot file, a second unique verification identifier of the target snapshot file is determined to obtain the preset identity-verification identifier correspondence.
3. The snapshot rollback verification method as described in claim 1, characterized in that, Before copying the data body of the file to be rewritten to the corresponding target storage area in the snapshot volume, the process also includes: Set the file read / write attribute of the file to be modified to read-only mode; Correspondingly, after determining that the replica data has passed the consistency check, the process also includes: Set the file read / write attributes of the file to be modified to read / write mode so that the data in the file can be modified.
4. The snapshot rollback verification method as described in claim 3, characterized in that, When determining that the first unique check identifier and the second unique check identifier are different, the following is included: Upon receiving a data modification instruction for the file to be modified, the execution of the action corresponding to the data modification instruction is stopped; The copy data in the target storage area is deleted and the control prompt module issues an alarm prompt so that the data body of the file to be rewritten can be copied again and a new consistency check can be performed.
5. The snapshot rollback verification method as described in claim 2, characterized in that, After determining the second unique verification identifier of the target file to be snapshotted based on the preset verification digest generation algorithm and the data ontology of the target file to be snapshotted, the method further includes: Determine the name of the second unique verification identifier; The file extended attribute function component is invoked to store the name and value of the second unique verification identifier as metadata information of the target snapshot file in the form of key-value pairs.
6. The snapshot rollback verification method as described in claim 1, characterized in that, The preset verification digest generation algorithm is a hash encryption algorithm.
7. The snapshot rollback verification method as described in claim 1, characterized in that, Based on the first unique verification identifier of each copy of the data and the preset overall verification strategy, the first overall verification identifier of the copy data corresponding to the upper-level management file is determined, including: The first unique verification identifier of each copy of the data is incremented; Based on the first accumulation result after accumulation and the preset check digest generation algorithm, the first check total identifier corresponding to the total copy data of the upper-level management file is determined; The second step of pre-determining the total verification identifier includes: The second unique verification identifier of each of the data bodies is accumulated; Based on the second accumulation result after accumulation and the preset verification digest generation algorithm, the second verification total identifier corresponding to the total data of the data ontology of the upper-level management file is determined.
8. A snapshot rollback verification device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the snapshot rollback verification method as described in any one of claims 1 to 7 when executing the computer program.
9. A distributed file system, characterized in that, include: The copying unit is used to copy the data body of the file to be modified to the corresponding target storage area in the snapshot volume when it receives a control instruction representing the creation of a snapshot copy of the file to be modified in the source volume at the current moment. The first determining unit is configured to determine the first unique verification identifier corresponding to the replica data based on the replica data in the target storage area and a preset verification digest generation algorithm. The second determining unit is used to determine the second unique verification identifier corresponding to the data body based on the identity identifier of the file to be rewritten and the preset identity-verification identifier correspondence. The second unique verification identifier is predetermined based on the data body and the preset verification digest generation algorithm. The first judgment unit is used to determine whether the first unique verification identifier and the second unique verification identifier are the same; if so, the verification pass unit is triggered. The verification unit is used to determine that the replica data has passed the consistency check; The total number of files to be rewritten is multiple, and each file to be rewritten belongs to the same upper-level management file; The distributed file system is further configured to: when it is determined that the consistency verification of the replica data corresponding to all the files to be rewritten has passed, determine the first total verification identifier of the replica data corresponding to the upper-level managed file based on the first unique verification identifier of each replica data and the preset total verification strategy; Based on the identity identifier of the upper-level management file and the preset identity-verification identifier correspondence, a second verification identifier corresponding to the total data of the data body of the upper-level management file is determined. The second verification identifier is predetermined based on the second unique verification identifier of each data body and the preset total verification strategy. It is then determined whether the first verification identifier and the second verification identifier are the same. If so, it is determined that the upper-level management file has passed the consistency verification.
Citation Information
Patent Citations
Method and device for quickly verifying consistency of multi-duplicate storage
CN107203345A
File integrity detection method and device
CN112883427A