A method, apparatus, device, and product for file storage and backup

By adopting a variable-length chunking method based on RK algorithm in the cStor distributed cloud storage system, the problem of low data re-deletion rate in the insertion and write-deletion scenarios is solved, and a more efficient backup process is achieved and the user experience is improved.

CN115454714BActive Publication Date: 2025-06-10CHINA TELECOM CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210908292.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-06-10
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

The cStor distributed cloud storage system has a low data deletion rate in the insertion and write scenarios, resulting in too long backup time and poor user experience.

Method used

The variable-length chunking method based on the RK algorithm is adopted to match the data block content of the backup data to be stored and the data content in the preset repository, and divide the data blocks to be processed to improve the data deletion rate.

Benefits of technology

It improves the data de-deletion rate, reduces the space occupancy rate of backup data, reduces the backup time, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115454714B_ABST
    Figure CN115454714B_ABST
Patent Text Reader

Abstract

The present invention discloses a file storage and backup method, apparatus, device and product, relating to the field of data processing. The method includes: determining that the data to be stored and backed up is not transmitted for the first time, matching the data content corresponding to the data block in the data to be stored and backed up with a preset storage repository based on the RK algorithm, determining the block boundary, and dividing the data to be stored and backed up based on the block boundary to obtain data blocks to be processed; determining the hash value of the data blocks to be processed, and determining the data fingerprint of the data blocks to be processed based on the hash value; matching the data fingerprints corresponding to all the data blocks to be processed with a preset fingerprint database, and determining whether the data blocks to be processed need to be backed up according to the matching situation; storing and backing up the data blocks to be processed that have been determined to need to be backed up. The block division result of the present invention is more reasonable, effectively improving the data deduplication rate, reducing the space occupancy rate of the backup data, reducing the backup time, and bringing a better experience to users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a file storage and backup method, device, equipment and product. Background Art

[0002] The cStor distributed cloud storage system uses file backup to ensure data reliability and security. After multiple backups, a large amount of duplicate data is generated. This backup method uses incremental backup to solve problems such as insufficient data storage space and low storage efficiency. However, there are still the following problems:

[0003] The backup method of the cStor distributed cloud storage system divides the data to be backed up into blocks with a fixed data length. It has a good deduplication effect for the scenario of overwrite write, but has a poor deduplication effect in the scenarios of insert write and delete write, resulting in too long backup time and poor user experience. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a file storage and backup method, device, equipment and product to solve the problems of low data deduplication rate and slow backup speed caused by the fixed-length data block division method.

[0005] According to a first aspect, an embodiment of the present invention provides a file storage and backup method, the method comprising:

[0006] Determine that the data to be stored and backed up is not transmitted for the first time, match the data content corresponding to the data blocks in the data to be stored and backed up with a preset storage repository based on the RK algorithm, determine the block boundary, and divide the data to be stored and backed up based on the block boundary to obtain data blocks to be processed; the data content includes data length and hash value, and the preset storage repository stores the data content corresponding to the backed-up data blocks; the data blocks to be processed are composed of several data blocks;

[0007] Determine the hash value of the data blocks to be processed, and determine the data fingerprint of the data blocks to be processed based on the hash value;

[0008] Match the data fingerprints corresponding to all the data blocks to be processed with a preset fingerprint database, and determine whether the data blocks to be processed need to be backed up according to the matching situation; the preset fingerprint database stores the data fingerprints corresponding to the backed-up data blocks;

[0009] Store and back up the data blocks to be processed that have been determined to need to be backed up.

[0010] In combination with the first aspect, in the first implementation manner of the first aspect, before the step of determining that the backup data to be stored is not the first time to be transmitted, and based on the RK algorithm, matching the data content corresponding to the data blocks included in the backup data to be stored with a preset repository, determining the block boundaries, and dividing the backup data to be stored based on the block boundaries to obtain the data blocks to be processed, the method further includes the following steps:

[0011] Determine that the backup data to be stored is the first time to be transmitted, determine the block boundaries based on a preset data length, and divide the backup data to be stored based on the block boundaries to obtain reference data blocks; the reference data blocks are composed of several data blocks;

[0012] Determine the hash value of the reference data blocks, and determine the data fingerprints of the reference data blocks based on the hash value;

[0013] Store and back up all the reference data blocks, store the data content corresponding to the backed-up data blocks in the preset database, and store the data fingerprints corresponding to the backed-up data blocks in the preset fingerprint database.

[0014] In combination with the first aspect, in the second implementation manner of the first aspect, the step of determining that the backup data to be stored is not the first time to be transmitted, and based on the RK algorithm, matching the data content corresponding to the data blocks included in the backup data to be stored with a preset repository, determining the block boundaries, and dividing the backup data to be stored based on the block boundaries to obtain the data blocks to be processed specifically includes:

[0015] Determine that the backup data to be stored is not the first time to be transmitted, and based on the RK algorithm, compare whether the data content corresponding to the data blocks included in the backup data to be stored is equal to the data content corresponding to the backed-up data blocks in the preset database;

[0016] Determine that there are data blocks with equal data content, and determine the first block boundary and the second block boundary; the first block boundary is the boundary between the foremost data block of the determined data blocks and the previous data block adjacent to the foremost data block, and the second block boundary is the boundary between the rearmost data block of the determined data blocks and the next data block adjacent to the rearmost data block, and the first block boundary and the second block boundary appear in pairs;

[0017] Based on the first block boundary and the second block boundary, divide the backup data to be stored to obtain the data blocks to be processed.

[0018] In combination with the second implementation manner of the first aspect, in the third implementation manner of the first aspect, the step of dividing the backup data to be stored based on the first block boundary and the second block boundary to obtain the data blocks to be processed specifically includes:

[0019] Before determining the first block boundary, if there is no block boundary, the data block before the first block boundary is divided into a to-be-processed data block;

[0020] After determining the second block boundary, if there is no block boundary, the data block after the second block boundary is divided into a to-be-processed data block;

[0021] Before determining the second block boundary, if there is a block boundary, the data between the second block boundary and the nearest previous first block boundary to the second block boundary is divided into the to-be-processed data block.

[0022] Combined with the first aspect, in the fourth embodiment of the first aspect, the data fingerprints corresponding to all the to-be-processed data blocks are matched with a preset fingerprint database, and it is determined whether the to-be-processed data blocks need to be backed up according to the matching situation; the data fingerprints corresponding to the already backed-up data blocks are stored in the preset fingerprint database, and specifically include:

[0023] The to-be-stored backup data is transmitted into the fingerprint database;

[0024] Compare whether the data fingerprints corresponding to all the to-be-processed data blocks are equal to the data fingerprints corresponding to the already backed-up data blocks in the preset fingerprint database, and determine whether each to-be-processed data block into which the to-be-stored backup data is divided needs to be backed up; the backup situation of the to-be-processed data block is represented based on a bitmap, where 1 indicates that there is a data fingerprint corresponding to the to-be-processed data block in the preset fingerprint database, and 0 indicates that there is no data fingerprint corresponding to the to-be-processed data block in the preset fingerprint database.

[0025] Combined with the first aspect, in the fifth embodiment of the first aspect, after the step of storing and backing up the to-be-processed data blocks that have been determined to need to be backed up, the method further includes:

[0026] Update and store the data content corresponding to the already backed-up data block into the preset database, and update and store the data fingerprint corresponding to the already backed-up data block into the preset fingerprint database.

[0027] According to the second aspect, an embodiment of the present invention further provides a file storage and backup device, and the device includes:

[0028] A partitioning module, configured to determine that the backup data to be stored is not being transmitted for the first time, match the data content corresponding to the data blocks included in the backup data to be stored with a preset repository based on the RK algorithm, determine a chunk boundary, and partition the backup data to be stored based on the chunk boundary to obtain data blocks to be processed; the data content includes a data length and a hash value, and the preset repository stores the data content corresponding to the backed-up data blocks;

[0029] A determination module, configured to determine the hash value of the data block to be processed, and determine the data fingerprint of the data block to be processed based on the hash value;

[0030] A matching module, configured to match the data fingerprints corresponding to all the data blocks to be processed with a preset fingerprint database, and determine whether the data blocks to be processed need to be backed up according to the matching result; the preset fingerprint database stores the data fingerprints corresponding to the backed-up data blocks;

[0031] A backup module, configured to store and back up the data blocks to be processed that have been determined to need to be backed up.

[0032] According to a third aspect, an embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, the steps of any of the above-mentioned file storage and backup methods are implemented.

[0033] According to a fourth aspect, an embodiment of the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned file storage and backup methods are implemented.

[0034] According to a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned file storage and backup methods are implemented.

[0035] The file storage and backup method, device, equipment, and product provided by the present invention perform variable-length chunking on data according to the data content of the backup data to be stored based on the BK algorithm, making the chunking result of the backup data to be stored more reasonable, and the chunking result satisfying multiple file writing scenarios, effectively improving the data deduplication rate, reducing the space occupancy rate of the backup data, reducing the backup time, and bringing a better experience to users. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings. The drawings are schematic and should not be construed as imposing any limitation on the present invention. In the drawings:

[0037] Figure 1Shows one of the schematic flowcharts of the file storage and backup method provided by the present invention;

[0038] Figure 2 Shows another schematic flowchart of the file storage and backup method provided by the present invention;

[0039] Figure 3 Shows a third schematic flowchart of the file storage and backup method provided by the present invention;

[0040] Figure 4 Shows the schematic diagram of data block division for the file storage and backup method provided by the present invention in the overwrite write scenario;

[0041] Figure 5 Shows the schematic diagram of data block division for the file storage and backup method provided by the present invention in the insert write scenario;

[0042] Figure 6 Shows the schematic diagram of data block division for the file storage and backup method provided by the present invention in the delete write scenario;

[0043] Figure 7 Shows the schematic structural diagram of the file storage and backup device provided by the present invention;

[0044] Figure 8 Shows the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] Currently, the distributed storage architecture has become a preferred solution for cloud computing, and its file storage has advantages such as being shareable and inexpensive. In a file storage system, data is organized into files for storage and access. Among them, the cStor distributed cloud storage system can support protocols such as the Network File System (NFS) and the Common Internet File System (CIFS), and at the same time support multi-platform access, with advantages such as elastic scalability and easy user operation.

[0047] The cStor distributed cloud storage system backs up data through the following steps:

[0048] Chunk the file data by a fixed length. Specifically, to chunk the data by a fixed length, it is necessary to first specify a data length L. When the file data is first input, the file data is sliced into data chunks with a data length of L, and the data fingerprint corresponding to each data chunk is calculated based on the Hash value; when the file is input again, it is still sliced into data chunks with a data length of L, and the data fingerprint corresponding to each data chunk is calculated based on the hash value in the same way.

[0049] It can be seen that in the cStor distributed cloud storage system, in scenarios such as overwrite writing of file data, since the data length of the file data does not change, the data deduplication rate is relatively high and a good deduplication effect can be obtained. However, in the scenarios of insert writing and delete writing of file data, the number of fixed-length divided data chunks will increase / decrease, and the low data deduplication rate results in a poor deduplication effect. Therefore, the chunking method with a fixed length has a problem of a large difference in the deduplication effect in different scenarios, and it is more suitable for scenarios where the file length hardly changes, such as volume backup. As a result, the existing file data partitioning method has poor generality and cannot handle complex file storage scenarios.

[0050] To address the above problems, the following Figure 1 describes the file storage and backup method of the present invention. This method aims to solve the problems of low data deduplication rate and slow backup speed caused by using the data chunking method with a fixed length. The method includes:

[0051] S40. Determine that the data to be stored and backed up is not input for the first time. Based on the RK algorithm, match the data content corresponding to the data chunks in the data to be stored and backed up with the preset storage repository to determine the chunking boundary, and divide the data to be stored and backed up based on the chunking boundary to obtain the data chunks to be processed. In this application, the data content includes the data length and the Hash value, and the preset storage repository stores the data content corresponding to the backed-up data chunks, that is, the data length and the Hash value of the backed-up data chunks are stored in the preset storage repository. The data chunks to be processed are composed of several data chunks.

[0052] In this embodiment, for the first full backup, that is, the data to be stored and backed up input for the first time, the data is chunked according to the existing fixed length. For example, according to the specified preset data length L, when the data to be stored and backed up is input for the first time, the data to be stored and backed up is sliced into data chunks with a data length of L.

[0053] Preferably, the preset data length L can be set according to the total data length of the data to be stored and backed up, so that the data length of each data chunk obtained by the first division is equal, that is, nL is the total data of the data to be stored and backed up.

[0054] In some possible embodiments, the value of the preset data length L is a positive integer, such as 1, 2, 3..., n. That is, the preset data length L corresponds to the number of data blocks of the reference length. The reference length can be set in advance. For example, 32 bytes is used as a reference length. That is, when divided according to L = 4, each obtained data block contains 4 data blocks of the reference length and is composed of 4 data blocks of the reference length.

[0055] In this embodiment, for the to-be-stored backup data that is not transmitted for the first time, the block boundary includes a pair of (set) first block boundaries and second block boundaries. Among them, the first block boundary is the boundary between the foremost data block of the determined data blocks and the previous data block adjacent to the foremost data block, and the second block boundary is the boundary between the rearmost data block of the determined data blocks and the next data block adjacent to the rearmost data block. Only when the data length and Hash value of the data block in the to-be-stored backup data transmitted this time are both equal to those of the backed-up data block, at this time, the block boundary will be generated.

[0056] S50. Determine the Hash value of the to-be-processed data block, and determine the data fingerprint of the to-be-processed data block based on the Hash value. Use the Hash function MD5 to calculate and generate the unique data fingerprint FP of each reference data block.

[0057] S60. Match the data fingerprints corresponding to all the to-be-processed data blocks with the preset fingerprint database, and determine whether the to-be-processed data blocks need to be backed up according to the matching situation. In this embodiment, the preset fingerprint database stores the data fingerprints corresponding to the backed-up data blocks.

[0058] Specifically, the matching situations are as follows: There is a backed-up data block in the preset fingerprint database whose data fingerprint is exactly the same as that of the to-be-processed data block, that is, it can be matched with one of the backed-up data blocks. At this time, it indicates that the to-be-processed data block does not need to be backed up; there is no backed-up data block in the preset fingerprint database whose data fingerprint is exactly the same as that of the to-be-processed data block, that is, it cannot be matched with any of the backed-up data blocks. At this time, it indicates that the to-be-processed data block needs to be backed up.

[0059] S70. Store and back up the to-be-processed data blocks that have been determined to need to be backed up.

[0060] Based on the matching situation in step S60, determine the to-be-processed data blocks that need to be backed up, and perform this incremental backup.

[0061] The file storage and backup method provided by this application performs variable-length chunking on data based on the BK algorithm according to the data content of the data to be stored and backed up, making the chunking result of the data to be stored and backed up more reasonable, and the chunking result meets multiple file writing scenarios, effectively improving the data deduplication rate, reducing the space occupancy rate of the backup data, reducing the backup time, and bringing a better experience to users.

[0062] Please refer to Figures 2 to 6 , the method further includes the following steps:

[0063] S10. Determine that the data to be stored and backed up is first input. Based on the preset data length L, determine the chunking boundary, and divide the data to be stored and backed up based on the chunking boundary to obtain the reference data chunks.

[0064] Specifically, start dividing from the data block at the very front end of the data to be stored and backed up, and divide one reference data block every interval of the preset data length L. That is, Length(D 1 ) = Length(D 2 ) = … = Length(D N ) = L, where D i is the i-th reference data block divided.

[0065] In this embodiment, for the data to be stored and backed up that is first input, the chunking boundary is the boundary between the data blocks determined every interval of the preset data length L starting from the first data block (i.e., the data block corresponding to the nL-th data block) and the next data block adjacent to this / these data blocks. Divide the data to be stored and backed up from each chunking boundary to obtain n reference data blocks.

[0066] S20. Determine the Hash value of the reference data block, and determine the data fingerprint of the reference data block based on the Hash value.

[0067] After that, use the Hash function MD5 to calculate and generate the unique data fingerprint FP for each reference data block. The data fingerprint FP is specifically:

[0068] V i = Hash(D i )

[0069] where V i is the data fingerprint FP of the i-th reference data block.

[0070] In this embodiment, in order to reduce the subsequent calculation complexity, when calculating the Hash value corresponding to each reference data block, the Hash value is also modulo with a preset value H (i.e., a fixed value H), so that the finally calculated Hash value of the divided data blocks is distributed in the interval (0, H).

[0071] V1 = Hash((D 1 ) mod H, V 2 = Hash((D 2 ) mod H……

[0072] V i = Hash((D i ) mod H, V N = Hash((D N ) mod H

[0073] For each benchmark data block D obtained by partitioning i Its corresponding data length and Hash value can be obtained:

[0074] (V 1 , L 1 ), (V 2 , L 2 )……(V N , L N )

[0075] S30. Store and back up all benchmark data blocks, perform a full backup of the first incoming file data, store the data content corresponding to the backed-up data blocks in a preset database, and store the data fingerprints corresponding to the backed-up data blocks in a preset fingerprint database. The information stored in the preset database and the preset fingerprint database is used for the next incremental backup.

[0076] Please refer to Figure 3 , and in this method, step S40 specifically includes:

[0077] S41. Determine that the data to be stored and backed up is not the first incoming. Based on the RK algorithm for string matching, compare whether the data content corresponding to the data blocks included in the data to be stored and backed up is equal to the data content corresponding to the backed-up data blocks in the preset database.

[0078] S42. Determine that there are data blocks with equal data content, and determine the first block boundary and the second block boundary.

[0079] In this embodiment, for the data to be stored and backed up that is not the first incoming, the block boundaries include a pair of (set) first block boundaries and second block boundaries. Among them, the first block boundary is the boundary between the foremost data block of the determined data blocks and the previous data block adjacent to the foremost data block, and the second block boundary is the boundary between the rearmost data block of the determined data blocks and the next data block adjacent to the rearmost data block.

[0080] S43. Based on the first block boundary and the second block boundary, partition the data to be stored and backed up to obtain the data blocks to be processed.

[0081] Similar to the division of the first-incoming backup data to be stored, the division of the non-first-incoming backup data to be stored also starts from the frontmost data block. Since there is no data block in front of the frontmost data block, the first first-block boundary is the front boundary of the frontmost data block. When Length(f) = L 1 and Hash(f) mod H = V 1 a first second-block boundary will be generated, where f is at least one data block.

[0082] The next division or determination of the block boundary starts from the second-block boundary generated last time. When the following conditions are met:

[0083] Length(f) = L i and Hash(f) mod H = W i i ∈ (1, N) will generate a new pair of block boundaries. If Length(f) = L i does not satisfy Hash(f) mod H = V when i or Hash(f) mod H = V i does not satisfy Length(f) = L when i , no block boundary will be generated. And so on until the data lengths and data fingerprints FP of all the backed-up data blocks in the previous block division are traversed.

[0084] In this embodiment, the second-block boundary of a certain data block to be processed may coincide with the first-block boundary of the adjacent subsequent data block to be processed, and the first-block boundary of a certain data block to be processed may also coincide with the second-block boundary of the adjacent previous data block to be processed. Taking the first incremental backup after a full backup as an example for illustration, since during the full backup, the data lengths of each reference data block are equal, taking L i as 4 for example, when the data in the backup data to be stored changes and is transmitted again, if the data content of the first 2L does not change, then the first two data blocks to be processed obtained by division are the first two reference data blocks during the full backup. At this time, the second-block boundary of the first data block to be processed is also the first-block boundary of the second data block to be processed. For these two adjacent data blocks to be processed, the second-block boundary of the previous data block to be processed is also the first-block boundary of the subsequent data block to be processed, and these two block boundaries coincide. For the next incremental backup after an incremental backup, there are also two adjacent data blocks to be processed, where the second-block boundary of the previous data block to be processed is also the first-block boundary of the subsequent data block to be processed.

[0085] More specifically, please refer to Figures 4 to 6 In this method, step S43 includes the following steps:

[0086] S431. Determine that there is no block boundary before the first block boundary, and divide the data blocks before the first block boundary into a data block to be processed. If there is no block boundary before the first block boundary and there is also no data block, then the divided data block to be processed is empty, and it is also possible not to divide the data block in this special case.

[0087] S432. Determine that there is no block boundary after the second block boundary, and divide the data blocks after the second block boundary into a data block to be processed. If there is no block boundary after the second block boundary and there is also no data block, then the divided data block to be processed is empty, and it is also possible not to divide the data block in this special case.

[0088] S433. Determine that there is a block boundary before the second block boundary, and divide the data between the second block boundary and the previous first block boundary closest to the second block boundary into a data block to be processed.

[0089] Step S433 can be specifically divided into two cases: 1) Determine that there is a block boundary before the second block boundary, and the data between the second block boundary and the previous first block boundary closest to the second block boundary is not empty, and divide the data between these two block boundaries into a data block to be processed; 2) Determine that there is a block boundary before the second block boundary, and the data between the second block boundary and the previous first block boundary closest to the second block boundary is empty, then the divided data block to be processed is empty, and it is also possible not to divide the data block in this special case.

[0090] In the case of 1), when the data between the second block boundary and the previous first block boundary closest to the second block boundary is divided according to the reference length, the divided data blocks can be non-integers. For example, the data between these two block boundaries can contain 3.5 (reference length) data blocks and is composed of 3.5 (reference length) data blocks.

[0091] The following illustrates the data block division and storage backup of the present application through a specific application scenario.

[0092] In the scenario of overwrite writing, as Figure 4 shown, when the file data is first block-divided, four reference data blocks are generated, and their corresponding Hash values are: V 1 , V 2 , V 3 , V 4 . When the file data is transmitted again, due to overwrite writing, a new data block is generated, and the Hash value corresponding to this new data block is V 2′, but the rest of the data blocks exist and their Hash values and data lengths remain unchanged, so there is no need for re - backup. Therefore, in the case of overwrite writing, when the file data is transmitted again, four data blocks to be processed are still generated based on the RK algorithm, and three of them do not need to be backed up again. Only the new data block generated by the overwrite writing needs to be stored and backed up to complete the incremental backup, and the data deduplication rate is relatively low.

[0093] In the case of insert writing, as Figure 5 shown, when the file data is first chunked, four benchmark data blocks are generated, and their corresponding Hash values are: FP 1 , FP 2 , FP 3 , FP 4 . When the file data is transmitted again, since some data content is added, the data length of the file data changes compared with the full - volume backup. After re - chunking based on the RK algorithm, four data blocks to be processed with different lengths are generated. But except for the second data block (the Hash value of this second data block is FP 2 ′), the rest of the data blocks exist and their Hash values and data lengths remain unchanged, so there is no need for re - backup. Therefore, in the case of insert writing, when the file data is transmitted again, four data blocks to be processed are still generated based on the RK algorithm, and three of them do not need to be backed up again. Only the data block with the changed data length generated by the insert writing needs to be stored and backed up to complete the incremental backup, and the data deduplication rate is relatively low.

[0094] In the case of delete writing, as Figure 6 shown, when the file data is first chunked, four benchmark data blocks are generated, and their corresponding Hash values are: FP 1 , FP 2 , FP 3 , FP 4 . When the file data is transmitted again, since some data content is deleted, the data length of the file data changes compared with the full - volume backup. After re - chunking based on the RK algorithm, four data blocks to be processed with different lengths are also generated. But except for the third data block (the Hash value of this third data block is FP 3 ′), the rest of the data blocks exist and their Hash values and data lengths remain unchanged, so there is no need for re - backup. Therefore, in the case of delete writing, when the file data is transmitted again, four data blocks to be processed are still generated based on the RK algorithm, and three of them do not need to be backed up again. Only the data block with the changed data length generated by the delete writing needs to be stored and backed up to complete the incremental backup, and the data deduplication rate is relatively low.

[0095] After each storage backup of file data, the next incremental backup is based on the data generated by the previous full backup / incremental backup. Therefore, this method further includes:

[0096] S80. Update and store the data content corresponding to the backed-up data block in a preset database, and update and store the data fingerprint corresponding to the backed-up data block in a preset fingerprint database.

[0097] That is, after each backup of the data to be stored and backed up is completed, the data content and data fingerprint of each divided reference data block / processing data block are updated.

[0098] In this embodiment, the Hash value obtained by taking the remainder of the data block by a preset value H can be used as the data fingerprint FP of the data block.

[0099] Currently, the cStor distributed cloud storage system uses a linear query method to query the data fingerprints of data blocks one by one to return query results to determine whether a data block needs to be backed up. Such a query method results in low query efficiency, long file backup time, and low backup efficiency. Considering the above problems, step S60 in this method specifically includes:

[0100] S61. Pass the data to be stored and backed up into the fingerprint database.

[0101] S62. Compare whether the data fingerprints corresponding to all processing data blocks are equal to the data fingerprints corresponding to the backed-up data blocks in the preset fingerprint database, and determine whether each processing data block into which the data to be stored and backed up is divided needs to be backed up. In this embodiment, the backup status of the processing data block is represented based on a bitmap. 1 means that there is a data fingerprint corresponding to the processing data block in the preset fingerprint database, that is, 1 means that storage backup is not required; 0 means that there is no data fingerprint corresponding to the processing data block in the preset fingerprint database, that is, 0 means that storage backup is required.

[0102] Send the file to be stored and backed up into the preset fingerprint database, and then batch retrieve all the divided processing data blocks of the file to be stored and backed up into the preset fingerprint database. Use a bitmap to indicate whether the corresponding processing data block has existed before: 1 means it exists, 0 means it does not exist, and the results are returned simultaneously.

[0103] For example, if file A is divided into ten data blocks and the 3rd and 9th data blocks have not been backed up, the returned result is: 1101111101.

[0104] In this application, the problem of low query efficiency caused by serial query is optimized. The bitmap is used to represent whether the logical position data block exists, and concurrent query and batch result return can be performed, greatly improving the query efficiency, shortening the upload time of the backup, accelerating the backup speed, and improving the security and reliability of file data.

[0105] The file storage and backup device provided by the present invention will be described below. The file storage and backup device described below can be correspondingly referred to the file storage and backup method described above.

[0106] The following combines Figure 7 The file storage and backup device of the present invention will be described. The device aims to solve the problems of low data deduplication rate and slow backup speed caused by the fixed-length data chunking method. The device includes:

[0107] A partitioning module 40, which is used to determine that the data to be stored and backed up is not the first time it is input, match the data content corresponding to the data block in the data to be stored and backed up with a preset storage repository based on the RK algorithm, determine the chunking boundary, and partition the data to be stored and backed up based on the chunking boundary to obtain data blocks to be processed. In this application, the data content includes the data length and the Hash value, and the data content corresponding to the backed-up data block is stored in the preset storage repository, that is, the data length and the Hash value of the backed-up data block are stored in the preset storage repository.

[0108] In this embodiment, for the first full backup, that is, the data to be stored and backed up input for the first time, the data is chunked according to the existing fixed length. For example, according to the specified preset data length L, when the data to be stored and backed up is input for the first time, the data to be stored and backed up is cut into data blocks with a data length of L.

[0109] Preferably, the data length L can be set according to the total data length of the data to be stored and backed up, so that the data length of each data block obtained by the first partitioning is equal, that is, nL is the total data of the data to be stored and backed up.

[0110] In this embodiment, for the data to be stored and backed up that is not input for the first time, the chunking boundary includes a pair of first chunking boundaries and second chunking boundaries (set). Among them, the first chunking boundary is the boundary between the frontmost data block of the determined data block and the previous data block adjacent to the frontmost data block, and the second chunking boundary is the boundary between the rearmost data block of the determined data block and the next data block adjacent to the rearmost data block. Only when the data length and the Hash value corresponding to the data block in the data to be stored and backed up input this time are both equal to the data length and the Hash value corresponding to the backed-up data block, at this time, the chunking boundary will be generated.

[0111] A determination module 50 is configured to determine the Hash value of a data block to be processed and determine the data fingerprint of the data block to be processed based on the Hash value. The data fingerprint FP unique to each reference data block is calculated and generated by using the Hash function MD5.

[0112] A matching module 60 is configured to match the data fingerprints corresponding to all data blocks to be processed with a preset fingerprint database, and determine whether the data blocks to be processed need to be backed up according to the matching result. In this embodiment, the preset fingerprint database stores the data fingerprints corresponding to the backed-up data blocks.

[0113] Specifically, the matching results are divided into: there is a backed-up data block in the preset fingerprint database whose data fingerprint is exactly the same as that of the data block to be processed, that is, it can be matched with one of the backed-up data blocks. At this time, it indicates that the data block to be processed does not need to be backed up; there is no backed-up data block in the preset fingerprint database whose data fingerprint is exactly the same as that of the data block to be processed, that is, it cannot be matched with any backed-up data block. At this time, it indicates that the data block to be processed needs to be backed up.

[0114] A backup module 70 is configured to store and back up the data blocks to be processed that have been determined to need backup.

[0115] Based on the matching result of the matching module 60, determine the data blocks to be processed that need to be backed up, and perform the incremental backup for this time.

[0116] The file storage and backup device provided by this application divides data into variable-length blocks according to the data content of the data to be stored and backed up based on the BK algorithm, making the block division result of the data to be stored and backed up more reasonable, and the block division result meets various file writing scenarios, effectively improving the data deduplication rate, reducing the space occupancy rate of the backup data, reducing the backup time, and bringing a better experience to users.

[0117] Figure 8 Illustrates a schematic physical structure diagram of an electronic device, as Figure 8 shown. The electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the file storage and backup method, and the method includes:

[0118] Determine that the backup data to be stored is not the first time it is introduced. Based on the RK algorithm, match the data content corresponding to the data blocks included in the backup data to be stored with a preset repository to determine the block boundaries, and divide the backup data to be stored based on the block boundaries to obtain data blocks to be processed; the data content includes the data length and the hash value, and the preset repository stores the data content corresponding to the backed-up data blocks;

[0119] Determine the hash value of the data block to be processed, and determine the data fingerprint of the data block to be processed based on the hash value;

[0120] Match the data fingerprints corresponding to all the data blocks to be processed with a preset fingerprint database, and determine whether the data blocks to be processed need to be backed up according to the matching situation; the preset fingerprint database stores the data fingerprints corresponding to the backed-up data blocks;

[0121] Store and back up the data blocks to be processed that have been determined to need to be backed up.

[0122] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0123] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the file storage and backup method provided by the above-mentioned various methods. The method includes:

[0124] Determine that the backup data to be stored is not the first time it is introduced. Based on the RK algorithm, match the data content corresponding to the data blocks included in the backup data to be stored with a preset repository to determine the block boundaries, and divide the backup data to be stored based on the block boundaries to obtain data blocks to be processed; the data content includes the data length and the hash value, and the preset repository stores the data content corresponding to the backed-up data blocks;

[0125] Determine the hash value of the data block to be processed, and determine the data fingerprint of the data block to be processed based on the hash value;

[0126] Match the data fingerprints corresponding to all the data blocks to be processed with a preset fingerprint database, and determine whether the data blocks to be processed need to be backed up according to the matching situation; the preset fingerprint database stores the data fingerprints corresponding to the backed-up data blocks;

[0127] Store and back up the data blocks to be processed that have been determined to need backup.

[0128] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a file storage and backup method provided by the above-mentioned various methods. The method includes:

[0129] Determine that the data to be stored and backed up is not the first time it is introduced. Based on the RK algorithm, match the data content corresponding to the data blocks included in the data to be stored and backed up with a preset storage repository, determine the block boundary, and divide the data to be stored and backed up based on the block boundary to obtain data blocks to be processed; the data content includes the data length and the hash value, and the preset storage repository stores the data content corresponding to the backed-up data blocks;

[0130] Determine the hash value of the data block to be processed, and determine the data fingerprint of the data block to be processed based on the hash value;

[0131] Match the data fingerprints corresponding to all the data blocks to be processed with a preset fingerprint database, and determine whether the data blocks to be processed need to be backed up according to the matching situation; the preset fingerprint database stores the data fingerprints corresponding to the backed-up data blocks;

[0132] Store and back up the data blocks to be processed that have been determined to need backup.

[0133] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0134] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0135] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A file storage and backup method, characterized in that, the method includes: Determine that the data to be stored and backed up is not the first time it is transmitted. Based on the RK algorithm, match the data content corresponding to the data blocks included in the data to be stored and backed up with a preset repository, determine the block boundaries, and divide the data to be stored and backed up based on the block boundaries to obtain data blocks to be processed, including: Determine that the data to be stored and backed up is not the first time it is transmitted. Based on the RK algorithm, compare whether the data content corresponding to the data blocks included in the data to be stored and backed up is equal to the data content corresponding to the already backed-up data blocks in the preset repository; Determine that there are data blocks with equal data content, and determine the first block boundary and the second block boundary; The first block boundary is the boundary between the foremost data block among the determined data blocks and the previous data block adjacent to the foremost data block, and the second block boundary is the boundary between the rearmost data block among the determined data blocks and the next data block adjacent to the rearmost data block. The first block boundary and the second block boundary appear in pairs; Based on the first block boundary and the second block boundary, divide the data to be stored and backed up to obtain data blocks to be processed; The data content includes the data length and the hash value, and the preset repository stores the data content corresponding to the already backed-up data blocks; The data blocks to be processed are composed of several data blocks; Determine the hash value of the data blocks to be processed, and determine the data fingerprint of the data blocks to be processed based on the hash value; Match the data fingerprints corresponding to all the data blocks to be processed with a preset fingerprint database, and determine whether the data blocks to be processed need to be backed up according to the matching situation; The preset fingerprint database stores the data fingerprints corresponding to the already backed-up data blocks; Store and back up the data blocks to be processed that have been determined to need to be backed up.

2. The file storage and backup method according to claim 1, characterized in that, before the step of determining that the data to be stored and backed up is not the first time it is transmitted, based on the RK algorithm, matching the data content corresponding to the data blocks included in the data to be stored and backed up with a preset repository, determining the block boundaries, and dividing the data to be stored and backed up based on the block boundaries to obtain data blocks to be processed, the method further includes the following steps: Determine that the data to be stored and backed up is the first time it is transmitted. Based on a preset data length, determine the block boundaries, and divide the data to be stored and backed up based on the block boundaries to obtain reference data blocks; The reference data blocks are composed of several data blocks; Determine the hash value of the reference data blocks, and determine the data fingerprint of the reference data blocks based on the hash value; Store and back up all the reference data blocks, store the data content corresponding to the already backed-up data blocks in the preset repository, and store the data fingerprints corresponding to the already backed-up data blocks in the preset fingerprint database.

3. The file storage and backup method according to claim 1, characterized in that, the step of dividing the data to be stored and backed up based on the first block boundary and the second block boundary to obtain data blocks to be processed specifically includes: Before determining the first block boundary and when there is no block boundary, divide the data block before the first block boundary into a to-be-processed data block; After determining the second block boundary and when there is no block boundary, divide the data block after the second block boundary into a to-be-processed data block; Before determining the second block boundary and when there is a block boundary, divide the data between the second block boundary and the nearest previous first block boundary to the second block boundary into the to-be-processed data block.

4. The file storage and backup method according to claim 1, wherein, match the data fingerprints corresponding to all the to-be-processed data blocks with a preset fingerprint database, and determine whether the to-be-processed data blocks need to be backed up according to the matching situation; the preset fingerprint database stores the data fingerprints corresponding to the already backed-up data blocks, specifically including: transmit the data to be stored and backed up into the fingerprint database; compare whether the data fingerprints corresponding to all the to-be-processed data blocks are equal to the data fingerprints corresponding to the already backed-up data blocks in the preset fingerprint database, and determine whether each of the to-be-processed data blocks into which the data to be stored and backed up is divided needs to be backed up; the backup situation of the to-be-processed data blocks is represented based on a bitmap, where 1 indicates that there is a data fingerprint corresponding to the to-be-processed data block in the preset fingerprint database, and 0 indicates that there is no data fingerprint corresponding to the to-be-processed data block in the preset fingerprint database.

5. The file storage and backup method according to claim 1, wherein, after the step of storing and backing up the to-be-processed data blocks that have been determined to need to be backed up, the method further includes: update and store the data content corresponding to the already backed-up data blocks into the preset storage repository, and update and store the data fingerprints corresponding to the already backed-up data blocks into the preset fingerprint database.

6. A file storage and backup device, wherein, the device includes: A partitioning module, configured to determine that the backup data to be stored is not being transmitted for the first time, match the data content corresponding to the data blocks included in the backup data to be stored with a preset storage repository based on the RK algorithm, determine chunk boundaries, and partition the backup data to be stored based on the chunk boundaries to obtain data blocks to be processed, including: determining that the backup data to be stored is not being transmitted for the first time, and based on the RK algorithm, comparing whether the data content corresponding to the data blocks included in the backup data to be stored is equal to the data content corresponding to the already backed-up data blocks in the preset storage repository; determining that there are data blocks with equal data content, and determining a first chunk boundary and a second chunk boundary; the first chunk boundary is the boundary between the foremost data block among the already determined data blocks and the previous data block adjacent to the foremost data block, and the second chunk boundary is the boundary between the rearmost data block among the already determined data blocks and the next data block adjacent to the rearmost data block, and the first chunk boundary and the second chunk boundary appear in pairs; partitioning the backup data to be stored based on the first chunk boundary and the second chunk boundary to obtain data blocks to be processed; The data content includes a data length and a hash value, and the preset storage repository stores the data content corresponding to the already backed-up data blocks; the data blocks to be processed are composed of a plurality of data blocks; A determination module, configured to determine the hash value of the data blocks to be processed, and determine the data fingerprint of the data blocks to be processed based on the hash value; A matching module, configured to match the data fingerprints corresponding to all the data blocks to be processed with a preset fingerprint database, and determine whether the data blocks to be processed need to be backed up according to the matching result; the preset fingerprint database stores the data fingerprints corresponding to the already backed-up data blocks; A backup module, configured to store and back up the data blocks to be processed that have been determined to need to be backed up.

7. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, the steps of the file storage and backup method according to any one of claims 1 to 5 are implemented.

8. A non-transitory computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the steps of the file storage and backup method according to any one of claims 1 to 5 are implemented.

9. A computer program product, including a computer program, wherein, when the computer program is executed by a processor, the steps of the file storage and backup method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Scalable deduplication system with small blocks

    CN103814361A

  • A data backup and recovery method based on source data re-deletion

    CN109101365A