File deduplication method and system, electronic device and computer storage medium
Through grouping based on file attribute parameters and filtering of preset duplication conditions, the problem of difficult to guarantee file deduplication efficiency and accuracy in the prior art is solved, and efficient and accurate file deduplication operation is achieved.
Patent Information
- Application Number
- PCT/CN2024/093400
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-02
- Filing Date
- 2024-05-15
- Publication Date
- 2025-05-08
AI Technical Summary
The prior art is difficult to ensure efficiency and accuracy in the process of file deduplication, especially when the number of files is large and uncontrollable, it is easy to cause memory overflow, large amount of calculation and misjudgment.
By grouping files according to the attribute parameters of the file, and filtering out file groups that need to be deduplicated in combination with preset duplication conditions, this method can reduce file groups that do not need to be deduplicated and improve deduplication efficiency and accuracy.
This method can improve the accuracy and efficiency of file deduplication, reduce the amount of calculation, avoid misjudgment, and realize parallel filtering operations for file groups for deduplication.
Smart Images

Figure CN2024093400_08052025_PF_FP_ABST
Abstract
Description
File deduplication method and system, electronic device, and computer storage medium Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a file deduplication method and system, electronic equipment, and computer storage medium. Background Art
[0002] With the continuous development of internet technology and the rise of cloud computing, the number of files that need to be stored is increasing, and their variety is also increasing. However, as the number of files continues to grow, storage systems often contain a large number of duplicate files, such as multiple identical files uploaded from the same source or from different sources. This large number of duplicate files consumes a lot of memory and hinders various file operations. Therefore, file deduplication is essential.
[0003] There are three commonly used deduplication algorithms in the existing technology: the longest common substring method, the cosine similarity method, and the Simhash algorithm. In practice, the longest common substring method is prone to memory overflow when the number of files is large and uncontrollable, which greatly reduces deduplication efficiency. The cosine similarity method uniformly vectorizes all files. Once a new file is added, all files must be revectorized. When deduplicating a large number of files, a large number of sparse vectors will be generated, resulting in a large amount of computation and hindering the efficiency of file deduplication. The Simhash algorithm causes partial loss of file information during the dimensionality reduction process, which is prone to misjudgment. In particular, shorter files may be mistakenly identified as having high similarity, which is not conducive to improving the accuracy of file deduplication.
[0004] In summary, for storage systems that store a large number of files, it is difficult to ensure the efficiency and accuracy of file deduplication during the file deduplication process. Therefore, it is particularly important to propose a new file deduplication method to improve the efficiency and accuracy of file deduplication.
[0005] Summary of the Invention
[0006] The present invention provides a file deduplication method and system, electronic equipment, and computer storage medium, which can help improve the efficiency and accuracy of file deduplication.
[0007] In order to solve the above technical problems, the first aspect of the present invention discloses a file deduplication method, which includes:
[0008] For a plurality of files to be deduplicated, grouping all the files according to a first attribute parameter of each file to obtain a plurality of first file groups;
[0009] Filtering out at least one second file group that meets a first preset repetition condition from all the first file groups;
[0010] For each of the second file groups, grouping all files in the second file group according to the second attribute parameter of each file in the second file group to obtain at least one sub-file group corresponding to the second file group;
[0011] For each of the second file groups, according to the second preset duplication condition, a sub-file group to be deduplicated is screened from all sub-file groups corresponding to the second file group, wherein the sub-file group to be deduplicated is a file group on which a file deduplication operation is required;
[0012] Perform a file deduplication operation on each of the file groups to be deduplicated.
[0013] A second aspect of the present invention discloses an electronic device, comprising:
[0014] a memory storing executable program code;
[0015] a processor coupled to the memory;
[0016] The processor calls the executable program code stored in the memory to execute the file deduplication method disclosed in the first aspect of the present invention.
[0017] A third aspect of the present invention discloses a computer storage medium, wherein the computer storage medium stores computer instructions. When the computer instructions are called, they are used to execute the file deduplication method disclosed in the first aspect of the present invention.
[0018] The fourth aspect of the present invention discloses a file deduplication system, which includes at least the electronic device described in the second aspect of the present invention, and a NAS device communicatively connected to the electronic device, and a number of files to be deduplicated are stored on the NAS device. When the electronic device obtains the several files to be deduplicated, the electronic device performs deduplication operations on the several files to be deduplicated according to any one of the file deduplication methods described in the first aspect of the present invention.
[0019] Compared with the prior art, the present invention has the following beneficial effects:
[0020] In the present invention, for a number of files, all files can be grouped according to the first attribute parameter of each file, and then at least one second file group that meets the first preset duplication condition can be screened out, that is: after a grouping based on the attribute parameters of the file is implemented, the file group that does not need to be deduplicated is eliminated, and only the file group that needs to be deduplicated is screened out, which is conducive to reducing the subsequent file volume. After eliminating the file group that does not need to be deduplicated, all files in each remaining file group are grouped according to the second attribute parameter of each file in the remaining file group, and then the file group that needs to be deduplicated is finally screened out according to the second preset duplication condition. This method of implementing file grouping based on the attribute parameters of the file and determining the file group to be deduplicated in combination with the preset duplication condition is conducive to improving the accuracy of the determined file group to be deduplicated, and thus is conducive to improving the accuracy and efficiency of file deduplication. In addition, based on the grouping method, parallel screening operations of the file group to be deduplicated can also be implemented, which is conducive to improving the efficiency of determining the file group to be deduplicated, and thus is conducive to further improving the efficiency of file deduplication. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] FIG1 is a schematic diagram of a flow chart of a file deduplication method disclosed in an embodiment of the present invention;
[0023] FIG2 is a flow chart of another method for deduplicating files disclosed in an embodiment of the present invention;
[0024] FIG3 is a schematic structural diagram of an electronic device disclosed in an embodiment of the present invention;
[0025] FIG4 is a schematic structural diagram of a file deduplication system disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The present invention discloses a file deduplication method and system, an electronic device, and a computer storage medium. These methods can group files based on their attribute parameters and determine a file group to be deduplicated in combination with a preset duplication condition, thereby improving the accuracy of the determined file group to be deduplicated, and thereby improving the accuracy and efficiency of file deduplication. Furthermore, based on the grouping method, parallel screening operations can be performed on the file group to be deduplicated, thereby improving the efficiency of determining the file group to be deduplicated, and thereby further improving the efficiency of file deduplication. Each of these methods is described in detail below.
[0027] Example 1
[0028] Please refer to Figure 1, which is a flowchart of a file deduplication method disclosed in an embodiment of the present invention. The file deduplication method described in Figure 1 can be applied to a file deduplication device, which is used to implement the screening of the file group to be deduplicated or further implement file deduplication after implementing the screening of the file group to be deduplicated. Optionally, the file deduplication device can be integrated into a NAS device, or can be integrated into other electronic devices that are connected to the NAS device and are used to deduplicate files in the NAS device, which is beneficial for ensuring the security of data in the file while implementing file deduplication. As shown in Figure 1, the file deduplication method may include:
[0029] 101. For a plurality of files to be deduplicated, group all the files according to a first attribute parameter of each file to obtain a plurality of first file groups.
[0030] In an embodiment of the present invention, optionally, for a plurality of files to be deduplicated, each file may have at least corresponding basic attribute parameters. Optionally, the basic attribute parameters corresponding to the file may include a combination of one or more of the following: file size, file modification time, file source, and file unique identifier. Further optionally, each file may also have corresponding extended attribute parameters. The extended attribute parameters corresponding to the file may include the file modification time and / or a checksum parameter. Further optionally, the checksum parameter of the file may include two checksum parameters, which may be an odd checksum value and an even checksum value, respectively.
[0031] In an embodiment of the present invention, for multiple files to be deduplicated, the file deduplication device can first quickly store all the files to be deduplicated and record the corresponding file modification time in the extended attribute parameters corresponding to the files. The file modification time in the extended attribute parameters corresponding to the files is mainly used to determine whether the file needs to recalculate the verification parameters during the next deduplication calculation, which is conducive to ensuring the accuracy of the file's verification parameters.
[0032] In the embodiment of the present invention, file duplication can be understood as files that are identical or highly similar, which means that one of the conditions for file duplication is that the file sizes are identical or very close. Based on this, the first attribute parameter of the file can be the file size, and in actual applications, basic attribute parameters including file size can be directly obtained through the API interface of the storage system. Optionally, for multiple files to be deduplicated, all files are grouped according to the first attribute parameter of each file to obtain multiple first file groups, which may include:
[0033] For multiple files to be deduplicated, all files are grouped according to a preset file size grouping condition and the file size of each file to obtain at least one first file group. For any two files that do not belong to the same first file group, the difference in file size between the two files is greater than a preset difference threshold. When the number of files in a first file group is greater than or equal to 2, for any two files in the first file group, the difference in file size between the two files is less than or equal to a preset difference threshold. Optionally, the preset difference threshold is preferably 0. This file size-based grouping method is advantageous for grouping files with a certain probability of duplication, thereby facilitating the subsequent rapid and accurate screening of the second file group.
[0034] Optionally, for the plurality of files to be deduplicated, before grouping all the files according to the first attribute parameter of each file to obtain a plurality of first file groups, the method may further include the following operations:
[0035] Non-critical content filtering is performed on the file content of each file to update each file. This helps reduce the impact of non-critical content on file size when grouping files based on file size, thereby improving the accuracy of file grouping based on file size.
[0036] 102. Filter out at least one second file group that meets a first preset repetition condition from all first file groups.
[0037] In an embodiment of the present invention, as an optional implementation, the above-mentioned step of selecting at least one second file group that meets the first preset repetition condition from all first file groups may include:
[0038] If there are target file groups with a file quantity greater than or equal to a first preset quantity threshold in all first file groups, all target file groups are determined as second file groups that meet the first preset duplication condition.
[0039] In this optional embodiment, the first preset number threshold is an integer greater than or equal to 2, preferably 2, which is conducive to improving the comprehensiveness and accuracy of the screened second file group.
[0040] As can be seen, this optional embodiment can eliminate file groups that do not require deduplication based on the number of files in the first file group, which is beneficial for reducing the number of files in subsequent file screening operations, thereby improving the efficiency of subsequent file screening. In addition, the first preset number threshold is preferably 2, which is beneficial for improving the comprehensiveness and accuracy of the second file group screened for subsequent processing, thereby improving the accuracy of the final determined file group to be deduplicated.
[0041] 103. For each second file group, group all files in the second file group according to the second attribute parameter of each file in the second file group to obtain at least one sub-file group corresponding to the second file group.
[0042] In this embodiment of the present invention, for any second file group obtained after step 102, all files in the second file group have the same or highly similar sizes, and the number of files in the second file group is greater than or equal to the first preset number threshold. Furthermore, for each second file group, each file in the second file group is grouped according to its second attribute parameter to obtain multiple file groups under the second file group. For ease of distinction, the multiple file groups under each second file group can be understood as sub-file groups corresponding to each second file group.
[0043] In an embodiment of the present invention, optionally, the second attribute parameter of each file may include a first check parameter and a second check parameter. Further, the first check parameter may be an odd check value and the second check parameter may be an even check value, or the first check parameter may be an even check value and the second check parameter may be an odd check value.
[0044] In an embodiment of the present invention, as an optional implementation, for each second file group, grouping all files in the second file group according to the second attribute parameter of each file in the second file group to obtain at least one sub-file group corresponding to the second file group may include:
[0045] For each second file group, all files in the second file group are grouped according to a preset verification parameter grouping condition and the first verification parameter and the second verification parameter of each file in the second file group to obtain at least one sub-file group; the preset verification parameter grouping condition is used to indicate that all files in the second file group corresponding to the same first verification parameter and the same second verification parameter are divided into the same sub-file group.
[0046] In this optional embodiment, based on the aforementioned step 102, it can be seen that all files in the second file group have the same or highly similar file sizes, preferably completely identical file sizes. The files in the second file group are only files with a certain possibility of being duplicates. After obtaining the second file group, for any second file group, further judgment is made based on the two verification parameters of the files. For truly duplicate files, the two verification parameters of the files are respectively the same. In this optional embodiment, files with the same verification parameters are divided into a sub-file group. In other words, for files to be duplicates of each other, the following three conditions must be met simultaneously: the file sizes are the same (or highly similar); the first verification parameters are the same between the files; and the second verification parameters are the same between the files.
[0047] For example, assuming that a second file group includes file 1, file 2, file 3, and file 4, the verification parameters of file 1 and file 2 are respectively the same (that is, the first verification parameter of file 1 is the same as the first verification parameter of file 2, and the second verification parameter of file 1 is the same as the second verification parameter of file 2), and the verification parameter of file 3 is not the same as the verification parameter of file 4, the verification parameter of file 3 is not the same as the verification parameter of files 1 and 2, and the verification parameter of file 4 is not the same as the verification parameter of files 1 and 2, then the sub-file groups corresponding to the second file group may include:
[0048] Sub-file group A consisting of file 1 and file 2; or,
[0049] File 1 and file 2 form sub-file group A, file 3 itself forms sub-file group B, and file 4 itself forms sub-file group C.
[0050] It should be noted that if a file in a sub-file group is unique, it means that there is no duplicate file corresponding to the file, that is, the sub-file group is not an object that needs to be deduplicated.
[0051] Further optionally, for each second file group, all files in the second file group are grouped according to a preset verification parameter grouping condition and the first verification parameter and the second verification parameter of each file in the second file group to obtain at least one sub-file group, which may include:
[0052] According to the first verification parameter (or second verification parameter) of each file in the second file group, all files with the same first verification parameter (or second verification parameter) are grouped into the same file group to obtain multiple third file groups, thereby eliminating files with unique first verification parameters (or second verification parameters);
[0053] For each third file group, based on the second verification parameter (or first verification parameter) of each file in the third file group, all files with the same second verification parameter (or first verification parameter) are grouped into the same file group, and each file with a unique second verification parameter (or first verification parameter) is grouped into an independent file group, thereby obtaining multiple sub-file groups; or,
[0054] For each third file group, according to the second verification parameter (or first verification parameter) of each file in the third file group, all files with the same second verification parameter (or first verification parameter) are divided into the same file group to eliminate files with unique first verification parameters.
[0055] It can be seen that in this optional embodiment, since each file has two verification parameters, when further grouping, hierarchical grouping can be performed according to the two verification parameters respectively, that is: after grouping or group screening according to one of the verification parameters to obtain the grouping results, the aforementioned grouping results are further grouped according to the other verification parameter, which is conducive to improving the grouping accuracy to a certain extent.
[0056] 104. For each second file group, according to a second preset duplication condition, a sub-file group to be deduplicated is screened from all sub-file groups corresponding to the second file group, where the sub-file group to be deduplicated is the file group requiring a file deduplication operation.
[0057] In the embodiment of the present invention, optionally, for each second file group, filtering out a target sub-file group to be deduplicated from all sub-file groups corresponding to the second file group according to the second preset duplication condition may include:
[0058] For each second file group, if there is a target sub-file group with a file number greater than or equal to a second preset number threshold among the sub-file groups corresponding to the second file group, all target sub-file groups are determined as sub-file groups to be deduplicated. The second preset number threshold is an integer greater than or equal to 2, preferably 2, which helps improve the accuracy of screening the sub-file groups to be deduplicated.
[0059] 105. Perform file deduplication operations on each file group to be deduplicated.
[0060] It should be noted that when there are multiple second file groups mentioned above, the files in each second file group can be grouped in parallel, which is conducive to improving the grouping efficiency; alternatively, the files in some of the second file groups can be grouped first, and after the grouping, the files in another part of the second file groups can be grouped, and so on. This serial grouping method can improve the grouping efficiency to a certain extent, and can also reduce the occurrence of inaccurate grouping due to insufficient processing resources of the file deduplication device and a large number of second file groups that need to be grouped.
[0061] It can be seen that the implementation of the method described in the embodiment of the present invention can realize the grouping of files based on the attribute parameters of the files and determine the file group to be deduplicated in combination with the preset duplication conditions, which is conducive to improving the accuracy of the determined file group to be deduplicated, and further conducive to improving the accuracy and efficiency of file deduplication. In addition, based on the grouping method, it is also possible to realize the parallel screening operation of the file group to be deduplicated, which is conducive to improving the efficiency of determining the file group to be deduplicated, and further conducive to further improving the efficiency of file deduplication.
[0062] Example 2
[0063] Please refer to Figure 2, which is a flowchart of another file deduplication method disclosed in an embodiment of the present invention. The file deduplication method described in Figure 2 can be applied to a file deduplication device, which is used to implement the screening of the file group to be deduplicated or further implement file deduplication after implementing the screening of the file group to be deduplicated. Optionally, the file deduplication device can be integrated into a NAS device, or can be integrated into other electronic devices that are connected to the NAS device and are used to deduplicate files in the NAS device, which is beneficial for ensuring the security of data in the file while implementing file deduplication. As shown in Figure 2, the file deduplication method may include:
[0064] 201. For a plurality of files to be deduplicated, group all the files according to a first attribute parameter of each file to obtain a plurality of first file groups.
[0065] 202. Filter out at least one second file group that meets a first preset repetition condition from all first file groups.
[0066] In the embodiment of the present invention, for the description of steps 201-202, please refer to the description of steps 101-102 in the first embodiment, and the embodiment of the present invention will not repeat them.
[0067] 203. For each second file group, group all files in the second file group according to a preset verification parameter grouping condition and the first verification parameter and the second verification parameter of each file in the second file group to obtain at least one sub-file group.
[0068] The preset verification parameter grouping condition is used to indicate that all files in the second file group corresponding to the same first verification parameter and the same second verification parameter are grouped into the same sub-file group.
[0069] 204. For each second file group, if there is a target sub-file group with a file quantity greater than or equal to a second preset quantity threshold in the sub-file group corresponding to the second file group, determine all target sub-file groups as sub-file groups to be deduplicated.
[0070] 205. Perform a file deduplication operation on each file group to be deduplicated.
[0071] It should be noted that, for other related descriptions of step 203 and step 204, please refer to the corresponding descriptions in the first embodiment, and the embodiment of the present invention will not repeat them.
[0072] It can be seen that the implementation of the method described in the embodiment of the present invention can realize the grouping of files based on the attribute parameters of the files and determine the file group to be deduplicated in combination with the preset duplication conditions, which is conducive to improving the accuracy of the determined file group to be deduplicated, and further conducive to improving the accuracy and efficiency of file deduplication. In addition, based on the grouping method, it is also possible to realize the parallel screening operation of the file group to be deduplicated, which is conducive to improving the efficiency of determining the file group to be deduplicated, and further conducive to further improving the efficiency of file deduplication. In addition, after determining the file group to be deduplicated, it is also possible to realize the file deduplication operation of the file group to be deduplicated, further expanding the intelligent function of the file deduplication equipment, and is conducive to improving the efficiency and convenience of file deduplication.
[0073] In an optional embodiment, performing a file deduplication operation on each to-be-deduplicated file group may include:
[0074] For each sub-file group to be deduplicated, the deduplication strategy corresponding to the file group to be deduplicated is determined based on the number of files in the file group to be deduplicated, the basic attribute parameters and extended attribute parameters of each file in the file group to be deduplicated, and the deduplication operation is performed on the file group to be deduplicated according to the deduplication strategy corresponding to the file group to be deduplicated.
[0075] It can be seen that this optional embodiment can match the deduplication strategy according to the multi-dimensional parameters of the file group to be deduplicated when implementing file deduplication, thereby realizing personalized file deduplication operations.
[0076] In another optional embodiment, before step 203, the method may further include the following operations:
[0077] For each file in each second file group, a first verification parameter and a second verification parameter of the file are determined respectively.
[0078] In this optional embodiment, for each file in each second file group, determining the first verification parameter and the second verification parameter of the file may include:
[0079] For each file in each second file group, if the file modification time in the basic attribute parameters corresponding to the file is inconsistent with the file modification time in the extended attribute parameters corresponding to the file, the file modification time in the extended attribute parameters corresponding to the file is updated according to the file modification time in the basic attribute parameters corresponding to the file, and after the update, the first verification parameter and the second verification parameter of the file are calculated respectively according to the preset verification value calculation method.
[0080] It can be seen that when determining the verification parameters for implementing file grouping, this optional embodiment first performs a consistency judgment on the file modification time in the basic attribute parameters corresponding to the file and the file modification time in the extended attribute parameters corresponding to the file. If the two are inconsistent, it means that the file has been modified and the first verification parameter and the second verification parameter of the file need to be recalculated, which is beneficial to improving the reliability of the first verification parameter and the second verification parameter used for file grouping, and thus helps to improve the accuracy of file grouping.
[0081] Furthermore, for each file in each second file group, determining the first verification parameter and the second verification parameter of the file may further include:
[0082] For each file in each second file group, if the file modification time in the basic attribute parameters corresponding to the file is consistent with the file modification time in the extended attribute parameters corresponding to the file, and the first verification parameter and the second verification parameter exist in the extended attribute parameters of the file, then the first verification parameter and the second verification parameter in the extended attribute parameters of the file are respectively determined as the first verification parameter and the second verification parameter of the file;
[0083] For each file in each second file group, if the file modification time in the basic attribute parameters corresponding to the file is consistent with the file modification time in the extended attribute parameters corresponding to the file, and the first verification parameter and the second verification parameter do not exist in the extended attribute parameters of the file, then the first verification parameter and the second verification parameter of the file are calculated respectively according to the preset verification value calculation method.
[0084] It can be seen that when determining the verification parameters for implementing file grouping, this optional embodiment first performs a consistency judgment on the file modification time in the basic attribute parameters corresponding to the file and the file modification time in the extended attribute parameters corresponding to the file. If the two are consistent, it means that the file has not been modified. At this time, if the first verification parameter and the second verification parameter of the file have been calculated, they can be directly read from the extended attribute parameters corresponding to the file without repeated calculation. If they have not been calculated or the corresponding verification parameters do not exist in the extended attribute parameters, the verification parameters are calculated again, which is beneficial to improving the efficiency of determining the first verification parameter and the second verification parameter of the file while ensuring the reliability of the verification parameters.
[0085] In yet another optional embodiment, the method may further include:
[0086] For any file, after calculating the first verification parameter and the second verification parameter of the file, the calculated first verification parameter and the second verification parameter of the file are updated to the extended attribute parameters of the file. In this way, when the file needs to be deduplicated and screened later, if the file content changes, the verification parameter in the extended attribute parameter can be directly read without further calculation. This is conducive to improving the efficiency of determining the first verification parameter and the second verification parameter of the file while ensuring the reliability of the verification parameter.
[0087] In yet another optional embodiment, for each file in each second file group, respectively calculating the first verification parameter and the second verification parameter of the file according to a preset verification value calculation method may include:
[0088] For each verification parameter to be calculated corresponding to the file, after opening the file, the matching bit data is read from the file in a loop according to the set cache space, and cached in the cache space in a loop. Each time the bit data is cached, the MD5 calculation is performed on the bit data until the MD5 values of all the bit data in the file that match the verification parameter to be calculated are calculated.
[0089] When the check parameter to be calculated is an odd check value, the bit data is the data at the odd position of the file; when the check parameter to be calculated is an even check value, the bit data is the data at the even position of the file.
[0090] It can be seen that when calculating the corresponding check parameters (odd check value or even check value), this optional embodiment cyclically reads the corresponding bit data from the file according to the set cache space and cyclically calculates the corresponding md5 value for the cached bit data. This segmented reading and cyclic cumulative calculation method can reduce the occurrence of situations where the memory cannot be loaded due to the file being too large, which is beneficial to improving the calculation efficiency and accuracy of the check parameters, and thus is beneficial to improving the efficiency and accuracy of file grouping based on the check parameters.
[0091] Example 3
[0092] Please refer to Figure 3, which is a structural diagram of an electronic device disclosed in an embodiment of the present invention. The electronic device described in Figure 3 is used to implement the screening of the file group to be deduplicated or further implement file deduplication after implementing the screening of the file group to be deduplicated. Optionally, the electronic device can be a NAS device or other electronic device that is connected to the NAS device and is used to deduplicate files in the NAS device, which is beneficial to ensure the security of the data in the file during the screening of the file group to be deduplicated or the further implementation of file deduplication after implementing the screening of the file group to be deduplicated. As shown in Figure 3, the electronic device may include:
[0093] A memory 401 storing executable program code;
[0094] a processor 402 coupled to the memory 401;
[0095] The processor 402 calls the executable program code stored in the memory 401 to execute the steps of the file deduplication method described in the first or second embodiment of the present invention.
[0096] Example 4
[0097] An embodiment of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the steps of the file deduplication method described in the first or second embodiment of the present invention.
[0098] Example 5
[0099] Please refer to Figure 4, which is a schematic diagram of the structure of a file deduplication system disclosed in an embodiment of the present invention. As shown in Figure 4, the file deduplication system includes at least the electronic device described in Example 3 and a NAS device in communication with the electronic device, wherein the NAS device stores a plurality of files to be deduplicated. When the electronic device obtains the plurality of files to be deduplicated, the electronic device performs a deduplication operation on the plurality of files to be deduplicated according to the file deduplication method described in Example 1 or Example 2.
[0100] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus the necessary general hardware platform, or of course, by means of hardware. Based on this understanding, the above technical solution, in essence, or the portion that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
Claims
1. A file deduplication method, characterized in that: The method comprises: For a plurality of files to be deduplicated, grouping all the files according to a first attribute parameter of each of the files to obtain a plurality of first file groups; Filtering out at least one second file group satisfying a first preset repetition condition from all the first file groups; For each of the second file groups, grouping all files in the second file group according to the second attribute parameter of each file in the second file group to obtain at least one sub-file group corresponding to the second file group; For each of the second file groups, according to the second preset duplication condition, a sub-file group to be deduplicated is screened out from all sub-file groups corresponding to the second file group, the sub-file group to be deduplicated being a file group on which a file deduplication operation is required; Perform a file deduplication operation on each of the to-be-deduplicated file groups respectively.
2. The file deduplication method according to claim 1, characterized in that: The second attribute parameter of each file in the second file group includes a first verification parameter and a second verification parameter; Wherein, for each of the second file groups, grouping all files in the second file group according to the second attribute parameter of each file in the second file group to obtain at least one sub-file group corresponding to the second file group includes: For each of the second file groups, all files in the second file group are grouped according to a preset verification parameter grouping condition and a first verification parameter and a second verification parameter of each file in the second file group to obtain at least one sub-file group; the preset verification parameter grouping condition is used to indicate that all files corresponding to the same first verification parameter and second verification parameter in the second file group are divided into the same sub-file group.
3. The file deduplication method according to claim 1, characterized in that: For each of the second file groups, according to the second preset duplication condition, a target sub-file group to be deduplicated is selected from all sub-file groups corresponding to the second file group, including: For each of the second file groups, if there is a target sub-file group whose number of files is greater than or equal to a second preset number threshold in the sub-file group corresponding to the second file group, all the target sub-file groups are determined as sub-file groups to be deduplicated.
4. The file deduplication method according to claim 2, characterized in that: The first attribute parameter of each of the files at least includes the file size of each of the files; Among them, for the plurality of files to be deduplicated, all the files are grouped according to the first attribute parameter of each of the files to obtain at least one first file group, including: For a plurality of files to be deduplicated, group all the files according to a preset file size grouping condition and the file size of each of the files to obtain at least one first file group; For any two files that do not belong to the same first file group, the difference in file sizes between the two is greater than a preset difference threshold; when the number of files in a certain first file group is greater than or equal to 2, for any two files in the first file group, the difference in file sizes between the two is less than or equal to the preset difference threshold.
5. The file deduplication method according to claim 4, characterized in that: The step of selecting at least one second file group satisfying a first preset repetition condition from all the first file groups includes: If there is a target file group whose file quantity is greater than or equal to a first preset quantity threshold in all the first file groups, all the target file groups are determined as second file groups that meet a first preset repetition condition.
6. The file deduplication method according to claim 4, characterized in that: The file has corresponding basic attribute parameters and extended attribute parameters; the first attribute parameter of the file is a parameter in the basic attribute parameters, and the second attribute parameter is a parameter in the extended attribute parameters; Furthermore, for each of the second file groups, before grouping all the files in the second file group according to the second attribute parameter of each file in the second file group to obtain at least one sub-file group corresponding to the second file group, the method further includes: For each file in each of the second file groups, respectively determine a first verification parameter and a second verification parameter of the file; The step of determining the first verification parameter and the second verification parameter of each file in each of the second file groups includes: For each file in each of the second file groups, if the file modification time in the basic attribute parameters corresponding to the file is inconsistent with the file modification time in the extended attribute parameters corresponding to the file, the file modification time in the extended attribute parameters corresponding to the file is updated according to the file modification time in the basic attribute parameters corresponding to the file, and after the update, the first verification parameter and the second verification parameter of the file are respectively calculated according to the preset verification value calculation method.
7. The file deduplication method according to claim 6, characterized in that: The step of determining the first verification parameter and the second verification parameter of each file in each of the second file groups respectively includes: For each file in each of the second file groups, if the file modification time in the basic attribute parameters corresponding to the file is consistent with the file modification time in the extended attribute parameters corresponding to the file, and the extended attribute parameters of the file contain the first verification parameter and the second verification parameter, then the first verification parameter and the second verification parameter in the extended attribute parameters of the file are respectively determined as the first verification parameter and the second verification parameter of the file; For each file in each of the second file groups, if the file modification time in the basic attribute parameters corresponding to the file is consistent with the file modification time in the extended attribute parameters corresponding to the file, and the first verification parameter and the second verification parameter do not exist in the extended attribute parameters of the file, then the first verification parameter and the second verification parameter of the file are calculated respectively according to the preset verification value calculation method.
8. The file deduplication method according to claim 6, characterized in that: The method further comprises: Corresponding to any of the files, after the first verification parameter and the second verification parameter of the file are calculated, the calculated first verification parameter and the second verification parameter of the file are updated to the extended attribute parameters of the file.
9. The file deduplication method according to claim 6, characterized in that: For each file in each of the second file groups, respectively calculating the first verification parameter and the second verification parameter of the file according to the preset verification value calculation method includes: For each verification parameter to be calculated corresponding to the file, after opening the file, the matching bit data is read from the file in a loop according to the set cache space, and the bit data is cached in a loop to the cache space, and each time the bit data is cached, the MD5 calculation is performed on the bit data until the MD5 values of all the bit data in the file that match the verification parameter to be calculated are calculated.
10. The file deduplication method according to claim 1, characterized in that: The performing the file deduplication operation on each of the to-be-deduplicated file groups respectively comprises: For each of the sub-file groups to be deduplicated, determine the deduplication strategy corresponding to the file group to be deduplicated based on the number of files in the file group to be deduplicated, the basic attribute parameters of each file in the file group to be deduplicated, and the extended attribute parameters, and perform deduplication operations on the file group to be deduplicated according to the deduplication strategy corresponding to the file group to be deduplicated.
11. The file deduplication method according to claim 10, characterized in that: Before the plurality of files to be deduplicated are grouped according to the first attribute parameter of each file to obtain a plurality of first file groups, the method further includes: For a plurality of files to be deduplicated, a non-critical content filtering operation is performed on the file content of each of the files to update each of the files.
12. An electronic device, characterized in that: The electronic device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the file deduplication method as described in any one of claims 1-11.
13. A file deduplication system, characterized in that: The file deduplication system includes at least the electronic device as described in claim 12, and a NAS device communicatively connected to the electronic device, and a number of files to be deduplicated are stored on the NAS device. When the electronic device obtains the number of files to be deduplicated, the electronic device performs a deduplication operation on the number of files to be deduplicated according to the file deduplication method as described in any one of claims 1-11.
14. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, which, when called, are used to execute the file deduplication method as described in any one of claims 1-11.
Citation Information
Patent Citations
File duplicate checking method and device
CN111309689A
Method for identifying duplicate files and electronic equipment
CN116414782A
File deduplication method, device and system, electronic equipment and computer storage medium
CN117687977A
Systems and methods for assessing upstream oil and gas electronic data duplication
US20180349054A1