File processing method and device

Through global hash index directory and hard link technology, the complexity of file comparison and deduplication in the existing technology is solved, and efficient utilization of disk space is achieved.

CN120255819AActive Publication Date: 2025-07-04BEIJING HAIYUDONGXIANG TECH CO LTD

Patent Information

Application Number
CN202510668508.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-04
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The prior art relies on databases when comparing files and deduplication, resulting in complex and not simple enough, making it difficult to effectively save disk space.

Method used

The global hash index directory is used instead of the database, and the same file exists by calculating the file hash value, and the hard links are used to share disk data to achieve optimization during file deduplication and upload.

Benefits of technology

No database is required, and file deduplication and upload optimization is achieved through the global hash index directory, which simplifies operations and effectively saves disk space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255819A_ABST
    Figure CN120255819A_ABST
Patent Text Reader

Abstract

The invention provides a file processing method and device, and the method comprises the steps: calculating a Hash value corresponding to a first file, judging whether a second file exists in a pre-created first global Hash index directory or not according to the Hash value corresponding to the first file, and if the second file exists in the first global Hash index directory, determining that the second file exists in the first global Hash index directory; if yes, replacing the first file with the hard link of the second file; and / or obtaining a hash value corresponding to the third file, judging whether a fourth file exists in a pre-created second global hash index directory or not according to the hash value corresponding to the third file, if the fourth file exists in the second global hash index directory, creating the third file, and performing hard linking on the third file to the fourth file, so that the disk space can be saved; and the realization is simpler.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and particularly relates to a method and apparatus for file processing. Background Art

[0002] In order to save disk space, there are many scenarios where file comparison / deduplication is required. For example, for different versions of game data files stored in a cloud storage server (the game data files are shared by all users), deduplication is performed so that only one copy of the content of the game data file is retained in the cloud storage server. Another example is that before uploading a certain version of the game data file (the game data file is shared by all users) to the cloud storage server, it is detected whether there is a file with the same content stored in the cloud storage server. If there is a file with the same content in the cloud storage server, the game data file is not uploaded to the cloud storage server, otherwise, the game data file is uploaded to the cloud storage server so that only one copy of the content of the game data file is retained in the cloud storage server, etc. Existing solutions for these scenarios often use a database to store the md5 values of files, and then perform file comparison based on the md5 values stored in the database. According to the comparison result, if the md5 values of different files are the same, hard links are used to make different files share the same disk data. Since the existing solutions rely on a database, the implementation is relatively complex, and when a file needs to be deleted, the database needs to be notified to delete the corresponding data record, which makes the overall implementation of the existing solutions relatively complex.

[0003] Therefore, how to provide an optimized file processing solution that can save disk space and is simpler to implement has become a technical problem to be solved urgently. Summary of the Invention

[0004] In view of the technical problems existing in the prior art, embodiments of the present application provide a method and apparatus for file processing.

[0005] In a first aspect, embodiments of the present application provide a method for file processing, including: Calculating a hash value corresponding to a first file, and determining whether a second file exists in a pre-created first global hash index directory according to the hash value corresponding to the first file. If the second file exists in the first global hash index directory, replacing the first file with a hard link to the second file, where the relative path of the second file with respect to the first global hash index directory contains the hash value corresponding to the first file; and / or Obtain the hash value corresponding to the third file, and determine whether there is a fourth file in the pre-created second global hash index directory according to the hash value corresponding to the third file. If there is a fourth file in the second global hash index directory, create the third file and hard link the third file to the fourth file, where the relative path of the fourth file with respect to the second global hash index directory contains the hash value corresponding to the third file.

[0006] In a second aspect, an embodiment of the present application further provides a file processing device, including: A processing unit, configured to calculate the hash value corresponding to the first file, and determine whether there is a second file in the pre-created first global hash index directory according to the hash value corresponding to the first file. If there is a second file in the first global hash index directory, replace the first file with a hard link of the second file, where the relative path of the second file with respect to the first global hash index directory contains the hash value corresponding to the first file; and / or obtain the hash value corresponding to the third file, and determine whether there is a fourth file in the pre-created second global hash index directory according to the hash value corresponding to the third file. If there is a fourth file in the second global hash index directory, create the third file and hard link the third file to the fourth file, where the relative path of the fourth file with respect to the second global hash index directory contains the hash value corresponding to the third file.

[0007] In summary, the file processing method and device provided by the embodiments of the present application do not use a database, but rely on pre-created global hash index directories (i.e., the first global hash index directory and / or the second global hash index directory): when performing file deduplication, determine whether there is a second file in the first global hash index directory. If there is a second file in the first global hash index directory, it means that there is a file with the same content as the first file, then replace the first file with a hard link of the second file, so that the first file and the second file share the same disk data; when uploading a new file (i.e., the third file), determine whether there is a fourth file in the second global hash index directory. If there is a fourth file in the second global hash index directory, it means that there is a file with the same content as the third file, then do not upload the third file, but create the third file and hard link the third file to the fourth file, so that the third file and the fourth file share the same disk data. In this way, disk space can be saved, and since no database is used, the implementation is simpler compared to the existing solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 It is a schematic flowchart of an embodiment of a file processing method provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of an embodiment of a file processing device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0009] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. It should be understood that the accompanying drawings in this application are only for the purposes of illustration and description, and are not used to limit the protection scope of this application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application.

[0010] In addition, the described embodiments are only some embodiments of this application, rather than all embodiments. The components of the embodiments of this application usually described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of this application claimed, but only represents selected embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of this application.

[0011] It should be noted that the term "including" will be used in the embodiments of this application to indicate the existence of the subsequently stated features, but does not exclude the addition of other features.

[0012] Refer to Figure 1 As shown, the embodiments of this application provide a schematic flowchart of a method for file processing. The method includes: S10. Calculate the hash value corresponding to the first file, and determine whether there is a second file in the pre-created first global hash index directory according to the hash value corresponding to the first file. If there is a second file in the first global hash index directory, replace the first file with a hard link to the second file, where the relative path of the second file relative to the first global hash index directory contains the hash value corresponding to the first file; and / or S11. Obtain the hash value corresponding to the third file, and determine whether there is a fourth file in the pre-created second global hash index directory according to the hash value corresponding to the third file. If there is a fourth file in the second global hash index directory, create the third file and hard link the third file to the fourth file, where the relative path of the fourth file relative to the second global hash index directory contains the hash value corresponding to the third file.

[0013] In this embodiment, it should be noted that step S10 is applicable to the scenario of file deduplication. For example, if it is necessary to deduplicate files in at least one directory (for example, the at least one directory is a file directory shared by all users in a cloud storage server, and files such as game resources and SO libraries shared by all users are stored under this at least one directory), then step S10 can be executed for each file in the at least one directory. In step S10, if a second file exists in the first global hash index directory, the first file is replaced with a hard link to the second file; otherwise, if the second file does not exist in the first global hash index directory, a second file is created in the first global hash index directory, and the second file is hard-linked to the first file. The relative path of the second file with respect to the first global hash index directory contains the hash value corresponding to the first file. For example, assume that it is necessary to deduplicate files in the directory / game_resources. There are 5 files in the directory / game_resources, namely / game_resources / 1.txt, / game_resources / 2.txt, / game_resources / 3.txt, / game_resources / 4.txt, and / game_resources / 5.txt. Then, a first global hash index directory / global_hash_index_dir can be created. Assume that the hash values of 1.txt, 2.txt, 3.txt, 4.txt, and 5.txt are AAAAAAAA, BBBBBBBB, CCCCCC, AAAAAAAA, and BBBBBBBB respectively. For 1.txt, since there is no file named AAAAAAAA in / global_hash_index_dir, it is necessary to create the file / global_hash_index_dir / AAAAAAAA and hard-link / global_hash_index_dir / AAAAAAAA to / game_resources / 1.txt. At this time, / game_resources / 1.txt and / global_hash_index_dir / AAAAAAAA share the same disk data (i.e., the disk data of / game_resources / 1.txt).Similarly, for 2.txt, the file / global_hash_index_dir / BBBBBBBB needs to be created, and a hard link of / global_hash_index_dir / BBBBBBBB to / game_resources / 2.txt is required. At this time, / game_resources / 2.txt and / global_hash_index_dir / BBBBBBBB share the same disk data; for 3.txt, the file / global_hash_index_dir / CCCCCCCC needs to be created, and a hard link of / global_hash_index_dir / CCCCCCCC to / game_resources / 3.txt is required. At this time, / game_resources / 3.txt and / global_hash_index_dir / CCCCCCCC share the same disk data. For 4.txt, since the file / global_hash_index_dir / AAAAAAAA exists, the hard link of / game_resources / 4.txt needs to be replaced with / global_hash_index_dir / AAAAAAAA. At this time, / game_resources / 1.txt, / global_hash_index_dir / AAAAAAAA, and / game_resources / 4.txt share the same disk data. For 5.txt, since the file / global_hash_index_dir / BBBBBBBB exists, the hard link of / game_resources / 5.txt needs to be replaced with / global_hash_index_dir / BBBBBBBB. At this time, / game_resources / 2.txt, / global_hash_index_dir / BBBBBBBB, and / game_resources / 5.txt share the same disk data. In this example, the hash value corresponding to the first file is the hash value of the first file. For example, the relative path of the second file / global_hash_index_dir / AAAAAAAA to the first global hash index directory / global_hash_index_dir is AAAAAAAA (i.e., the hash value of the first file / game_resources / 1.txt). The hash value corresponding to the first file may include the hash value of all the contents of the first file, or the hash value of the first data, or the hash value of a part of the contents of the first file, or the hash value of the second data. Among them, the first data may include the hash value of all the contents of the first file, and the second data includes the hash value of a part of the contents of the first file.In addition to the hash value of all the contents of the first file, the hash value corresponding to the first file may also include the hash value of at least one basic attribute and / or extended attribute of the first file. The basic attributes of the first file include the access permission mode of the first file, the owner user of the first file, the user group group to which the first file belongs, and the last modification time of the first file, etc. For example, the hash value of all the contents of the first file 6.txt is DDDDDDDD, the access permission mode of 6.txt is 123, the owner user of 6.txt is 10000, the user group group to which 6.txt belongs is 10000, and the first global hash index directory is / global_hash_index_dir. Then the path of the second file corresponding to 6.txt can be / global_hash_index_dir / DDDDDDDD or / global_hash_index_dir / DDDDDDDD_123_10000_10000 (the symbol "_" in / global_hash_index_dir / DDDDDDDD_123_10000_10000 can be replaced by other delimiters) or / global_hash_index_dir / DDDDDDDD / 123_10000_10000, etc. In addition to the hash value of all the contents of the first file, the first data may also include the hash value of at least one basic attribute and / or extended attribute of the first file. For example, assuming that the hash value of all the extended attributes of the first file 6.txt is EEEEE, the first data corresponding to the first file 6.txt may include DDDDDDDD, 123, 10000, 10000, and EEEEE. The path of the second file corresponding to 6.txt can be / global_hash_index_dir / FFFFFFFF, where FFFFFFFF is the hash value of the first data.The hash value corresponding to the first file may include the hash values of partial contents of the first file. For example, at least one part of the content is extracted from the first file, and the hash value of each extracted part is calculated separately to obtain at least one hash value. Then, the hash value of the partial contents of the first file includes the at least one hash value. In addition, the hash value corresponding to the first file may also include the hash value of all the contents of the first file and / or the hash value of at least one basic attribute of the first file and / or the hash value of the extended attribute of the first file. For example, the first file 6.txt is divided into two parts according to the size of 1M. The hash value of the first part is MMMMM, and the hash value of the second part is NNNNN. Then, the path of the second file corresponding to 6.txt can be / global_hash_index_dir / MMMMM / NNNNN or / global_hash_index_dir / MMMMM / NNNNN_123_10000_10000 or / global_hash_index_dir / MMMMM / NNNNN / 123_10000_10000, etc. In addition to including the hash value of the partial contents of the first file, the second data may also include the hash value of all the contents of the first file and / or the hash value of at least one basic attribute of the first file and / or the hash value of the extended attribute of the first file. For example, the second data corresponding to 6.txt may include MMMMM, NNNNN, and EEEEE. The path of the second file corresponding to 6.txt can be / global_hash_index_dir / GGGGGGGG, where GGGGGGGG is the hash value of the second data. By calculating the hash value of the partial contents of the first file and using the hash value of the partial contents of the first file in the judgment logic, the operation efficiency can be improved.

[0014] Step S11 is for the scenario of uploading a file. For example, it is necessary to upload a file in one device to another device (such as uploading a new version of the game data file to the cloud storage server, where the cloud storage server stores the old version of the game data file). In this scenario, step S11 can be applied to the other device. Obtaining the hash value corresponding to the third file in step S11 may include receiving the hash value corresponding to the third file uploaded by other devices, or receiving the content of the third file uploaded by other devices and calculating the hash value corresponding to the third file based on the content of the third file. The hash value corresponding to the third file may include the hash value of all the content of the third file, or the hash value of the third data, or the hash value of a part of the content of the third file, or the hash value of the fourth data, where the third data includes the hash value of all the content of the third file, and the fourth data includes the hash value of a part of the content of the third file. The judgment logic in step S11 is the same as that in step S10, which will not be elaborated here again. If the fourth file exists in the second global hash index directory, it is necessary to determine the absolute path of the third file according to the relative path of the third file uploaded by other devices (this is prior art and will not be elaborated here), create the third file according to the absolute path of the third file, and hard link the third file to the fourth file so that the third file and the fourth file share the same disk data (i.e., the disk data of the fourth file). The second global hash index directory is the same as or different from the first global hash index directory.For example, assume that it is necessary to upload the third file 7.txt to the device applied in step S11. The hash value, access permission mode, owner user, and group to which 7.txt belongs are BBBBBBBB, 336, 11000, and 11000 respectively. There are 4 files in the second global hash index directory / global_hash_index_dir, namely / global_hash_index_dir / AAAAAAAA_220_10000_10000, / global_hash_index_dir / BBBBBBBB_336_11000_11000, / global_hash_index_dir / CCCCCCCC_440_12000_12000, and / global_hash_index_dir / BBBBBBBB_461_13000_13000 (that is, the form of the file path in the second global hash index directory is / global_hash_index_dir / hash value_access permission mode_owner user_group to which it belongs). Then, since there is a fourth file / global_hash_index_dir / BBBBBBBB_336_11000_11000 corresponding to 7.txt, 7.txt is no longer uploaded to the device applied in step S11. Instead, 7.txt is created and hard-linked to / global_hash_index_dir / BBBBBBBB_336_11000_11000 so that 7.txt and / global_hash_index_dir / BBBBBBBB_336_11000_11000 share the same disk data. If the fourth file does not exist in the second global hash index directory, the fourth file is received and stored, the fourth file is created in the second global hash index directory, and the fourth file is hard-linked to the third file.For example, assume that it is necessary to upload the third file 8.txt to the device used in step S11. The hash value, access permission mode, owner user, and user group to which 8.txt belongs are BBBBBBBB, 425, 14000, and 14000 respectively. There are 4 files in total under the second global hash index directory / global_hash_index_dir, namely / global_hash_index_dir / AAAAAAAA_256_10000_10000, / global_hash_index_dir / BBBBBBBB_312_11000_11000, / global_hash_index_dir / CCCCCCCC_370_12000_12000, and / global_hash_index_dir / BBBBBBBB_425_13000_13000. Since there is no file named BBBBBBBB_425_14000_14000 under / global_hash_index_dir, it is necessary to upload 8.txt to the device used in step S11, store 8.txt, create / global_hash_index_dir / BBBBBBBB_425_14000_14000, and hard link / global_hash_index_dir / BBBBBBBB_425_14000_14000 to 8.txt, so that / global_hash_index_dir / BBBBBBBB_425_14000_14000 and 8.txt share the same disk data.

[0015] The file processing method provided by the embodiments of the present application does not use a database, but relies on a pre-created global hash index directory (i.e., the first global hash index directory and / or the second global hash index directory): when performing file deduplication, it is judged whether a second file exists in the first global hash index directory. If a second file exists in the first global hash index directory, it means that there is a file with the same content as the first file, then the first file is replaced with a hard link to the second file, so that the first file and the second file share the same disk data; when uploading a new file (i.e., the third file), it is judged whether a fourth file exists in the second global hash index directory. If a fourth file exists in the second global hash index directory, it means that there is a file with the same content as the third file, then the third file is not uploaded, but the third file is created and hard linked to the fourth file, so that the third file and the fourth file share the same disk data. In this way, disk space can be saved, and since no database is used, the implementation is simpler compared with the existing solutions.

[0016] Based on the foregoing method embodiments, the method may further include: Deleting a file whose nlink attribute in the inode pointed to by the dentry in the first global hash index directory and / or the second global hash index directory is 1; or Deleting a file whose nlink attribute in the inode pointed to by the dentry in the first global hash index directory and / or the second global hash index directory is 1 and the time interval between the ctime attribute and the current time is greater than a preset threshold.

[0017] In this embodiment, it should be noted that for a file in the first global hash index directory, if the nlink attribute in the inode pointed to by the dentry of the file is 1, it means that except for this file, there is no other file's dentry pointing to this inode. At this time, this file needs to be deleted. In order to avoid the content of this file being used shortly after the file is deleted, an additional condition can be added on the basis that the nlink attribute in the inode pointed to by the dentry is 1: the time interval between the ctime attribute in the inode pointed to by the dentry and the current time is greater than a preset threshold (such as 24 hours). If there is a file in the first global hash index directory that meets these two conditions, the file in the first global hash index directory that meets these two conditions can be deleted. The processing logic for deleting files in the second global hash index directory is the same as that for deleting files in the first global hash index directory, and will not be elaborated here.

[0018] Referring to Figure 2 As shown, it is a schematic structural diagram of a file processing device provided by an embodiment of the present application. The device includes: A processing unit 20, configured to calculate a hash value corresponding to a first file, and determine whether a second file exists in a pre-created first global hash index directory according to the hash value corresponding to the first file. If the second file exists in the first global hash index directory, replace the first file with a hard link to the second file, where the relative path of the second file with respect to the first global hash index directory includes the hash value corresponding to the first file; and / or obtain a hash value corresponding to a third file, and determine whether a fourth file exists in a pre-created second global hash index directory according to the hash value corresponding to the third file. If the fourth file exists in the second global hash index directory, create the third file and hard link the third file to the fourth file, where the relative path of the fourth file with respect to the second global hash index directory includes the hash value corresponding to the third file.

[0019] The file processing device provided by the embodiment of the present application does not use a database, but relies on a pre-created global hash index directory (i.e., the first global hash index directory and / or the second global hash index directory): when performing file deduplication, it is determined whether a second file exists in the first global hash index directory. If the second file exists in the first global hash index directory, it indicates that there is a file with the same content as the first file, then the first file is replaced with a hard link to the second file, so that the first file and the second file share the same disk data; when uploading a new file (i.e., the third file), it is determined whether a fourth file exists in the second global hash index directory. If the fourth file exists in the second global hash index directory, it indicates that there is a file with the same content as the third file, then the third file is not uploaded, but the third file is created and the third file is hard-linked to the fourth file, so that the third file and the fourth file share the same disk data. In this way, disk space can be saved, and since no database is used, the implementation is simpler compared to the existing solution.

[0020] Based on the foregoing device embodiment, the hash value corresponding to the first file may include the hash value of all the contents of the first file or the hash value of the first data or the hash value of a part of the contents of the first file or the hash value of the second data, where the first data includes the hash value of all the contents of the first file, and the second data includes the hash value of a part of the contents of the first file; The hash value corresponding to the third file may include the hash value of all the contents of the third file or the hash value of the third data or the hash value of a part of the contents of the third file or the hash value of the fourth data, where the third data includes the hash value of all the contents of the third file, and the fourth data includes the hash value of a part of the contents of the third file.

[0021] Based on the foregoing device embodiment, the processing unit may further be configured to: If the second file does not exist in the first global hash index directory, create the second file in the first global hash index directory and hard-link the second file to the first file.

[0022] Based on the foregoing device embodiment, the processing unit may further be configured to: If the fourth file does not exist in the second global hash index directory, receive the fourth file, store the fourth file, create the fourth file in the second global hash index directory, and hard-link the fourth file to the third file.

[0023] Based on the foregoing device embodiment, the device may further include: A deletion unit is configured to delete a file whose nlink attribute in the inode pointed to by the dentry in the first global hash index directory and / or the second global hash index directory is 1; or delete a file whose nlink attribute in the inode pointed to by the dentry in the first global hash index directory and / or the second global hash index directory is 1 and the time interval between the ctime attribute and the current time is greater than a preset threshold.

[0024] The device for file processing provided by the embodiments of the present application has the same implementation process as the method for file processing provided by the embodiments of the present application, and can achieve the same effects as the method for file processing provided by the embodiments of the present application, and will not be elaborated herein.

[0025] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for file processing, characterized in that, including: calculating a hash value corresponding to a first file, determining whether a second file exists in a pre-created first global hash index directory according to the hash value corresponding to the first file, and if the second file exists in the first global hash index directory, replacing the first file with a hard link to the second file, where a relative path of the second file with respect to the first global hash index directory includes the hash value corresponding to the first file; and / or obtaining a hash value corresponding to a third file, determining whether a fourth file exists in a pre-created second global hash index directory according to the hash value corresponding to the third file, and if the fourth file exists in the second global hash index directory, creating the third file and hard-linking the third file to the fourth file, where a relative path of the fourth file with respect to the second global hash index directory includes the hash value corresponding to the third file.

2. The method according to claim 1, characterized in that, The hash value corresponding to the first file includes a hash value of all contents of the first file or a hash value of first data or a hash value of partial contents of the first file or a hash value of second data, where the first data includes a hash value of all contents of the first file and the second data includes a hash value of partial contents of the first file; The hash value corresponding to the third file includes a hash value of all contents of the third file or a hash value of third data or a hash value of partial contents of the third file or a hash value of fourth data, where the third data includes a hash value of all contents of the third file and the fourth data includes a hash value of partial contents of the third file.

3. The method according to claim 1, characterized in that, further including: if the second file does not exist in the first global hash index directory, creating the second file in the first global hash index directory and hard-linking the second file to the first file.

4. The method according to claim 1, characterized in that, further including: if the fourth file does not exist in the second global hash index directory, receiving the fourth file, storing the fourth file, creating the fourth file in the second global hash index directory, and hard-linking the fourth file to the third file.

5. The method according to any one of claims 1 to 4, characterized in that further including: deleting a file in which an nlink attribute of an inode pointed to by a dentry in a directory entry in the first global hash index directory and / or the second global hash index directory is 1; or deleting a file in which an nlink attribute of an inode pointed to by a dentry in a directory entry in the first global hash index directory and / or the second global hash index directory is 1 and a time interval between a ctime attribute and the current time is greater than a preset threshold.

6. A device for file processing, characterized in that, including: A processing unit is configured to calculate the hash value corresponding to the first file, and determine whether a second file exists in a pre-created first global hash index directory according to the hash value corresponding to the first file. If the second file exists in the first global hash index directory, replace the first file with a hard link to the second file, where the relative path of the second file with respect to the first global hash index directory contains the hash value corresponding to the first file; and / or obtain the hash value corresponding to the third file, and determine whether a fourth file exists in a pre-created second global hash index directory according to the hash value corresponding to the third file. If the fourth file exists in the second global hash index directory, create the third file and hard link the third file to the fourth file, where the relative path of the fourth file with respect to the second global hash index directory contains the hash value corresponding to the third file.

7. The device according to claim 6, characterized in that, The hash value corresponding to the first file includes the hash value of all contents of the first file, or the hash value of the first data, or the hash value of part of the contents of the first file, or the hash value of the second data, where the first data includes the hash value of all contents of the first file, and the second data includes the hash value of part of the contents of the first file; The hash value corresponding to the third file includes the hash value of all contents of the third file, or the hash value of the third data, or the hash value of part of the contents of the third file, or the hash value of the fourth data, where the third data includes the hash value of all contents of the third file, and the fourth data includes the hash value of part of the contents of the third file.

8. The device according to claim 6, characterized in that, The processing unit is further configured to: If the second file does not exist in the first global hash index directory, create the second file in the first global hash index directory and hard link the second file to the first file.

9. The device according to claim 6, characterized in that, The processing unit is further configured to: If the fourth file does not exist in the second global hash index directory, receive the fourth file and store it, create the fourth file in the second global hash index directory, and hard link the fourth file to the third file.

10. The device according to any one of claims 6 to 9, characterized in that, It further includes: A deletion unit is configured to delete a file whose nlink attribute is 1 in the inode pointed to by the dentry in the first global hash index directory and / or the second global hash index directory; or delete a file whose nlink attribute is 1 and the time interval between the ctime attribute and the current time is greater than a preset threshold in the inode pointed to by the dentry in the first global hash index directory and / or the second global hash index directory.

Citation Information

Patent Citations

  • File system duplicate removal method and device based on cloud storage

    CN103136243A

  • Log-structured storage systems

    CN111886582A

  • File redundancy removal method for terminal equipment, terminal equipment and storage medium

    CN114860677A

  • File storage and file reading method, device and system

    CN116010362A

  • Container resource occupation overhead optimization method and system, medium and computer equipment

    CN118708290A

Cited By

  • Remote file fast synchronization method and device

    CN120602476A