Design method for managing duplicated files of disk under Windows

By using SQLite index database and MD5 hashing technology, combining file type and fixed shard content hashing to dynamically monitor USN log update files, the problem of long comparison time and misjudgment in the existing disk duplicate file management methods is solved, and duplicate files are quickly identified and managed, avoiding wasted storage space.

CN120045523APending Publication Date: 2025-05-27TIANJIN AUTOHOME DATA INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510044812.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing disk duplicate file management methods have a long comparison time when the number of files is large, and it is prone to misjudgment of duplicate files, resulting in wasted storage space.

Method used

Using SQLite index database, file information is obtained through MFT and MD5 hash is generated, combining file type, size and fixed shard content hash, dynamically monitor USN log update files, and quickly identify and manage duplicate files.

Benefits of technology

Improves the speed of large file hash generation, quickly query and identify duplicate files, avoids waste of storage space, and reduces server load.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention belongs to the technical field of disk duplicate file management, and discloses a design method for disk duplicate file management under Windows, which comprises the following steps: step 1, a user selects a disk path or a drive needing to be managed; 2, based on the disk path selected in the step 1, checking whether an index database file exists in the current path or not. According to the scheme, by utilizing the step 4, the message transmission quantity of the server and the client is reduced, and meanwhile, the load on the server is reduced, so that the throughput of the server is improved; by means of the step 5, the Hash production speed of the large file is increased; by means of the step 6, the updated file can be rapidly obtained, the index database is updated, meanwhile, the purposes of rapidly querying and recognizing the duplicate file are achieved based on rapid retrieval of the index database, the problems that in an existing scheme, the duplicate file comparison time is long, and duplicate file misjudgment is likely to occur are solved, and the user experience is improved. And the waste of storage space caused by repeated files is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of disk duplicate file management, and in particular relates to a design method for disk duplicate file management under Windows. Background Art

[0002] With the development of technology, the number of files is increasing, which leads to the fact that mobile phones or computers can no longer meet the storage needs of a large number of files. We generally use large-capacity external storage, such as mobile hard disks and other devices to store these files. Therefore, we often face the scenario of transferring files from mobile phones or computers to external storage, and we also face the problem of wasted disk space caused by file duplication.

[0003] There are two existing methods for managing duplicate disk files. The first method uses a disk file traversal method and determines duplicate files by comparing them with existing file information such as file name and file type. This method may encounter the problem of long comparison time when the file volume is large. Secondly, when files with the same name are compared through file name comparison, they may be mistakenly judged as duplicate files, and there is a possibility of accidental deletion. The second method uses a disk file traversal method and compares files by comparing file hashes. File hashing is a time-consuming operation. When the file volume is large or there are large files, it is easy to cause the comparison process to take too long, so it needs to be improved. Summary of the invention

[0004] The purpose of the present invention is to provide a design method for disk duplicate file management under Windows to solve the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solution: a design method for disk duplicate file management under Windows, comprising the following steps:

[0006] Step 1: The user selects the disk path or drive letter to be managed;

[0007] Step 2: Based on the disk path selected in step 1, check whether the index database file already exists in the current path;

[0008] Step 3: If the index database file does not exist, first create the index database file in the selected path. The index database uses SQLite.

[0009] Step 4: Obtain all files in the path through MFT, and index all file information involved into the file index database, including file name, file size, extension, file hash, etc.

[0010] Step 5: The file hash adopts the MD5 algorithm. In order to avoid the problem of long hashing time for large files, the file is hashed based on the file type, file size and the hashing method of the fixed fragment content of the file;

[0011] Step 6: For newly added or updated files in the specified path, dynamically monitor them by means of USN log files, so as to quickly and efficiently re-index the newly added and updated files;

[0012] Step 7: If the user needs to query the duplicate files in the path, all the duplicate files can be filtered out by simply using the data in the file index database;

[0013] Step 8: When the user adds a new file to the path, a hash is automatically generated for the file and then compared with the file index database. If a duplicate file is matched, the user is reminded. If no duplicate file exists, the new file is added to the file index database.

[0014] Preferably, the index database described in step 2 includes three types: clustered index, non-clustered index and combined index, and the combined index consists of multiple fields.

[0015] Preferably, the SQLite described in step three adopts an embedded design, and the SQLite can run on Windows, Linux and Mac OS operating systems.

[0016] Preferably, the MFT (Master File Table) described in step 4 is composed of MFT items, and each MFT item occupies 1024 bytes of space.

[0017] Preferably, the file hash described in step five includes common algorithms such as MD5, SHA-1 and SHA-256, and the MD5 is represented in the form of a 32-bit hexadecimal number.

[0018] Preferably, the file types described in step five include document files, image files, audio files, video files, compressed files, spreadsheet files and slide files, and the file types are divided into different categories according to content format, purpose and related standards.

[0019] Preferably, the USN file described in step six has an automatic storage function, and the information is automatically stored every time the USN file is changed.

[0020] Preferably, the reminder to the user in step eight uses email and communication APP to send notification content, and the notification content can be sent in one of the two ways selected by the notifier.

[0021] Preferably, the file size described in step 4 and step 5 depends on the amount of data contained in the file, and the file size is expressed in bytes.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] This solution utilizes step four to reduce the number of messages transmitted between the server and the client, while reducing the load on the server, thereby improving the server's throughput; utilizes step five to improve the hash production speed of large files; utilizes step six to quickly obtain updated files and update the index database, while achieving the purpose of quickly querying and identifying duplicate files based on the rapid retrieval of the index database, solving the problems of long duplicate file comparison time and easy misjudgment of duplicate files in the existing solution, and avoiding the waste of storage space caused by duplicate files. DETAILED DESCRIPTION

[0024] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in the field without making any creative work shall fall within the scope of protection of the present invention.

[0025] The embodiment of the present invention provides a design method for managing duplicate files on a disk under Windows, comprising the following steps:

[0026] Step 1: The user selects the disk path or drive letter to be managed;

[0027] Step 2: Based on the disk path selected in step 1, check whether the index database file already exists in the current path;

[0028] Step 3: If the index database file does not exist, first create the index database file in the selected path. The index database uses SQLite.

[0029] Step 4: Obtain all files in the path through MFT, and index all file information involved into the file index database, including file name, file size, extension, file hash, etc.

[0030] Step 5: The file hash adopts the MD5 algorithm. In order to avoid the problem of long hashing time for large files, the file is hashed based on the file type, file size and the hashing method of the fixed fragment content of the file;

[0031] Step 6: For newly added or updated files in the specified path, dynamically monitor them by means of USN log files, so as to quickly and efficiently re-index the newly added and updated files;

[0032] Step 7: If the user needs to query the duplicate files in the path, all the duplicate files can be filtered out by simply using the data in the file index database;

[0033] Step 8: When the user adds a new file to the path, a hash is automatically generated for the file and then compared with the file index database. If a duplicate file is matched, the user is reminded. If no duplicate file exists, the new file is added to the file index database.

[0034] This solution can efficiently retrieve duplicate files that already exist in the specified path, and when adding a new file to the path, it can also quickly prompt whether the file already exists in the path, avoiding the waste of storage space caused by duplicate files, while freeing up disk storage space and reducing the burden on computer operation.

[0035] The index database in step 2 includes three types: clustered index, non-clustered index and joint index, and the joint index is composed of multiple fields.

[0036] By using the index database, data that meets the query conditions can be quickly located, thereby improving the efficiency of the query.

[0037] Among them, the SQLite in step three adopts an embedded design, and SQLite can run on Windows, Linux and Mac OS operating systems.

[0038] SQLite is designed to process indexed databases more quickly.

[0039] Among them, the MFT (Master File Table) in step 4 is composed of MFT items, and each MFT occupies 1024 bytes of space.

[0040] MFT is an important data structure used to store file and directory information in the NTFS file system.

[0041] Among them, the file hash in step five includes common algorithms such as MD5, SHA-1 and SHA-256, and MD5 is represented in the form of 32-bit hexadecimal numbers.

[0042] By comparing hash values, files with the same content can be quickly found.

[0043] Among them, the file types in step five include document files, image files, audio files, video files, compressed files, spreadsheet files and slide files. The file types are divided into different categories according to content format, purpose and related standards.

[0044] The Windows operating system and applications identify file types by file extensions and open and process files accordingly.

[0045] Among them, the USN file in step six has an automatic storage function, and the information is automatically stored every time the USN file is changed.

[0046] This design avoids the loss of USN files and the impact on the file index database.

[0047] Among them, step eight reminds users to use email and communication APP to send notification content, and the notification content can be sent in a way that the notifier can choose one of them.

[0048] This design allows the notifier to choose the method of sending the notification content according to his or her own needs, thereby making it convenient for the notifier to receive and be prompted.

[0049] The file size of step 4 and step 5 depends on the amount of data contained in the file, and the file size is expressed in bytes.

[0050] Bytes play a key role in computer storage and data transmission and are used to measure file size, memory capacity, network bandwidth, etc.

[0051] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. Although embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations may be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A design method for managing duplicate files on a disk under Windows, characterized in that: The following steps are involved: Step 1: The user selects the disk path or drive letter to be managed; Step 2: Based on the disk path selected in step 1, check whether the index database file already exists in the current path; Step 3: If the index database file does not exist, first create the index database file in the selected path. The index database uses SQLite. Step 4: Obtain all files in the path through MFT, and index all file information involved into the file index database, including file name, file size, extension, file hash, etc. Step 5: The file hash adopts the MD5 algorithm. In order to avoid the problem of long hashing time for large files, the file is hashed based on the file type, file size and the hashing method of the fixed fragment content of the file; Step 6: For newly added or updated files in the specified path, dynamically monitor them by means of USN log files, so as to quickly and efficiently re-index the newly added and updated files; Step 7: If the user needs to query the duplicate files in the path, all the duplicate files can be filtered out by simply using the data in the file index database; Step 8: When the user adds a new file to the path, a hash is automatically generated for the file and then compared with the file index database. If a duplicate file is matched, the user is reminded. If no duplicate file exists, the new file is added to the file index database.

2. A design method for managing duplicate files on a disk under Windows according to claim 1, characterized in that: The index database described in step 2 includes three types: clustered index, non-clustered index and combined index, and the combined index is composed of multiple fields.

3. The design method for managing duplicate files on a disk under Windows according to claim 1, characterized in that: The SQLite described in step 3 adopts an embedded design, and the SQLite can run on Windows, Linux and Mac OS operating systems.

4. The design method for managing duplicate files on a disk under Windows according to claim 1, characterized in that: The MFT (Master File Table) described in step 4 is composed of MFT items, and each MFT item occupies 1024 bytes of space.

5. The design method for managing duplicate files on a disk under Windows according to claim 1, characterized in that: The file hash described in step five includes common algorithms such as MD5, SHA-1 and SHA-256, and the MD5 is represented in the form of a 32-bit hexadecimal number.

6. The design method for managing duplicate files on a Windows disk according to claim 1, characterized in that: The file types described in step five include document files, image files, audio files, video files, compressed files, spreadsheet files and slide files. The file types are divided into different categories according to content format, purpose and related standards.

7. The design method for managing duplicate files on a disk under Windows according to claim 1, characterized in that: The USN file described in step 6 has an automatic storage function, and the information is automatically stored every time the USN file is changed.

8. The design method for managing duplicate files on a disk under Windows according to claim 1, characterized in that: The reminder user described in step eight uses email and communication APP to send notification content, and the notification content can be sent in one of the ways selected by the notifier.

9. The design method for managing duplicate files on a disk under Windows according to claim 1, characterized in that: The file size described in step 4 and step 5 depends on the amount of data contained in the file, and the file size is expressed in bytes.