Recognition method and device of duplicate file, storage medium and computer equipment
By quickly positioning the different locations of files and extracting type characteristics through the prior knowledge base, the problem of high calculation cost and low efficiency when identifying and removing duplicate files in the prior art is solved, and the rapid and accurate file comparison and deduplication effect is achieved.
Patent Information
- Application Number
- CN202411915103.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is highly cost-effective in identifying and removing duplicate files, especially when processing large data sets or media files.
Quickly locate possible different locations in files through the prior knowledge base, reduce unnecessary full file scanning, and reduce computing resource consumption; extract type characteristics at different locations, and analyze only different locations, greatly improving the speed of file comparison.
It realizes the rapid and accurate judgment of whether a file is a duplicate file, reduces the consumption of computing resources, improves the speed of file comparison, and is suitable for multiple file types and has strong versatility.
Smart Images

Figure CN119938622A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and device for identifying duplicate files, a storage medium, and a computer device. Background Art
[0002] In the field of information technology, with the rapid growth of data volume, file deduplication has become a key task in data management. Duplicate files not only take up a lot of storage space and increase storage costs, but also bring inconvenience to data management and maintenance, especially when dealing with large data sets or scenarios containing a large number of media files such as audio and video. This problem is particularly prominent.
[0003] Traditional file deduplication methods mainly rely on comparing file hash values or comparing file contents byte by byte. However, these traditional methods have exposed many defects and shortcomings in practical applications. First, although byte-by-byte comparison can ensure the accuracy of file deduplication, its computational cost is high, especially when dealing with large files or a large number of files. It is inefficient and difficult to meet the speed and efficiency requirements of modern data management. Secondly, although calculating hash values has improved the efficiency of deduplication to a certain extent, for large files, the calculation of hash values is still a time-consuming and resource-intensive process, and although the possibility of hash conflicts is extremely low, its potential impact still needs to be considered. Summary of the invention
[0004] In view of this, the present application provides a duplicate file identification method and device, storage medium, and computer equipment, which can quickly locate the locations where there may be differences in the file through a priori knowledge base, reduce unnecessary full file scanning, and reduce the consumption of computing resources; by extracting the type features at the difference locations, only the difference locations are analyzed, greatly improving the speed of file comparison. In addition, the difference locations corresponding to different file types are stored in the priori knowledge base, making the method applicable to a variety of file types and highly versatile.
[0005] According to one aspect of the present application, a method for identifying duplicate files is provided, comprising:
[0006] Acquire two target files to be identified for repetitiveness, and respectively determine the storage space occupied by each target file, wherein the two target files have the same file type;
[0007] In the case where the storage spaces occupied by the two target files are equal, based on the file types of the two target files, querying the difference positions corresponding to the file types from a priori knowledge base, and extracting the type features corresponding to each target file at the difference positions respectively;
[0008] Based on the type feature, it is determined whether the two target files are duplicate files.
[0009] According to another aspect of the present application, a duplicate file identification device is provided, comprising:
[0010] A storage space determination module, used to obtain two target files to be identified for duplication, and respectively determine the storage space occupied by each target file, wherein the two target files are of the same file type;
[0011] A difference position determination module, used for, when the storage spaces occupied by the two target files are equal, querying the difference positions corresponding to the file types from a priori knowledge base based on the file types of the two target files, and respectively extracting the type features corresponding to each target file at the difference positions;
[0012] The duplicate file determination module is used to determine whether the two target files are duplicate files based on the type feature.
[0013] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned duplicate file identification method is implemented.
[0014] According to another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the above-mentioned duplicate file identification method when executing the program.
[0015] By means of the above technical scheme, the present application provides a method and device for identifying duplicate files, a storage medium, and a computer device. First, two target files to be identified for duplication are obtained, wherein the two target files to be identified for duplication are files of the same file type. Then, the storage space occupied by each of the two target files is determined. If the storage space occupied by the two target files is equal, then further, the difference position matching the file type of the two target files can be found from the prior knowledge base. Subsequently, according to the difference position provided by the prior knowledge base, the type features at the difference position are extracted from the two target files respectively. Further, the type features of the two target files are compared. If there are differences in the type features of the two target files at the difference position, it can be directly stated that the two target files are not duplicate files. The embodiment of the present application quickly locates the position where there may be differences in the file through the prior knowledge base, reduces unnecessary full file scanning, and reduces the consumption of computing resources; by extracting the type features at the difference position, only the difference position is analyzed, which greatly improves the speed of file comparison. In addition, the prior knowledge base stores the difference positions corresponding to different file types, so that the method is applicable to a variety of file types and has strong versatility.
[0016] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 A schematic diagram of a flow chart of a method for identifying duplicate files provided in an embodiment of the present application is shown;
[0019] Figure 2 A schematic diagram of the structure of a duplicate file identification device provided in an embodiment of the present application is shown;
[0020] Figure 3 A schematic diagram of the device structure of a computer device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0021] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict.
[0022] In this embodiment, a method for identifying duplicate files is provided. Figure 1 As shown, the method includes:
[0023] Step 101, obtaining two target files to be identified for duplication, and determining the storage space occupied by each target file, wherein the two target files are of the same file type.
[0024] Step 102, when the storage space occupied by the two target files is equal, based on the file types of the two target files, query the difference positions corresponding to the file types from the prior knowledge base, and extract the type features corresponding to each target file at the difference positions respectively.
[0025] Step 103: Based on the type feature, determine whether the two target files are duplicate files.
[0026] A method for identifying duplicate files provided by an embodiment of the present application can quickly and accurately determine whether two files are duplicate files. When performing duplicate identification on files, first, two target files to be identified for repeatability are obtained, and these two target files can be candidate files that the user or the system believes may have repeatability. It should be noted that in order to ensure the effectiveness of duplicate file identification, the two target files to be identified for repeatability are files of the same file type. For example, assuming that the two target files to be identified for repeatability are file 1 and file 2, respectively, if file 1 is a text file, then file 2 should also be a text file. Here, file types can include text files, picture files, audio files, video files, etc. Next, determine the size of the storage space occupied by each of the two target files. If the sizes are different, they are definitely not the same file. Only when the storage space of the two target files is equal, further repeatability checks are performed.
[0027] If the storage space occupied by the two target files is equal, then further, the difference positions that match the file types of the two target files can be found from the prior knowledge base. Among them, the pre-established prior knowledge base contains information about the positions (or areas) where the content differences are most likely to occur in files of different formats under each file type. This is based on statistical laws derived from a large number of file analyses, and aims to improve recognition efficiency. Here, the difference position can be one or more. Subsequently, based on the difference position provided by the prior knowledge base, the type features at the difference positions are extracted from the two target files respectively. These type features can be the hash value of a part of the target file, a specific byte sequence, or other indicators that can reflect the uniqueness of the target file content.
[0028] After extracting the type features, the type features of the two target files are further compared. If the type features of the two target files at the difference position are exactly the same, it cannot be said that the two target files are non-duplicate files, and further judgment is required; if the type features of the two target files at the difference position are different, it can be directly said that the two target files are not duplicate files.
[0029] By applying the technical solution of this embodiment, first, two target files to be identified for repetitiveness are obtained, wherein the two target files to be identified for repetitiveness are files of the same file type. Next, the size of the storage space occupied by each of the two target files is determined. If the storage space occupied by the two target files is equal, then further, the difference position matching the file type of the two target files can be found from the prior knowledge base. Subsequently, according to the difference position provided by the prior knowledge base, the type features at the difference position are extracted from the two target files respectively. Further, the type features of the two target files are compared. If there are differences in the type features of the two target files at the difference position, it can be directly explained that the two target files are not duplicate files. The embodiment of the present application quickly locates the position where there may be differences in the file through the prior knowledge base, reduces unnecessary full file scanning, and reduces the consumption of computing resources; by extracting the type features at the difference position, only the difference position is analyzed, which greatly improves the speed of file comparison. In addition, the prior knowledge base stores the difference positions corresponding to different file types, so that the method is applicable to a variety of file types and has strong versatility.
[0030] In an embodiment of the present application, optionally, step 103 includes: when the type features corresponding to the two target files are different, determining that the two target files are not duplicate files; when the type features corresponding to the two target files are the same, continuing to determine whether there are local differences between the two target files according to the metadata feature comparison strategy and / or the local hash comparison strategy, and if there are local differences between the two target files, determining that the two target files are not duplicate files.
[0031] In this embodiment, if the type characteristics of the two target files are different, then it can be immediately determined that the two target files are not duplicate files, because there are differences in their contents at key positions; if the type characteristics of the two target files are the same, it means that their contents at key positions are consistent, but this is not enough to determine that they are exactly the same target files. Because target files often contain a large amount of data, and the type characteristics only reflect part of it. At this time, the metadata characteristics of the two target files can be further compared according to the metadata feature comparison strategy. Metadata usually includes information such as the creation time, modification time, size, and permissions of the file. Although metadata does not directly reflect the file content, it can provide additional clues about the file identity. If any metadata feature of the two target files is different, then it can be considered that they are not duplicate files.
[0032] In addition, a local hash comparison strategy can be used. This strategy involves calculating hash values for certain parts of the file (such as key data blocks, specific chapters, or randomly selected fragments) and comparing these hash values. If two files show differences in the local hash comparison, they can be considered not to be duplicate files.
[0033] Specifically, you can choose to use the metadata feature comparison strategy, the local hash comparison strategy, or use both strategies for comprehensive judgment as needed. If one of the strategies indicates that there are differences in the files, then it can be finally determined that the two target files are not duplicate files. The embodiment of the present application can more accurately identify duplicate files through multi-level and multi-angle judgment strategies such as type feature comparison, metadata feature comparison and / or local hash comparison, and avoid direct byte-by-byte comparison of the full text through these methods, thereby improving the accuracy and efficiency of file deduplication.
[0034] In an embodiment of the present application, optionally, the "continue to determine whether there are local differences between the two target files according to the metadata feature comparison strategy and / or the local hash comparison strategy" includes: determining the target metadata according to the file types of the two target files, and extracting the metadata features corresponding to each target file according to the target metadata, and when the metadata features of the two target files are different, determining that there are local differences between the two target files; and / or, randomly extracting a preset number of groups of data segments with the same position and the same size of the two target files, and calculating the local hash value corresponding to each data segment, and when there is a group of data segments with different local hash values in the two target files, determining that there are local differences between the two target files.
[0035] In this embodiment, the metadata feature comparison strategy and / or the local hash comparison strategy can be used to further determine whether there are local differences between the two target files. The specific implementation process of the two strategies is as follows:
[0036] For the metadata feature comparison strategy, first, determine the metadata that needs to be compared based on the file types of the two target files (such as text, pictures, audio, video, etc.). Different file types contain different metadata. For example, text files contain metadata such as author, creation date, and modification date; while audio files contain metadata such as artist, album name, and playback time. Next, extract the corresponding target metadata from the two target files respectively to form metadata features. After that, compare the metadata features of the two target files. If their metadata features are different (such as different creation dates, different authors, etc.), it is determined that the two target files have local differences, that is, they are not exactly the same files.
[0037] For the local hash comparison strategy, first, a preset number of groups of data segments with the same position and the same size are randomly extracted from the two target files. These data segments can be any part of the file, but in order to ensure the fairness of the comparison, they should have the same starting position and size in the two files. Then, the local hash value is calculated for each extracted data segment. The hash value is a technology that maps data of any length to a fixed-length value, which can be used to quickly compare whether the data is the same. Subsequently, the local hash values of the corresponding data segments in the two target files are compared. If there is any group of data segments with different local hash values, it is determined that the two target files have local differences, that is, they are not exactly the same files.
[0038] In an embodiment of the present application, optionally, after “continuing to determine whether there are local differences between the two target files”, the method further includes: if there are no local differences between the two target files, respectively calculating the global hash value corresponding to each target file, and determining whether there are differences in the global hash values; when there are differences in the global hash values, determining that the two target files are not duplicate files; when there are no differences in the global hash values, performing a byte-by-byte comparison on the two target files, and determining whether the two target files are duplicate files based on the comparison results.
[0039] In this embodiment, after determining that there are no local differences between the two target files, the global hash value of each target file is calculated separately. The global hash value is a unique identifier obtained by performing a hash operation on the entire file content, which reflects the overall content of the file. Afterwards, the global hash values of the two target files are compared. If there are differences in the global hash values, that is, the hash values of the two files are different, then it can be determined that the two target files are not duplicate files because they have differences in overall content.
[0040] If there is no difference in the global hash values of the two target files, that is, their hash values are exactly the same, then a byte-by-byte comparison is required to confirm whether they are exactly the same files. Because although the same global hash value means that the two files are very similar in overall content, there may still be very small differences (for example, misjudgment caused by the collision of the hash function). Byte-by-byte comparison refers to comparing each byte of the two files one by one to confirm whether they are exactly the same. According to the results of the byte-by-byte comparison, if every byte of the two target files is exactly the same, then it can be determined that the two target files are duplicate files; if there is any byte difference, then it can be determined that the two target files are not duplicate files.
[0041] In an embodiment of the present application, optionally, before step 101, the method further includes: obtaining file samples under each file type, wherein the file samples under each file type include samples of all file formats under the file type, and the file types include text files, picture files, audio files, and video files; for the file samples under each file type, identifying the file header features and data segment features corresponding to various file formats under the file samples, and determining the difference positions corresponding to each file format based on the file header features and the data segment features, and using the difference positions as the difference positions corresponding to the file type; and constructing the prior knowledge base based on the difference positions corresponding to each file type.
[0042] In this embodiment, the prior knowledge base can be determined as follows. First, obtain file samples, where the file samples cover all file formats under each file type. That is, for each file type (such as text files, picture files, audio files, video files, etc.), it is necessary to collect all possible file formats under this type as samples. For example, for text files, the corresponding file formats may include: (1).txt: the most basic text file format, widely used to store plain text information. (2).doc and .docx: Microsoft Word document formats, used to create and edit complex texts and documents. (3).pdf: Portable Document Format, used to share and print documents across platforms while maintaining the format and layout of the documents. (4).md: Markdown format, a lightweight markup language used to create plain text format documents that are easy to read and write. (5).html and .htm: Hypertext Markup Language format, used to create web pages and web page content. (6).rtf: Rich text format, documents that support text formatting, such as fonts, colors, paragraphs, etc.
[0043] For each file sample, identify the file header features and data segment features corresponding to various file formats under the file sample. Among them, the file header features are usually located at the beginning of the file, including the file identification information, type information, encoding information, etc., which is an important basis for identifying the file type. For example, for image files, the file header features can include magic values, such as jpg files usually start with FF D8 FF, and png files usually start with 89 50 4E 47; the data segment feature is the main content part of the file, containing actual data or information. For different types of files, the structure and content of the data segment will be very different. For example, the data segment features of a text file may include: line terminators: such as \r\n in Windows, \n in Unix / Linux, etc., these symbols mark the end of the text line; content structure: such as whether it contains specific markup languages (such as HTML, XML), programming languages (such as Python, Java) structural features. The data segment features of image files may include: color depth: the number of bits used for each pixel, such as 24-bit true color (8 bits for each color channel); compression algorithm: such as JPEG's discrete cosine transform (DCT) compression, PNG's lossless compression (such as DEFLATE), etc. The data segment features of audio files may include: audio data: the main body of an audio file is audio data, which can be PCM (pulse code modulation) raw data or compressed data (such as MPEG audio compression in MP3). The data segment features of video files may include: video frame data: the main body of a video file is continuous video frames, which can be uncompressed raw image data or compressed frame data (such as H.264, VP9 and other video coding standards).
[0044] After identifying the various features of the file, the difference positions corresponding to each file format under each file type are determined based on these features. The difference positions here refer to which features or positions are different between different file formats. These differences may be caused by different file formats, different encoding methods, different data organization methods, etc. The purpose of determining the difference positions is to more accurately identify the differences between different files in the subsequent repetitive recognition process, thereby improving the accuracy and efficiency of recognition.
[0045] Finally, the determined difference position is recorded as the difference position corresponding to the file type. That is, for each file type, a difference position database or table is established to store the difference position information between various file formats under this type. In this way, these difference position information can be directly consulted in the subsequent repetitive recognition process.
[0046] The embodiment of the present application obtains comprehensive and diverse file samples and deeply analyzes the file header features and data segment features of these samples to determine the difference positions between various file formats under different file types. This information can help to more accurately and quickly identify the differences between different files, thereby improving the accuracy and efficiency of recognition.
[0047] In an embodiment of the present application, optionally, the "obtaining two target files to be identified for repetitiveness" in step 101 includes: in response to an operation instruction to add files to be compared to a preset window, obtaining two files to be compared added to the preset window, and using the two files to be compared as two target files for repetitiveness identification; and / or, obtaining files of the same file type from a designated space according to a preset polling rule, and using every two files of the same file type as two target files for repetitiveness identification.
[0048] In this embodiment, two target files to be repeatedly identified can be determined by adding the files to be compared to a preset window. The preset window can be a file selection window in a graphical user interface (GUI) or a file input prompt in a command line interface (CLI). The user can add the files to be compared to the preset window. This adding process can be operations such as selecting, dragging, and uploading files, depending on the design of the preset window and the interactive mode of the user interface. Once the files are added to the preset window, these files can be obtained and used as target files for repeatability identification.
[0049] In addition, two target files to be repeatedly identified can be determined by obtaining files from a designated space according to preset polling rules. The designated space can be a folder, a database, a cloud storage service, or any other system that can store files, and this space contains a set of files that need to be repeatedly identified. Specifically, files can be obtained from this designated space according to preset polling rules. Since two files are selected from multiple files for comparison each time, a mechanism can be preset to group or pair these files. For example, it can be achieved by random selection, sequential pairing, etc. Once the files are obtained and grouped, the two files in each group can be used as target files for repeated identification.
[0050] The embodiments of the present application provide different ways to obtain two target files to be identified for repetitiveness. The first method relies on user interaction to specify the files to be compared, which is suitable for situations where user participation or flexible file selection is required. The second method is to automatically determine the target file, which is suitable for scenarios where a large number of files need to be processed and repetitiveness needs to be identified automatically.
[0051] In an embodiment of the present application, optionally, the method further includes: after determining that any two files are not duplicate files, generating a storage data according to the file identifiers and file generation time corresponding to the any two files, and storing the storage data in a target list; accordingly, after the "obtaining two target files to be identified for duplication" in step 101, the method further includes: based on the file identifiers of the two target files, querying through the target list whether the storage data corresponding to the two target files are included; when the target list does not include the storage data corresponding to the two target files, executing the step of separately determining the storage space occupied by each target file; when the target list includes the storage data corresponding to the two target files, but the generation time of at least one target file does not match the storage data, executing the step of separately determining the storage space occupied by each target file; when the target list includes the storage data corresponding to the two target files, and the generation time of the two target files matches the storage data, it is determined that the two target files are not duplicate files.
[0052] In this embodiment, unnecessary repeated checks on known non-duplicate files can be avoided. First, after determining that any two files are not duplicate files, storage data can be generated based on the file identification and generation time of the two non-duplicate files. Among them, the file identification is an identifier used to uniquely identify a file, which can be a file name, a file hash value, or other information that can ensure the uniqueness of the file; the file generation time is the timestamp when the file is created or last modified. Afterwards, these storage data are stored in a target list, which can be a database, a memory data structure, or any other system that can efficiently store and retrieve data. In this way, the target list can store information on a group of files that are determined to be non-duplicate files each time a repeatability identification is performed.
[0053] The target list can be used directly later. Specifically, after obtaining the two target files to be repetitively identified, the target list can be queried based on the file identifiers of the two target files to check whether the storage data corresponding to the two target files already exist in the target list. Specifically, if the file identifiers of the two files corresponding to a certain storage data in the target list are the same as the file identifiers of the above two target files, then it means that the storage data corresponding to the two target files already exist in the target list.
[0054] If there is no storage data corresponding to the two target files in the target list, this means that the two target files have not been verified as non-duplicate files before. Therefore, continue to perform the step of determining the storage space occupied by each target file respectively, and then perform duplicate detection normally according to the steps of steps 102 to 103.
[0055] If the target list contains storage data corresponding to the two target files, but the generation time of at least one target file does not match the time in the storage data, this means that the file may have been modified or updated. Therefore, it is also necessary to perform the steps of determining the storage space occupied by each target file, and then perform duplicate detection according to the steps 102 to 103 to re-verify the duplication of the file.
[0056] If the target list contains storage data corresponding to the two target files, and the generation time of the two target files is exactly the same as the time in the storage data, this means that the two files have been verified as non-duplicate files before and have not been modified since then. Therefore, it can be directly determined that the two target files are not duplicate files without further duplication check.
[0057] The embodiment of the present application implements an efficient file duplication identification optimization strategy by introducing storage data and using file identification and file generation time as key information. It avoids unnecessary repeated checks and improves the performance and efficiency of the system. At the same time, it also takes into account the possibility of file changes and ensures the accuracy and validity of the stored data by checking the file generation time. This strategy can greatly reduce the workload of duplication checks and improve the efficiency and accuracy of file management.
[0058] Further, as Figure 1 The specific implementation of the method, the embodiment of the present application provides a duplicate file identification device, such as Figure 2 As shown, the device comprises:
[0059] A storage space determination module, used to obtain two target files to be identified for duplication, and respectively determine the storage space occupied by each target file, wherein the two target files are of the same file type;
[0060] A difference position determination module, used for, when the storage spaces occupied by the two target files are equal, querying the difference positions corresponding to the file types from a priori knowledge base based on the file types of the two target files, and respectively extracting the type features corresponding to each target file at the difference positions;
[0061] The duplicate file determination module is used to determine whether the two target files are duplicate files based on the type feature.
[0062] Optionally, the duplicate file determination module is used to:
[0063] When the type features corresponding to the two target files are different, determining that the two target files are not duplicate files;
[0064] When the type features corresponding to the two target files are the same, the metadata feature comparison strategy and / or the local hash comparison strategy are used to continue to determine whether there are local differences between the two target files. If there are local differences between the two target files, it is determined that the two target files are not duplicate files.
[0065] Optionally, the duplicate file determination module is further used to:
[0066] Determining target metadata according to the file types of the two target files, and extracting metadata features corresponding to each target file according to the target metadata, and determining that there are local differences between the two target files when the metadata features of the two target files are different; and / or,
[0067] A preset number of data segments of the same position and size are randomly extracted from the two target files, and a local hash value corresponding to each data segment is calculated respectively. When a group of data segments with different local hash values exists in the two target files, it is determined that there is a local difference between the two target files.
[0068] Optionally, the device further comprises a full-text hash value comparison module; the full-text hash value comparison module is used to:
[0069] After continuing to determine whether there is a local difference between the two target files, if there is no local difference between the two target files, respectively calculating the global hash value corresponding to each target file, and determining whether there is a difference between the global hash values;
[0070] When there is a difference between the global hash values, determining that the two target files are not duplicate files;
[0071] When there is no difference between the global hash values, the two target files are compared byte by byte, and whether the two target files are duplicate files is determined according to the comparison result.
[0072] Optionally, the device further comprises a priori knowledge base construction module; the priori knowledge base construction module is used to:
[0073] Before obtaining the two target files to be identified for repetitiveness, obtaining file samples of each file type, wherein the file samples of each file type include samples of all file formats of the file type, and the file types include text files, picture files, audio files, and video files;
[0074] For a file sample under each file type, identify the file header features and data segment features corresponding to various file formats under the file sample, and determine the difference position corresponding to each file format according to the file header features and data segment features, and use the difference position as the difference position corresponding to the file type;
[0075] The prior knowledge base is constructed based on the difference position corresponding to each file type.
[0076] Optionally, the storage space determination module is used to:
[0077] In response to an operation instruction to add a file to be compared to a preset window, two files to be compared added to the preset window are acquired, and the two files to be compared are used as two target files for repeatability identification; and / or,
[0078] Files of the same file type are obtained from the specified space according to the preset polling rules, and every two files of the same file type are used as two target files for repeated identification.
[0079] Optionally, the device further comprises:
[0080] A storage module, configured to generate a piece of storage data according to the file identifiers corresponding to the two files and the file generation time after determining that the two files are not duplicate files, and store the storage data in a target list;
[0081] Accordingly, the device further includes a query module; the query module is used to:
[0082] After acquiring the two target files to be repetitively identified, querying whether the target list includes the storage data corresponding to the two target files based on the file identifiers of the two target files;
[0083] When the target list does not include the storage data corresponding to the two target files, executing the step of respectively determining the storage space occupied by each target file;
[0084] When the target list includes storage data corresponding to the two target files, but the generation time of at least one target file does not match the storage data, performing the step of respectively determining the storage space occupied by each target file;
[0085] When the target list includes storage data corresponding to the two target files, and the generation time of the two target files is consistent with the storage data, it is determined that the two target files are not duplicate files.
[0086] It should be noted that for other corresponding descriptions of the functional units involved in the duplicate file identification device provided in the embodiment of the present application, reference can be made to Figure 1 The corresponding description in the method will not be repeated here.
[0087] The present application also provides a computer device, which may be a personal computer, a server, a network device, etc. Figure 3 As shown, the computer device includes a bus, a processor, a memory and a communication interface, and may also include an input and output interface and a display device. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps in each method embodiment are implemented.
[0088] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0089] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium may be non-volatile or volatile, and stores a computer program thereon. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0090] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0091] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0092] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0093] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0094] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for identifying duplicate files, characterized in that: include: Acquire two target files to be identified for repetitiveness, and respectively determine the storage space occupied by each target file, wherein the two target files have the same file type; In the case where the storage spaces occupied by the two target files are equal, based on the file types of the two target files, querying the difference positions corresponding to the file types from a priori knowledge base, and extracting the type features corresponding to each target file at the difference positions respectively; Based on the type feature, it is determined whether the two target files are duplicate files.
2. The method according to claim 1, characterized in that The determining, based on the type feature, whether the two target files are duplicate files includes: When the type features corresponding to the two target files are different, determining that the two target files are not duplicate files; When the type features corresponding to the two target files are the same, the metadata feature comparison strategy and / or the local hash comparison strategy are used to continue to determine whether there are local differences between the two target files. If there are local differences between the two target files, it is determined that the two target files are not duplicate files.
3. The method according to claim 2, characterized in that The step of continuing to determine whether there are local differences between the two target files according to the metadata feature comparison strategy and / or the local hash comparison strategy includes: Determining target metadata according to the file types of the two target files, and extracting metadata features corresponding to each target file according to the target metadata, and determining that there are local differences between the two target files when the metadata features of the two target files are different; and / or, A preset number of data segments of the same position and size are randomly extracted from the two target files, and a local hash value corresponding to each data segment is calculated respectively. When a group of data segments with different local hash values exists in the two target files, it is determined that there is a local difference between the two target files.
4. The method according to claim 2 or 3, characterized in that: After continuing to determine whether there are local differences between the two target files, the method further includes: If there is no local difference between the two target files, respectively calculate the global hash value corresponding to each target file, and determine whether there is a difference between the global hash values; When there is a difference between the global hash values, determining that the two target files are not duplicate files; When there is no difference between the global hash values, the two target files are compared byte by byte, and whether the two target files are duplicate files is determined according to the comparison result.
5. The method according to claim 1, characterized in that: Before obtaining the two target files to be identified for repetitiveness, the method further includes: Obtaining file samples under each file type, wherein the file samples under each file type include samples of all file formats under the file type, and the file types include text files, picture files, audio files, and video files; For a file sample under each file type, identify the file header features and data segment features corresponding to various file formats under the file sample, and determine the difference position corresponding to each file format according to the file header features and data segment features, and use the difference position as the difference position corresponding to the file type; The prior knowledge base is constructed based on the difference position corresponding to each file type.
6. The method according to claim 1, characterized in that The step of obtaining two target files to be subjected to repeatability identification includes: In response to an operation instruction to add a file to be compared to a preset window, two files to be compared added to the preset window are acquired, and the two files to be compared are used as two target files for repeatability identification; and / or, Files of the same file type are obtained from the specified space according to the preset polling rules, and every two files of the same file type are used as two target files for repeated identification.
7. The method according to claim 1, characterized in that The method further comprises: After determining that any two files are not duplicate files, generating a piece of storage data according to the file identifiers corresponding to the any two files and the file generation time, and storing the storage data in the target list; Accordingly, after obtaining the two target files to be identified for repetitiveness, the method further includes: Based on the file identifiers of the two target files, querying through the target list whether the storage data corresponding to the two target files are included; When the target list does not include the storage data corresponding to the two target files, executing the step of respectively determining the storage space occupied by each target file; When the target list includes storage data corresponding to the two target files, but the generation time of at least one target file does not match the storage data, performing the step of respectively determining the storage space occupied by each target file; When the target list includes storage data corresponding to the two target files, and the generation time of the two target files is consistent with the storage data, it is determined that the two target files are not duplicate files.
8. A duplicate file identification device, characterized in that: include: A storage space determination module, used to obtain two target files to be identified for duplication, and respectively determine the storage space occupied by each target file, wherein the two target files are of the same file type; A difference position determination module, used for, when the storage spaces occupied by the two target files are equal, querying the difference positions corresponding to the file types from a priori knowledge base based on the file types of the two target files, and respectively extracting the type features corresponding to each target file at the difference positions; The duplicate file determination module is used to determine whether the two target files are duplicate files based on the type feature.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.