Duplicate file search method and electronic device

By performing preliminary screening of files and obtaining hash values ​​of files in a file group with the same hash value in parallel, the method solves the problem that the method for finding duplicate files in the prior art cannot determine duplicate files in a short time at a low performance cost, optimizes the performance of file search, achieves the effect of finding duplicate files at a low performance cost, and improves the utilization rate of storage resources and read and write performance of electronic devices.

CN117725026BActive Publication Date: 2025-09-26HONOR DEVICE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202311024645.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-14
Publication Date
2025-09-26
Estimated Expiration
2043-08-14

AI Technical Summary

Technical Problem

Existing technologies cannot find duplicate files in a short time with a low performance cost in scenarios with a large number of files, resulting in redundant storage resource usage and degraded read and write performance.

Method used

By performing preliminary screening based on file size, the first file group with the same file size is obtained, and the hash value is obtained in parallel using asynchronous threads, which shortens the scanning and calculation time and reduces the number of files and resource consumption for hash value calculation.

Benefits of technology

It can identify duplicate files in a short time at a low performance cost, reduce computing resource consumption, and optimize file search performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117725026B_ABST
    Figure CN117725026B_ABST
Patent Text Reader

Abstract

The present application discloses a method for searching for duplicate files and an electronic device, which relates to the technical field of electronic devices, including: the electronic device traverses the directories and files contained in the target directory, and obtains a first file group based on the size of each file; the first file group includes at least two files with the same file size. In the process of obtaining the first file group, the electronic device obtains the hash values ​​of the files contained in each first file group in parallel through an asynchronous thread, and determines that the files with the same hash value are duplicate files based on the hash value of each file. Preliminary screening based on file size can reduce the number of files used to calculate hash values, and parallel scanning of files and calculation of file sizes and hash values ​​can shorten the scanning and calculation time, reduce the resource consumption of hash calculations, greatly reduce the time used to find duplicate files, and optimize the performance of the electronic device in finding duplicate files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of electronic devices, and in particular to a method for searching for duplicate files and an electronic device. Background Art

[0002] During operation, electronic device systems or applications often generate identical files. These identical files can lead to redundant storage resources. Therefore, by finding and cleaning duplicate files, redundant storage resources can be freed up and file read and write performance can be optimized.

[0003] However, the methods for finding duplicate files in the prior art cannot achieve the purpose of finding duplicate files in a short time and at a low performance cost in a scenario with a large number of files. Summary of the Invention

[0004] An embodiment of the present application provides a method and electronic device for finding duplicate files. The electronic device can perform preliminary screening based on file size to obtain a first file group with the same file size. The hash values ​​of each file in each first file group are obtained in parallel through asynchronous threads and compared to identify duplicate files with the same hash value. The preliminary screening can reduce the number of files used to calculate hash values, reducing the resource consumption of hash calculations. The parallel scanning of files and calculation of file sizes and hash values ​​can shorten the scanning and calculation time, thereby optimizing the performance of the electronic device in finding duplicate files.

[0005] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions.

[0006] In a first aspect, a method for finding duplicate files is provided, the method comprising:

[0007] The electronic device traverses the directories and files contained in the target directory and obtains a first file group based on the size of each file; the first file group includes at least two files of the same size. While obtaining the first file group, the electronic device concurrently obtains hash values ​​of the files contained in each first file group via an asynchronous thread and, based on the hash values ​​of the files, determines that files with the same hash values ​​are duplicate files.

[0008] In this application, for all files in the same directory, the files can be preliminarily screened according to preset conditions. For example, by searching according to the size of each file, at least one first file group consisting of files of the same size can be obtained. For each file in the first file group, the hash value of each file is calculated, and files with the same hash value can be determined to be duplicate files with exactly the same content. In this solution, the preliminary screening can filter out files that do not have the same file size, thereby reducing the number of files for which hash values ​​are calculated. Performing file scanning and calculation of file size and hash values ​​in parallel can shorten the scanning and calculation time. Since calculating hash values ​​is a very complex calculation process, reducing the number of files for which hash values ​​are calculated greatly reduces the number of times hash values ​​are calculated, the computing resources occupied, etc., shortens the time for determining duplicate files based on hash values, and achieves the effect of determining duplicate files in a short time with lower performance.

[0009] In a possible implementation of the first aspect, traversing the directories and files contained in the target directory includes:

[0010] Traverse the directories and files contained in the target directory, add the directories in the first data structure to the end of the preset global scan linked list in the traversal order, and add the files in the second data structure to the end of the preset global scan linked list in the traversal order.

[0011] In this application, fentry is used to represent files. The data structure of fentry may include file number, file size (bytes), file modification time (mtime), file hash value, file name, pointer to parent directory, etc. Dentry is used to represent directories. The data structure of dentry may include directory depth, number of sub-items, directory number, directory name, pointer to parent directory, etc. Sub-items include sub-directories and files. When the electronic device traverses the sub-items contained in the target directory, it adds the directory in the form of dentry to the end of the preset global scan linked list, and adds the file in the form of fentry to the end of the preset global scan linked list, so as to facilitate the subsequent search operation of duplicate files based on the global scan linked list.

[0012] In another possible implementation of the first aspect, obtaining the first file group according to the size of each file includes:

[0013] Based on the file size of each file, each file is grouped to obtain at least one second file group; and from the at least one second file group, a first file group formed by files with the same file size is obtained.

[0014] In this application, an electronic device groups files based on file size. For example, files with sizes within the same threshold range can be grouped together, or a modulo operation can be performed on the file sizes based on a certain value to obtain multiple groups of files with the same remainder. Various methods for grouping files based on file size can roughly group files by file size, and then compare the file sizes of the files within each group, thereby improving the efficiency of file size comparison.

[0015] In another possible implementation of the first aspect, grouping the files based on the file sizes of the files to obtain at least one second file group includes:

[0016] When a file is added to a preset global scan linked list in a second data structure, the file size of the file is obtained; a remainder operation is performed based on the file size and a preset first value to obtain a remainder corresponding to the file; and at least one second file group formed by multiple files with the same remainder is obtained.

[0017] In the present application, the electronic device can roughly classify the files by file size by taking the modulus of the file size based on the file size and a preset first value, thereby obtaining multiple groups of files with the same modulus. Comparing the file sizes of the groups of files with the same modulus can improve the efficiency of file size comparison.

[0018] In another possible implementation of the first aspect, the preset first value is used to represent the number of red-black trees in a preset red-black forest; the root nodes of the red-black trees correspond to different remainders;

[0019] Acquiring at least one second file group formed by a plurality of files having the same remainder, comprising:

[0020] Based on the remainder of the file, the file is added to the corresponding red-black tree; if the node of the red-black tree to which the file is to be added already has a file, the file is added to the second file group corresponding to the existing file.

[0021] In the present application, a red-black forest consisting of N red-black trees can be set, and the root nodes of the N red-black trees are used to form an array. When inserting the fentry of a file into the corresponding position of the red-black tree, the file size can be used to perform a remainder operation on N first, and the remainder can be used as the array index to obtain the corresponding red-black tree root node, and then the fentry can be inserted into the red-black tree represented by this root node. That is, each red-black tree corresponds to a remainder value, and the files are grouped based on the file size. This ensures that the nodes in each red-black tree are roughly one-Nth of the total number of nodes in the red-black forest, reducing the depth of the red-black tree, and ensuring that files of the same size can be selected using the red-black forest. Setting a red-black forest consisting of multiple red-black trees can avoid the problem of reduced search efficiency caused by too many nodes leading to an excessive depth of the red-black forest.

[0022] In another possible implementation of the first aspect, obtaining, from at least one second file group, a first file group formed by files having the same file size includes:

[0023] Taking a file in the second file group as the first reference file, traverse the second file group to obtain a first matching file with the same file size as the reference file; if the file size of the first matching file is greater than or equal to a preset threshold, calculate a first check value of the target content in the first reference file and a second check value of the target content in the first matching file; the target content is used to represent the content of a preset number of bytes in the file; the first check value and the second check value are cyclic check codes.

[0024] If the first check value is the same as the second check value, the first data structure of the first matching file is added to the same-size linked list of the first reference file; the same-size linked list includes the first data structure of at least one first matching file with the same file size as the first reference file; the same-size linked list is used to represent the first file group.

[0025] The preset threshold value may be M, where M may represent the number of bytes. For example, M may be greater than 2. 13 (8192). The target content may include the first content and the second content in the file content. The first content may be the content from the 1st byte to the Nth byte in the file, and the first content may be the starting part of the content in the file. For example, the first content may be the content of 1-4096 bytes. The offset is 0 and the byte length is 4096. The second content may be determined based on M, and the second content may be the non-starting part of the content in the file. The offset of the second content may be M-4096, and the length may also be 4096. If M is 8192, then the second content is the content of bytes 4096-8192 in the file.

[0026] The electronic device calculates a check value based on the target content (eg, a total of 8192 bytes of data of the first content and the second content), and may calculate a CRC64 value of the target content as the check value (the check value may be referred to as a pattern).

[0027] In this application, for files with a file size greater than or equal to M, a checksum of the target content is calculated. If the checksums are different, then the target contents of the two files must be different, further preventing fentry files with the same file size but different actual contents from being added to the same-size linked list. At the same time, in the subsequent process of calculating file hash values ​​based on the same-size linked list by the electronic device, the number of files for which file hash values ​​are calculated can be effectively reduced, thereby reducing the computational burden of calculating file hash values.

[0028] In another possible implementation of the first aspect, the method further includes:

[0029] If the file size of the first matching file is smaller than a preset threshold, the first data structure of the first matching file is added to the same-size linked list of the first reference file.

[0030] In this application, for files whose file size is smaller than a threshold value M, the checksum of the target content is not calculated, which can save the consumption of computing resources and I / O bandwidth caused by calculating the checksum.

[0031] In another possible implementation of the first aspect, determining, based on the hash values ​​of the files, that files with the same hash values ​​are duplicate files includes:

[0032] Taking a file in the first file group as a second reference file, traversing the first file group to obtain a second matching file with the same hash value as the second reference file;

[0033] Among them, the first data structure of the second matching file is stored in the same hash value linked list of the second reference file; the same hash value linked list includes at least one first data structure of the second matching file with the same hash value as the second reference file; the second reference file and the second matching file are duplicate files.

[0034] In the present application, files with the same hash value are added to the same hash value linked list of the second reference file. The electronic device can traverse and store the hash values ​​of the files based on the same hash value linked list of the files, provide a basis for hash value search operations for non-first duplicate files, reduce the number of times the file hash values ​​are calculated, and thus reduce the computational burden of calculating the file hash values.

[0035] In another possible implementation of the first aspect, the method further includes:

[0036] The second reference file and the second matching file in the same hash value chain list of the second reference file are added to a preset duplicate file chain list.

[0037] In this application, the second reference file with the same hash value and the second matching file of the second reference file are added to a preset duplicate file linked list, which can provide data support for subsequent operations such as cleaning up duplicate files.

[0038] In another possible implementation of the first aspect, the method further includes:

[0039] The preset global scan linked list is traversed in reverse order, and the first data structure of the file without calculated hash value and the second data structure of the empty directory are deleted from the preset global scan linked list to obtain an updated global scan linked list.

[0040] In the present application, in an electronic device implementing a method for searching for duplicate files based on a global scan linked list, a same-size linked list, and a same-hash value linked list, the electronic device may also scan and delete fentry files in the global scan linked list for which hash values ​​have not been calculated, thereby optimizing the data validity of the global scan linked list. Files for which hash values ​​have not been calculated can be considered files that have been determined not to be duplicate files after the current scan is completed. Such files have not had their hash values ​​calculated and can be deleted from the global scan linked list.

[0041] In the present application, in another possible implementation of the first aspect, the method further includes:

[0042] Traverse the same hash value linked list of each second reference file and the updated global scan linked list, and store the first information of the directory and file in the updated global scan linked list into a preset first cache file; the first information includes an identifier, a parent directory number of the directory or file, and a directory name / file name; the identifier is used to characterize that the storage object is a file or directory; the hash value and modification time of the file in the updated global scan linked list are stored into a preset second cache file.

[0043] In the present application, the electronic device traverses the files and directories in the global scan list and can save the hash value of each file in a cache file in a compact manner. When the electronic device performs a non-initial repeated file scan, the hash value of the file in the cache file can be read, further reducing the consumption of computing resources used to calculate the file hash value.

[0044] In another possible implementation of the first aspect, obtaining hash values ​​of files included in the first file group in parallel through asynchronous threads includes:

[0045] If the second cache file does not exist, or the second cache file does not contain the hash value of the target file included in the first file group, or the first modification time of the target file in the second cache file is inconsistent with the second modification time of the target file in the first file group, the hash value of the file included in the first file group is calculated by an asynchronous thread;

[0046] If a second cache file exists, the second cache file contains hash values ​​of the files included in the first file group, the first modification time of the target file in the second cache file is consistent with the second modification time of the target file in the first file group, and the hash values ​​of the files corresponding to the files included in the first file group are obtained from the first cache file and the second cache file.

[0047] In this application, the absence of a second cache file indicates that this is the first time the electronic device is performing a duplicate file search operation. In this scenario, it is necessary to calculate the hash values ​​of the files contained in the first file group. The presence of a second cache file, and the hash value of the target file of the first file group contained in the second cache file, and the modification time of the target file has not changed, indicates that this is not the first time the electronic device is performing a duplicate file search operation, and the target file has not been modified. The hash value of the file in the cache file can be directly read, further reducing the consumption of computing resources used to calculate the file hash value.

[0048] In a second aspect, an electronic device is provided, comprising a memory and one or more processors; the memory is coupled to the processor; computer program code is stored in the memory, and the computer program code comprises computer instructions, which, when executed by the processor, enable the electronic device to perform a method as described in any one of the above-mentioned first aspects.

[0049] In a third aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium. When the computer-readable storage medium is executed on an electronic device, the electronic device can execute any one of the methods in the first aspect.

[0050] In a fourth aspect, a computer program product comprising instructions is provided, which, when executed on an electronic device, enables the electronic device to execute any one of the methods described in the first aspect.

[0051] In a fifth aspect, an embodiment of the present application provides a chip, the chip including a processor, the processor being used to call a computer program in a memory to execute the method of the first aspect.

[0052] It can be understood that the beneficial effects that can be achieved by the electronic device described in the second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect provided above can refer to the beneficial effects in the first aspect and any possible design method thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0054] Figure 2 A schematic diagram of a duplicate file search process provided in an embodiment of the present application;

[0055] Figure 3 A schematic diagram of the structure of a multi-branch tree corresponding to a target directory provided in an embodiment of the present application;

[0056] Figure 4A schematic diagram of the structure of a global scan linked table provided in an embodiment of the present application;

[0057] Figure 5 A schematic diagram of the structure of a red and black forest provided in an embodiment of the present application;

[0058] Figure 6 A schematic diagram of the structure of another red and black forest provided in an embodiment of the present application;

[0059] Figure 7 A schematic diagram of the structure of a duplicate file scanning result provided in an embodiment of the present application;

[0060] Figure 8 A schematic diagram of the structure of another global scan linked table provided in an embodiment of the present application;

[0061] Figure 9 A schematic diagram of the structure of a first cache file and a second cache file provided in an embodiment of the present application;

[0062] Figure 10 A schematic diagram of a module of an electronic device provided in an embodiment of the present application;

[0063] Figure 11 A schematic structural diagram of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0064] In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing specific embodiments, and are not intended to be used as limitations on the present application. As used in the specification and claims of the present application, the singular expressions "a", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear contrary indication in the context. It should also be understood that in the following embodiments of the present application, "at least one", "one or more" refer to one or more (including two). The term "and / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist; for example, A and / or B can represent: the situation where A exists alone, A and B exist at the same time, and B exists alone, wherein A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are a kind of "or" relationship.

[0065] References to "one embodiment" or "some embodiments" etc. described in this specification mean that the specific features, structures or characteristics described in conjunction with the embodiment are included in one or more embodiments of the present application. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. appearing in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in another way. The term "connected" includes direct and indirect connections, unless otherwise stated. "First" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated.

[0066] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0067] When running systems or applications on electronic devices, identical files may be generated. These identical files can cause redundant storage space usage and affect read and write performance. Therefore, it's necessary to find and identify duplicate files and clean them up to free up redundant storage resources and optimize read and write performance.

[0068] Existing technologies often determine whether there are duplicate files by traversing and matching the size or content of all files. For handheld consumer electronic products such as mobile phones and tablets, the methods for finding duplicate files provided by existing technologies cannot meet the performance and time constraints required by such electronic devices. In other words, it is impossible to achieve the effect of finding duplicate files in a short time with a low performance cost.

[0069] The embodiment of the present application provides a method for searching for duplicate files. For all files in the same directory, the files can be preliminarily screened according to preset conditions. For example, by searching according to the size of each file, at least one first file group consisting of files of the same size can be obtained. While the electronic device obtains the first file group, it can pull up an asynchronous thread in parallel to calculate the hash value of files of the same file size, greatly shortening the time for file scanning and data calculation; the preliminary screening can filter out files that do not have the same file size, which can greatly reduce the number of files for which hash values ​​are calculated. Since calculating hash values ​​is a very complex calculation process, reducing the number of files for which hash values ​​are calculated greatly reduces the number of times hash values ​​are calculated, the computing resources occupied, etc., shortening the time for determining duplicate files based on hash values, and achieving the effect of determining duplicate files in a short time with low performance.

[0070] The duplicate file search method provided in the embodiment of the present application can be applied to electronic devices. For example, Figure 1 A schematic diagram of the structure of an electronic device 100 is provided. The electronic device 100 in the embodiments of the present application may be a portable computer (such as a mobile phone), a tablet computer, a laptop computer, a personal computer (PC), a wearable electronic device (such as a smartwatch), an augmented reality (AR) or virtual reality (VR) device, an in-vehicle computer, or the like. The following embodiments do not impose any particular restrictions on the specific form of the electronic device.

[0071] Please refer to Figure 1 , which shows a structural block diagram of an electronic device (such as electronic device 100) provided in an embodiment of the present application. The electronic device 100 may include a processor 310, an external memory interface 320, an internal memory 321, a universal serial bus (USB) interface 330, a charging management module 340, a power management module 341, a battery 342, an antenna 1, a communication module 360, and the like.

[0072] The structure shown in the embodiment of the present invention does not limit the electronic device 100. The electronic device 100 may include more or fewer components than shown, or some components may be combined or separated, or arranged differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0073] The processor 310 may include one or more processing units. For example, the processor 310 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0074] The controller is the decision-maker that directs the various components of electronic device 100 to coordinate operations according to instructions. It serves as the nerve center and command center of electronic device 100. Based on instruction opcodes and timing signals, the controller generates operational control signals to control instruction fetching and execution.

[0075] Processor 310 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 310 is a high-speed cache memory that can store instructions or data that have just been used or are being recycled by processor 310. If processor 310 needs to use the same instruction or data again, it can directly access the memory. This avoids repeated accesses, reduces processor 310 latency, and thus improves system efficiency.

[0076] In some embodiments, the processor 310 may include an interface. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM interface, and / or a USB interface.

[0077] The interface connection relationship between the modules shown in the embodiment of the present invention is for illustrative purposes only and does not limit the structure of the electronic device 100. The electronic device 100 may adopt different interface connection methods or a combination of multiple interface connection methods in the embodiment of the present invention.

[0078] The charging management module 340 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 340 can receive charging input from the wired charger via the USB interface 330. In some wireless charging embodiments, the charging management module 340 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 342, the charging management module 340 can also provide power to the electronic device 100 via the power management module 341.

[0079] The power management module 341 is used to connect the battery 342, the charging management module 340, and the processor 310. The power management module 341 receives input from the battery 342 and / or the charging management module 340 and provides power to the processor 310, the internal memory 321, the external memory interface 320, and the communication module 360. The power management module 341 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some embodiments, the power management module 341 can also be provided in the processor 310. In some embodiments, the power management module 341 and the charging management module 340 can also be provided in the same device.

[0080] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the communication module 360, the modem and the baseband processor.

[0081] Antenna 1 is used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, a cellular network antenna can be reused as a wireless local area network diversity antenna. In some embodiments, the antenna can be used in conjunction with a tuning switch.

[0082] The modem may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium- or high-frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs the sound signal through an audio device. In some embodiments, the modem may be an independent device. In some embodiments, the modem may be independent of the processor 310 and be provided in the same device as other functional modules.

[0083] The communication module 360 ​​can provide a communication processing module for wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc., which are applied to the electronic device 100. The communication module 360 ​​can be one or more devices that integrate at least one communication processing module. The communication module 360 ​​receives electromagnetic waves via the antenna 1, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 310. The communication module 360 ​​can also receive the signal to be sent from the processor 310, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 1.

[0084] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the communication module 360 ​​so that the electronic device 100 can communicate with a network and other devices via wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning satellite system (satellite based augmentation systems, SBAS), a global navigation satellite system (GLONASS), a BeiDou navigation satellite system (BDS), a Quasi-Zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).

[0085] The external memory interface 320 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 310 via the external memory interface 320 to implement data storage. For example, in this embodiment, multiple directories and files can be stored in the external memory card. The duplicate file search method provided in this embodiment is performed on files in the external memory card.

[0086] The internal memory 321 can be used to store computer executable program codes, which include instructions. The processor 310 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 321. The memory 321 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the memory 321 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, other volatile solid-state storage devices, a universal flash storage (UFS), etc.

[0087] The embodiment of the present application provides a method for finding duplicate files, which is applied to an electronic device. The electronic device is the execution subject. The duplicate file finding process can refer to Figure 2 Shown, including:

[0088] S201: The electronic device traverses a target directory and obtains multiple files in the target directory.

[0089] The target directory is any root directory in the electronic device. The target directory may include at least one subdirectory and / or at least one file. A subdirectory may include at least one subdirectory and / or at least one file, and a subdirectory may also be an empty directory. For example, the file in this embodiment may be a document, image, video, audio, etc. containing multi-byte content. This embodiment is described using a document as an example.

[0090] For example, reference may be made to Figure 3 , Figure 3A multi-branch tree diagram of a target directory is given. Circular nodes are used to represent directories (or subdirectories), and rectangular nodes are used to represent files. Root node 1 represents the target directory (root directory). The target directory includes node 2 representing subdirectory 2, node 3 representing subdirectory 3, node 5 representing subdirectory 5, node 4 representing file 4, and node 6 representing file 6. Node 2 representing subdirectory 2 includes node 7 representing file 7 and node 8 representing subdirectory 8; node 8 representing subdirectory 8 includes node 12 representing file 12 and node 13 representing file 13. Node 3 representing subdirectory 3 includes node 10 representing file 10 and node 9 representing subdirectory 9; node 9 representing subdirectory 9 includes node 14 representing file 14 and node 15 representing file 15. Node 5 representing subdirectory 5 includes node 11 representing file 11. In this example, the multiple files acquired in the target directory include: file 4, file 6, file 7, file 10, file 11, file 12, file 13, file 14, and file 15.

[0091] The paths of files or subdirectories belonging to the same directory node are consistent. For example, the paths of file 14 and file 15 are both directory 1 / subdirectory 3 / subdirectory 9. The paths of subdirectory 8 and file 7 are both directory 1 / subdirectory 2.

[0092] For example, the electronic device can traverse the target directory by using a breadth traversal method or a depth traversal method. In this embodiment, the electronic device uses a multi-tree traversal process to traverse the target directory. After traversing the target directory, the electronic device can obtain all subdirectories and files contained in the target directory. Figure 3 The multi-branch tree of the target directory is traversed, and the result of each traversal is added to a preset linked list to obtain the overall traversal result of the target directory. The preset linked list can be a pre-set linked list for storing the traversal scan results of the target directory. For example, the preset linked list can be a preset global scan linked list.

[0093] Before traversing the target directory, the basic memory data structures for directories and files can be agreed upon. For example, a fentry is used to represent a file. The fentry data structure may include file number, file size (bytes), file modification time (mtime), file hash value, file name, and a pointer to the parent directory. A dentry is used to represent a directory. The dentry data structure may include directory depth, number of subitems, directory number, directory name, and a pointer to the parent directory. Subitems include subdirectories and files. Furthermore, multiple linked lists may be created for use in the duplicate file search process. For example, these multiple linked lists may include a global scan linked list, a same-size linked list, a same-hash linked list, and a global duplicate file linked list. The global scan linked list is used to store the target directory traversal scan results; the same-size linked list is used to store information about files with the same file size in the target directory; the same-hash linked list is used to store information about files with the same file hash value in the target directory; and the global duplicate file linked list is used to store information about files identified as duplicates in the target directory.

[0094] In this embodiment, the data structure of the file's fdentry can be added to the global scan linked list, the global duplicate file linked list, the same-size linked list, the same-hash linked list, and the red-black tree, and the fdentry data structure contains the index pointer members necessary for the data structure. The data structure of the directory's dentry can be added to the global scan linked list, and the dentry data structure contains the necessary pointer members. Optionally, in order to save memory, the directory number, directory depth, and number of sub-items in the dentry data structure can share an integer. For example, the directory number occupies a 64-bit wide character, the directory depth occupies 11 bits wide, and the number of sub-items occupies 53 bits wide. Since the directory number, directory depth, and number of sub-items are used at different times, the directory number, directory depth, and number of sub-items can share a 64-bit integer.

[0095] For example, the electronic device may first create a dentry (referring to the target directory) and add it to the empty global scan linked list, while setting the directory depth of the target directory to 0. The electronic device performs a breadth scan traversal on the target directory. If the current scan item is a directory, all subitems contained in the directory are added to the end of the global scan linked list in the scan order. Among them, for all subitems corresponding to the target, if the subitem is a directory, a dentry corresponding to the subitem is created and added to the end of the global scan linked list. If the subitem is a file, a fentry corresponding to the subitem is created and added to the end of the global scan linked list.

[0096] Optionally, when creating the data structure of each child item, set its pointer to the parent directory to the address of the current directory item, and set the directory depth of all child directories to the directory depth of their parent directory plus 1. Complete the scan of the current directory and set the current scan item to the next item in the global scan list.

[0097] For example, after scanning all sub-items in the target directory, we can get Figure 4 The global scan linked list shown in Figure 1 is shown in Figure 1. Circular nodes are used to represent directories, and rectangular nodes are used to represent files. Figure 4 This is just a schematic diagram of the global scan list. In practice, the global scan list stores the fdentry data structure of the file and the dentry data structure of the directory. Figure 4 , the backward arrows from node 1 to node 15 can represent the scanning order of breadth traversal, and the forward arrows of each child item are used to point to its corresponding parent directory. For example, the traversal order is node 1 to node 15. Among them, the parent directory of subdirectory 2 is the root directory, so node 2 representing subdirectory 2 points to node 1 representing the parent directory, and the directory depth of node 2 representing subdirectory 2 is 1. The parent directory of file 4 is the root directory, and node 4 representing file 4 points to node 1 representing the parent directory. The parent directory of subdirectory 8 is subdirectory 2, and node 8 representing subdirectory 8 points to node 2 representing subdirectory 2, and the directory depth of node 8 representing subdirectory 8 is 2. The parent directory of file 12 is directory 8, and node 12 representing file 12 points to node 8 representing directory 8, and so on.

[0098] The electronic device can obtain all files in the target directory by performing a broad scan on the target directory. Optionally, if the current scan item is a file, the electronic device can also add the file to a preset tree structure, thereby obtaining multiple files in the target directory.

[0099] S202: Acquire a first file group according to the size of each file; the first file group includes at least two files with the same file size.

[0100] In some embodiments, the electronic device can obtain the file sizes of all files in the target directory, traverse and match the file sizes of each file, and obtain 0, 1 or multiple file groups (first file groups) consisting of the same file size.

[0101] However, when there are a large number of files, traversing and matching the file sizes of each file consumes a long time and computing resources. In this embodiment, the electronic device can group the files according to their file size. For example, files with sizes within a first threshold range are grouped into a first group, files with sizes within a second threshold range are grouped into a second group, and so on. File sizes are then matched for each file within each group to identify files with the same file size, thereby obtaining a file group consisting of files with the same file size.

[0102] In actual implementation, for example, the electronic device may add multiple files to the corresponding nodes of the tree structure according to the file sizes, group the files based on the file sizes, and then further match the file sizes to obtain at least one first file group.

[0103] For example, to balance scanning of all files and matching file sizes, the preset tree structure can be a binary search tree, for example, a relatively balanced red-black tree. In some embodiments, when the tree structure is a red-black tree, during the electronic device's extensive scan of a target directory, if the current scan item is a file, the file is added to the corresponding location in the red-black tree. When adding a file to the red-black tree, the electronic device can calculate the file's size and insert the file's fentry into the corresponding node in the red-black tree based on the file's size.

[0104] Optionally, in a feasible implementation method of the electronic device inserting the file's fentry into the corresponding node in the red-black tree according to the file's size, the electronic device may set a red-black tree forest containing multiple red-black trees, each red-black tree corresponding to a file size range, and the electronic device inserts the file's fentry into the node corresponding to the red-black tree corresponding to the file size range in which the file size falls according to the file size. For example, a file whose file size falls within the first threshold range is inserted into the corresponding position in the first red-black tree, a file whose file size falls within the second threshold range is inserted into the corresponding position in the second red-black tree, and a file whose file size falls within the third threshold range is inserted into the corresponding position in the third red-black tree. In this way, each red-black tree will include at least one fentry of a file that falls within the corresponding file size range, and the electronic device can compare the file sizes of all files in the same red-black tree, determine files with the same file size, and obtain a first file group.

[0105] Optionally, in another feasible way in which the electronic device inserts the fentry of a file into the corresponding node in the red-black tree according to the size of the file, the electronic device can set a red-black forest consisting of N red-black trees, and use the root nodes of the N red-black trees to form an array. When inserting the fentry of the file into the corresponding position of the red-black tree, the file size can be first used to perform a remainder operation on N, and the corresponding red-black tree root node can be obtained using the remainder as the array index, and then the fentry is inserted into the red-black tree represented by this root node. That is, each red-black tree corresponds to a remainder value, so that files are grouped based on file size. This ensures that the nodes in each red-black tree are roughly one-Nth of the total number of nodes in the red-black forest, reduces the depth of the red-black tree, and ensures that files of the same size can be selected using the red-black forest.

[0106] In this embodiment, a red-black forest consisting of multiple red-black trees is set to avoid the problem of reduced search efficiency caused by too many nodes leading to an excessive depth of the red-black forest.

[0107] For example, the electronic device scans the target directory. If the scanned item is a file, the process of adding the file to the corresponding node of the red-black tree based on the file size can be referred to Figure 5 .in, Figure 5 A schematic diagram of a red-black forest consisting of eight red-black trees is provided. The root nodes of the eight red-black trees form the array [0, 1, 2, 3, 4, 5, 6, 7]. When the electronic device determines that the currently scanned item is a file, it calculates the file size, modulo the file size by the number of red-black trees (8), and, based on the remainder (0 / 1 / 2 / 3 / 4 / 5 / 6 / 7), inserts the file's fentry into the corresponding position of the red-black tree corresponding to the corresponding root node. If there is no fentry at the corresponding position in the red-black tree, the current fentry data structure is inserted into that position; if there is already a fentry at the corresponding position in the red-black tree, a fentry of the same file size is added to the preset linked list corresponding to the fentry at the current position. Here, the preset linked list can be a linked list of the same size.

[0108] For example, reference Figure 5 , the file size of file 1 is modulo 8, the remainder is 2, and the fentry of file 1 is inserted into the corresponding position of the red-black tree represented by the remainder 2. The file size of file 2 is modulo 8, the remainder is also 2, but the file size of file 2 is different from that of file 1, and the fentry of file 2 is inserted into another position of the red-black tree represented by the remainder 2 other than file 1. For example, the file size of file 3 is modulo 8, the remainder is also 2, the file size of file 3 is the same as that of file 2, but the fentry of file 2 has occupied a node, so the fentry of file 3 is added to the same-size linked list corresponding to file 2. Figure 5The processing of files 4 and 5 in the example is the same as that of file 3.

[0109] For example, the remainder of the file size of file 6 is also 2 when the modulo operation is performed on 8. However, the file size of file 6 is different from the file sizes of file 1 and file 2. Insert the fentry of file 6 at the other corresponding position in the red-black tree. The remainder of the file size of file 7 is also 2 when the modulo operation is performed on 8. However, the file size of file 7 is different from the file sizes of file 1, file 2, and file 6. Insert the fentry of file 7 at the other corresponding position in the red-black tree. The remainder of the file size of file 10 is also 2 when the modulo operation is performed on 8. The file size of file 10 is the same as the file size of file 7, but the fentry of file 7 has already occupied a node. Therefore, add the fentry of file 10 to the same-size linked list corresponding to file 7. Figure 5 The processing of file 11 in the example is the same as that of file 10.

[0110] Figure 5 In this example, the processing of file 8 is similar to that of file 7, and the processing of files 9, 12, and 13 is similar to that of file 10. Thus, a first file group is formed, consisting of multiple groups of files of the same file size. The corresponding same-size linked lists for each of the multiple first file groups include same-size linked list 1 for file 2, same-size linked list 2 for file 7, and same-size linked list 3 for file 8. Same-size linked list 1 includes the fentry corresponding to files 3, 4, and 5, which have the same file size as file 2; same-size linked list 2 includes the fentry corresponding to files 10 and 11, which have the same file size as file 7; and same-size linked list 3 includes the fentry corresponding to files 9, 12, and 13, which have the same file size as file 8.

[0111] In this embodiment, after determining that a node is occupied, the electronic device inserts files of the same size into the same-size linked list of files occupying the node, thereby implementing indexing of files of the same size. That is, if the electronic device determines that a same-size linked list exists for file 2, it can retrieve other files of the same size as file 2 from the same-size linked list, thus enabling rapid indexing of files of the same size.

[0112] After calculating the file size of each file in the target directory, at least one first file group with the same file size can be obtained based on the file size. The files in each first file group have the same file size.

[0113] Optionally, in some embodiments, there may be a problem where the files have the same size but different contents. For this scenario, the electronic device can also calculate a checksum of the target content in the file and use the checksum of the target content as a matching basis to match the target content consistency of the files of different file sizes.

[0114] In some embodiments, the electronic device may group files based on the file size ranges and the file size of each file to obtain a second file group containing files with file sizes within the same threshold range. Alternatively, the electronic device may group files based on the remainder obtained by performing a modulo operation on the first value based on the file size to obtain a second file group containing files with the same remainder. Furthermore, duplicate files with the same hash value are matched for content consistency to obtain a first file group containing files with the same file size.

[0115] For example, the electronic device may set a file size threshold M, where M may represent the number of bytes. For example, M may be greater than 2. 13 (8192) is a natural number. When inserting the fentry of the file to be inserted (the first matching file) into the corresponding position of the corresponding red-black tree, if there is no file with the same file size as the file to be inserted in the red-black tree, the fentry of the file is directly inserted into the corresponding position. If there is a file with the same file size as the file to be inserted in the red-black tree (referred to as the first reference file), the file size of the file to be inserted can be determined before adding the fentry of the file to be inserted to the same-size linked list of the reference files. If the file size of the file to be inserted is less than the threshold M, the fentry of the file to be inserted is directly added to the same-size linked list of the reference files. If the file size of the file to be inserted is equal to or greater than the threshold M, the target content in the file to be inserted is obtained and the first check value of the target content is calculated. The target content in the first reference file is obtained and the second check value of the target content is calculated. If the first check value is the same as the second check value, the fentry of the file to be inserted is added to the same-size linked list of the first reference file; if the first check value is different from the second check value, the fentry of the file to be inserted is inserted into another empty node in the red-black tree.

[0116] The target content may include the first content and the second content in the file content. The first content may be the content from the 1st byte to the Nth byte in the file, and the first content may be the starting part of the content in the file. For example, the first content may be content of 1-4096 bytes. The offset is 0 and the byte length is 4096. The second content may be determined based on M, and the second content may be the non-starting part of the content in the file. The offset of the second content may be M-4096, and the length may also be 4096. If M is 8192, then the second content is the content of bytes 4096-8192 in the file.

[0117] The electronic device calculates a check value based on the target content (eg, a total of 8192 bytes of data of the first content and the second content), and may calculate a CRC64 value of the target content as the check value (the check value may be referred to as a pattern).

[0118] If the pattern of the file to be inserted is the same as the pattern of the reference file, it means that the file to be inserted (the first matching file) and the first reference file have the same file size, and the target content used to calculate the pattern may be the same. It is very likely that the file to be inserted (the first matching file) and the first reference file have the same content. The fentry of the file to be inserted is added to the same-size linked list of the first reference file. If the pattern of the file to be inserted is different from the pattern of the first reference file, it means that the file to be inserted and the first reference file have the same file size, but the target content of the file to be inserted and the first reference file are inconsistent. If the target content is inconsistent, the entire content of the file to be inserted and the first reference file are definitely inconsistent. Therefore, the fentry of the file to be inserted is not added to the same-size linked list of the first reference file, and the subsequent hash value calculation and comparison of the file to be inserted is not performed. The fentry of the file to be inserted is added to other empty nodes in the red-black tree.

[0119] In this embodiment, the file size and pattern are used as the matching basis for adding files to the nodes of the red-black tree to further illustrate the process of generating the red-black tree. Figure 6 As shown. Setting Figure 6 The red-black tree shown is a remainder 2 red-black tree. In this red-black tree, root node 1, child nodes 2, and child nodes 3 each have a corresponding file's fentry added. The files corresponding to root node 1, child nodes 2, and child nodes 3 have different file sizes and patterns; however, the remainder calculated based on the file size and the number of red-black trees is always 2.

[0120] In the process of adding a file's fentry to a red-black tree node, for example, if the file size of the file to be inserted with a remainder of 2 is smaller than the file size of the root node, the left side of the root node is queried. The left side of the root node has a child node 2, and the fentry of the file has been added to the child node 2. If the file size of the file to be inserted is smaller than the file size of the file of the child node 2, and the position of the left child node of the child node 2 is empty, the file to be inserted is added to the left child node of the child node 2 (for example, Figure 6 If the file size of the file to be inserted is larger than the file size of the child node 2, and the position of the right child node of the child node 2 is empty, then the file to be inserted is added to the right child node of the child node 2 (for example, Figure 6If the file size of the file to be inserted is equal to the file size of child node 2, the pattern of the file to be inserted is compared with the pattern of the file already added to child node 2.

[0121] If the pattern of the file to be inserted is the same as the pattern of the file already added to child node 2, then the file to be inserted will be added to the same size linked list of files already added to child node 2. If the pattern of the file to be inserted is smaller than the pattern of the file already added to child node 2, and the position of the left child node of child node 2 is empty, then the file to be inserted will be added to the left child node of child node 2 (for example, Figure 6 Node 4 in the child node). If the pattern of the file to be inserted is larger than the pattern of the file already added in child node 2, and the position of the right child node of child node 2 is empty, then the file to be inserted is added to the right child node of child node 2 (for example, Figure 6 Node 5 in the example). This completes the goal of obtaining a linked list of files of the same size based on the file size and pattern as the initial screening criteria.

[0122] If the file size of the file to be inserted, with a remainder of 2, is larger than the file size of the root node, the query is performed to the right of the root node. The left side of the root node contains child node 3, which contains the file's fentry. The file size and pattern matching process for the file to be inserted based on the fentry of the file added to child node 3 can be compared with the file size and pattern matching process for the file added to child node 2 and the file to be inserted, which is not detailed here.

[0123] That is, in the above process of generating a red-black tree, the file size can be used as the query basis first. If the file sizes are different, the file to be inserted is inserted into the corresponding empty node in the red-black tree; if the file sizes are the same, the pattern is used as the matching basis. When the patterns are the same, the file to be inserted is added to the same-size linked list corresponding to the file of the current node; if the patterns are different, the file to be inserted is inserted into the corresponding empty node in the red-black tree.

[0124] In this embodiment, for files whose file size is less than a threshold value M, the checksum of their target content is not calculated, which can save the consumption of computing resources and I / O bandwidth caused by calculating the checksum. For files whose file size is greater than or equal to M, the checksum of the target content is calculated. If the checksums are different, it means that the target contents of the two files must be different, and all the contents of the two files must be different. This further prevents the fentry of files with the same file size but different actual content from being added to the same-size linked list. At the same time, in the subsequent process of electronic devices calculating file hash values ​​based on the same-size linked list, the number of files for which file hash values ​​are calculated can be effectively reduced, thereby reducing the computational burden of calculating file hash values.

[0125] The same-size linked list in this embodiment represents the first file group to a certain extent, and the same-size linked list includes multiple files of the same file size.

[0126] S203: Calculate hash values ​​of the files included in each of the first file groups.

[0127] In some embodiments, after obtaining at least one first file group, the electronic device calculates a hash value for each file in the first file group. The method for calculating the hash value may be Message Digest Algorithm (MD5), Secure Hash Algorithm 1 (SHA-1), or Secure Hash Algorithm 256 (SHA-256).

[0128] Corresponding to the above embodiment, when the electronic device inserts the file to be inserted into the red-black tree, if the node in the red-black tree has a fentry of the file (second reference file), that is, the file to be inserted is the same size as the file already in the node (belonging to the same first file group), then the first hash value of the second reference file and the second hash value of the file to be inserted are calculated. If the node in the red-black tree does not have a fentry of the file, the fentry of the file to be inserted is directly added to the empty node in the red-black tree. That is, there is currently no second reference file in the red-black tree with the same file size as the file to be inserted, that is, there is no first file group corresponding to the file to be inserted.

[0129] Optionally, in some embodiments, the electronic device calculates the file size of the file to be inserted and, when adding the file to be inserted to a non-empty node in the red-black tree, may launch an asynchronous thread to calculate the hash value of the file. Specifically, when the electronic device determines that a second reference file with the same file size as the file to be inserted exists in the red-black tree, it launches an asynchronous thread to calculate the hash value of the file.

[0130] After the electronic device inserts all files in the target directory into the red-black tree, it completes the size matching of all files and obtains at least one first file group. Simultaneously, the electronic device can calculate hash values ​​for files of the same size in parallel, enabling simultaneous multi-threaded processing of scanning files in the target directory, calculating file sizes, and calculating hash values ​​for files of the same size. This significantly reduces the computation time for the entire process of scanning files, calculating file sizes, and calculating hash values, thereby improving the efficiency of file scanning and related data calculation.

[0131] S204: Obtain duplicate files with the same hash value according to the hash value of each file.

[0132] In this embodiment, the electronic device traverses the red-black tree forest. For the fentry of the file at each node, if the same-size linked list of the file is not empty, the electronic device starts an asynchronous thread to traverse the fentry of all files in the same-size linked list. For example, the fentry of the file at the red-black tree node is used as the fentry of the second reference file, and the fentry of the second reference file is used as the current fentry. The fentry with the same file hash value as the current fentry is traversed and searched in its same-size linked list (fentry of the second matching file). Every time such a fentry is found, it is deleted from the original same-size linked list and then added to the same hash value linked list corresponding to the current fentry. In this way, the electronic device can find all fentry with the same hash value as the reference fentry, and the files corresponding to each fentry form duplicate files with the same hash value. The electronic device traverses the fentry of the files at each node in the red-black tree in the red-black tree forest, and obtains 0 groups, 1 group, or multiple groups of duplicate files with the same hash value.

[0133] In this embodiment, the same hash value linked list can be used to store fentry of duplicate files with the same hash value. The same hash value linked list of the second reference fentry stores the fentry of at least one file with the same hash value.

[0134] In some embodiments, the electronic device can also obtain duplicate files based on the fentry of the files in the same hash value linked list of each file. Optionally, the electronic device can also add the duplicate files to a preset global duplicate file linked list. Exemplarily, the electronic device traverses the fentry (second reference file) of the files of each node in the red-black tree. If the second reference file has a same hash value linked list, and the same hash value linked list is not empty, the fentry of the second reference file is added to the preset global duplicate file linked list until the electronic device has traversed all nodes in the red-black tree. The same hash value linked list of the second reference file is linked to the global duplicate file linked list, wherein the global duplicate file linked list can be a secondary linked list.

[0135] For example, Figure 7 A schematic diagram of a global duplicate file linked list is given. Among them, the files represented by the solid-line box are files that have the same hash value linked list in each node of the red-black tree traversed by the electronic device, and the same hash value linked list is not empty; the files represented by the dotted-line box are files in the same hash value linked list of the solid-line box files. That is, Figure 7 As shown, each column is a duplicate file group with the same hash value.

[0136] In an embodiment of the present application, files requiring hash value calculation are screened out based on the primary condition that the file size and pattern are equal, thereby reducing the number of files required for hash value calculation and lowering the computational load of the electronic device.

[0137] Optionally, in an electronic device implementing a duplicate file search method based on a global scan linked list, a same-size linked list, and a same-hash value linked list, the electronic device may also scan and delete fentry files in the global scan linked list for which hash values ​​have not been calculated, thereby optimizing the data validity of the global scan linked list. Files for which hash values ​​have not been calculated can be considered files that have been determined not to be duplicates after the current scan. Such files, for which hash values ​​have not been calculated, can be deleted from the global scan linked list.

[0138] Optionally, the electronic device deletes the fentry of the file whose hash value is not calculated in the global scan linked list in reverse order.

[0139] by Figure 4 Taking the global scan linked list as an example, during the duplicate file search process, the hash values ​​of files 4, 6, 7, 10, 12, and 13 in the global scan linked list are calculated. The hash values ​​of files 11, 14, and 15 are not calculated.

[0140] The electronic device searches file 15 in reverse order and determines that file 15 has no hash value. The electronic device deletes the fentry of file 15 in the global scan list and reduces the number of child items of its parent directory by 1.

[0141] The electronic device continues searching in reverse order until it reaches file 14. Determining that file 14 has no hash value, the electronic device deletes the fentry for file 14 from the global scan list and simultaneously reduces the number of children in its parent directory by 1. The parent directory of both files 14 and 15 is directory 9. After deleting the fentry for files 14 and 15, the number of children in directory 9 decreases from 2 to 0, making directory 9 empty. The electronic device can also delete the dentry for empty directory 9. In other words, if the number of children in a directory in the global scan list is 0, the electronic device can delete the directory.

[0142] The electronic device continues to search in reverse order to file 13, which has a hash value; the electronic device continues to search in reverse order to file 12, which has a hash value.

[0143] The electronic device continues searching in reverse order until it reaches file 11. File 11 has no hash value. The electronic device deletes the fentry of file 14 from the global scan list and reduces the number of children in its parent directory by 1. The parent directory of file 11 is directory 5. After deleting the fentry of file 14, the number of children in directory 5 decreases from 1 to 0, making directory 5 empty. The electronic device can then delete the dentry of empty directory 5.

[0144] File 7, file 6, and file 4 all have hash values, and the electronic device does not perform the deletion operation during the reverse query process. Therefore, the original global scan list is deleted from the fentry of files without hash values ​​and the dentry of directories with 0 subitems, and the following is obtained: Figure 8 Given a global scan list. Figure 8 and Figure 4 , you can see Figure 8 In the global scan list, the fentry for files 11, 14, and 15, which had no hash values, have been deleted. The dentry for directory 5 and the dentry for directory 9, which had zero child items, have also been deleted. Only the fentry for the file whose hash value was calculated and the dentry for its corresponding parent directory remain in the global scan list. This improves the validity of the files and directories in the global scan list.

[0145] Based on deleting the global scan linked list of empty directories and files without hash values, the electronic device can save the hash values ​​of the files in the global scan linked list.

[0146] In some embodiments, the electronic device traverses the files and directories in the global scan list and can compactly store the hash values ​​of each file in a cache file in the memory. When the electronic device performs a non-initial repeated file scan, the hash value of the file in the cache file can be read, further reducing the consumption of computing resources used to calculate the file hash value.

[0147] Optionally, to ensure the validity of the calculated hash value, the file's modification time can also be saved in the cache file when the file's hash value is saved. If the modification time remains unchanged, the file has not been modified, and the calculated hash value remains valid. If the modification time changes, the file content has been modified, and the hash value calculated based on the file content before the modification becomes invalid.

[0148] Exemplarily, the electronic device may create two cache files, wherein the first cache file is used to store information of a file or directory; and the second cache file is used to store a hash value and modification time of the file.

[0149] refer to Figure 9 , Figure 9 A schematic diagram of a first cache file and a second cache file is provided. The first cache file may also be called a name file, and the second cache file may also be called a hash file.

[0150] The electronic device sets a 64-bit wide header identification header data structure in the first cache file to assist in saving the directories or files in the global scan linked list. Among them, 1 bit in the header identification structure is used to represent the directory or file. For example, when the value of 1 bit is the first value, it indicates that the current information is the information of the file. The first value can be 1. When the value of 1 bit is the second value, it indicates that the current information is the information of the directory, and the second value can be 0. The 12 bits in the header identification structure are used to save the name length, and the remaining 51 bits in the header identification structure save the parent directory number of the directory or file. The information saved in the first cache file includes the header identification and the directory name / file name. Among them, the length of the directory name / file name is the length indicated by 12 bits. The second cache file is used to sequentially save the modification time and hash value of the file.

[0151] In some embodiments, the electronic device traverses the global scan linked list in positive order from the beginning, while using an integer as a directory number with an initial value of 0. Whenever a dentry is traversed, the value of the directory number is written into the directory number member of the dentry, and 1 is added to the directory number. The directory name length and parent directory number of the current dentry are written into the header, and the 1st bit of the header is set to the second value. The root directory in the global scan linked list has no parent directory, and the parent directory number in its header is written to zero. The header structure and the directory name of the current directory are written into the first cache file in sequence.

[0152] Whenever the electronic device traverses a fentry, it writes the file name length and parent directory number of the current fentry into the header and sets the first bit of the header to the first value. The header structure and the current file name are sequentially written to the first cache file. Simultaneously, the current file's modification time and hash value are sequentially written to the second cache file. This completes the traversal of the global scan linked list, ensuring the compact storage of the file's hash value and modification time in the second cache file.

[0153] When the next run is performed, the cached file is first read and a hash table is created that maps the file path to the file hash and the file modification time. In the above-mentioned "scan and hash merge method", the hash table is first queried before the asynchronous thread is started to calculate the file hash value of a fentry. If the hash value of the fentry exists in the hash table and the corresponding file modification time is the same as the current modification time of the file, it means that the cached hash is still valid. The cached hash value is then used as the file hash value of the current fentry, thus saving one hash operation. In actual operation, this method can avoid most hash operations and greatly improve the execution time after the first run (no cached file is available during the first run).

[0154] When an electronic device performs a non-first search operation for duplicate files, it can read the information in the first cache file and the second cache file, and establish a two-tuple hash table that maps the file path to the file hash value and the file modification time. When the electronic device needs to obtain the hash value of file A for hash value matching, and obtains duplicate files with the same hash value, the electronic device can first query the two-tuple hash table whether there is information corresponding to file A. If it exists, and the modification time of file A is consistent with the modification time of file A in the two-tuple hash table, it means that there is no content modification in file A, and the hash value of file A is still valid. The electronic device can directly read the hash value of file A for subsequent hash value matching operations, and the electronic device no longer needs to calculate the hash value based on the content of file A. In actual operation, most hash operations can be avoided, which can greatly improve the execution time of non-first duplicate file searches.

[0155] In addition, the first cache file and the second cache file respectively store variable-length file names and fixed-length file modification times and hash values, so that the process of saving and reconstructing the hash value of the file only requires reading the content of a specific length in sequence when necessary, with simple logic and high execution efficiency. The contents of the first cache file and the second cache file are all saved in binary format, with extremely high information density, which saves the storage space occupied by the cache file. The first cache file and the second cache file use directory numbering to retain the directory tree information in the process of sequentially saving directory information, which means that each file does not need to save its complete path, further saving the space occupied by the cache file.

[0156] In this embodiment, after obtaining the first file group based on file size, the electronic device obtains the corresponding hash value for each file in the first file group and matches it to obtain duplicate files with the same hash value. By filtering the files by size, the number of files used to calculate the hash value is greatly reduced. Moreover, when searching for duplicate files that are not the first time, the electronic device can also read the file's hash value from the cache file, further reducing the resource consumption of the hash calculation, effectively improving the efficiency of duplicate file search, and achieving the effect of duplicate file search in a short time with minimal performance consumption. This method can be applied to electronic devices with high timeliness and performance requirements.

[0157] Figure 10 A possible structural diagram of the electronic device involved in the above embodiments is shown. Figure 10 The electronic device 1000 shown includes a processor 1001 and a storage module 1002 .

[0158] The processor 1001 may be a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The processor may include an application processor and a baseband processor. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The storage module 1002 may be a memory.

[0159] For example, the processor 1001 may be Figure 1 The processor 310 shown; the storage module 1002 can be as follows Figure 1 The internal memory 321 shown. The electronic device provided in the embodiment of the present application can be Figure 1 The electronic device 100 is shown.

[0160] The present application also provides a chip system (eg, a system on a chip (SoC)). Figure 11 As shown, the chip system includes at least one processor 701 and at least one interface circuit 702. The processor 701 and the interface circuit 702 can be interconnected via lines. For example, the interface circuit 702 can be used to receive signals from other devices (such as a memory of an electronic device). For another example, the interface circuit 702 can be used to send signals to other devices (such as a processor 701 or a camera of an electronic device). Exemplarily, the interface circuit 702 can read an instruction stored in the memory and send the instruction to the processor 701. When the instruction is executed by the processor 701, the electronic device can execute the various steps in the above embodiments. Of course, the chip system can also include other discrete devices, which is not specifically limited in the embodiments of the present application.

[0161] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned electronic device, the electronic device executes the various functions or steps executed by the electronic device 100 in the above-mentioned method embodiment.

[0162] The present application also provides a computer program product, which, when executed on a computer, enables the computer to execute the functions or steps executed by the electronic device 100 in the above method embodiment. For example, the computer may be the above electronic device 100.

[0163] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0164] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0165] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0166] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0167] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0168] The above content is only a specific embodiment of this application, but the scope of protection of this application is not limited to this. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for searching duplicate files, characterized in that: include: Traversing the directories and files contained in the target directory, grouping the files based on the file size of each file, and obtaining at least one second file group; The remainders of the file sizes of at least two files in the second file group with respect to a preset first value are the same; Obtaining, from the at least one second file group, a first file group formed by files having the same file size; the first file group including at least two files having the same file size; when the file sizes of the at least two files having the same file size are greater than or equal to a preset threshold, target content check values ​​of the at least two files having the same file size are the same, and the target content check value is a cyclic check code of the target content; In the process of obtaining the first file group, hash values ​​of files included in the first file group are obtained in parallel through an asynchronous thread; wherein, if the hash value of a file of the first file group is included in the designated cache file and the hash value of the file has not been modified, the hash value of the file is obtained from the designated cache file through the asynchronous thread; if the hash value of a file of the first file group is not included in the designated cache file, or the hash value of the file in the designated cache file has been modified, the hash value of the file is calculated through the asynchronous thread; the storage format of the designated cache file is binary format; According to the hash values ​​of the files, files with the same hash values ​​are determined to be duplicate files.

2. The method according to claim 1, characterized in that The directories and files contained in the traversal target directory include: Traverse the directories and files contained in the target directory, add the directories in the form of a first data structure to the end of a preset global scan linked list in sequence according to the traversal order, and add the files in the form of a second data structure to the end of the preset global scan linked list in sequence according to the traversal order.

3. The method according to claim 1, characterized in that When the remainders of the file sizes of at least two files in the second file group with respect to a preset first value are the same, grouping the files based on the file sizes of the files to obtain at least one second file group includes: When adding the file to the preset global scan linked list in the second data structure, obtaining the file size of the file; Performing a remainder operation based on the file size and the preset first value to obtain a remainder corresponding to the file; At least one second file group formed by a plurality of files having the same remainder is obtained.

4. The method according to claim 3, characterized in that The preset first value is used to represent the number of red-black trees in the preset red-black forest; the root nodes of the red-black trees correspond to different remainders; The obtaining of at least one second file group formed by a plurality of files having the same remainder includes: Based on the remainder of the file, adding the file to the corresponding red-black tree; If a file already exists at the node of the red-black tree to which the file is to be added, the file is added to the second file group corresponding to the existing file.

5. The method according to any one of claims 1 to 4, characterized in that The step of obtaining the first file group formed by files having the same file size from the at least one second file group includes: Taking a file in the second file group as a first reference file, traversing the second file group to obtain a first matching file having the same file size as the reference file; If the file size of the first matching file is greater than or equal to the preset threshold, calculating a first checksum of the target content in the first reference file and a second checksum of the target content in the first matching file; the target content is used to represent a preset number of bytes of content in the file; If the first check value is the same as the second check value, the first data structure of the first matching file is added to the same-size linked list of the first reference file; the same-size linked list includes the first data structure of at least one first matching file with the same file size as the first reference file; the same-size linked list is the first file group.

6. The method according to claim 5, characterized in that The method further comprises: If the file size of the first matching file is smaller than the preset threshold, the first data structure of the first matching file is added to the same-size linked list of the first reference file.

7. The method according to any one of claims 1 to 4, characterized in that The step of determining, based on the hash values ​​of the files, that files with the same hash values ​​are duplicate files includes: Taking a file in the first file group as a second reference file, traversing the first file group to obtain a second matching file having the same hash value as the second reference file; Among them, the first data structure of the second matching file is stored in the same hash value linked list of the second reference file; the same hash value linked list includes the first data structure of at least one second matching file with the same hash value as the second reference file; the second reference file and the second matching file are duplicate files.

8. The method according to claim 7, characterized in that The method further comprises: The second reference file and the second matching file in the same hash value linked list of the second reference file are added to a preset duplicate file linked list.

9. The method according to claim 7, characterized in that The method further comprises: The preset global scan linked list is traversed in reverse order, and the first data structure of the file without calculated hash value and the second data structure of the empty directory are deleted from the preset global scan linked list to obtain an updated global scan linked list.

10. The method according to claim 9, characterized in that The method further comprises: Traversing the same hash value linked list and the updated global scan linked list of each second reference file, and storing first information of the directory and file in the updated global scan linked list into a preset first cache file; the first information includes an identifier, a parent directory number of the directory or file, and a directory name / file name; the identifier is used to indicate that the storage object is a file or directory; The hash value and modification time of the file in the updated global scan linked list are stored in a preset second cache file.

11. The method according to claim 10, characterized in that The obtaining hash values ​​of the files included in the first file group in parallel through asynchronous threads includes: If the second cache file does not exist, or the second cache file does not contain the hash value of the target file included in the first file group, or the first modification time of the target file in the second cache file is inconsistent with the second modification time of the target file in the first file group, calculating the hash value of the file included in the first file group by an asynchronous thread; If the second cache file exists, the second cache file contains the hash value of the target file included in the first file group, the first modification time of the target file in the second cache file is consistent with the second modification time of the target file in the first file group, and the hash value of the file corresponding to the file included in the first file group is obtained from the first cache file and the second cache file.

12. An electronic device, characterized in that: The electronic device includes a memory and one or more processors; the memory is coupled to the processor; computer program code is stored in the memory, and the computer program code includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the method as described in any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that The method comprises computer instructions, which, when executed on an electronic device, enable the electronic device to execute the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • File comparison method and device for HDFS (Hadoop Distributed File System)

    CN103377251A

  • Data compression method and device based on Cassandra database and storage medium

    CN110727685A

  • Data retrieval method and device

    CN111061680A

  • Algorithm for distributed asynchronous parallel detection of MPTCP protocol multipath attack

    CN113067819A

  • Harassment mark data processing method and device, electronic equipment and medium

    CN113722419A