Repeated file detection method and device, storage medium and computer equipment
By classifying and grouping files, and extracting specific and general feature vectors, the problems of poor accuracy and high computational cost in the prior art are solved, and efficient and accurate repeated file detection is achieved.
Patent Information
- Application Number
- CN202411915100.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has poor accuracy when detecting and managing duplicate files, and is very computationally consuming, and cannot effectively process non-text files.
By first classifying according to file type and grouping according to file size, and then extracting features of specific feature dimensions and common feature dimensions for each target file group, the target feature vector of the file is finally obtained, and the repetition of the file is detected through the target feature vector.
It effectively improves the accuracy and efficiency of repeated file detection and can cope with repeated detection of files of different file types.
Smart Images

Figure CN119938608A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a duplicate file detection method and device, a storage medium, and a computer device. Background Art
[0002] In the digital age, with the rapid development of information technology, people have generated a large number of electronic files in their work and life. These files may include documents, pictures, audio, video and other forms, and they play a vital role in different application scenarios. However, with the sharp increase in the number of files, file management and storage face huge challenges, among which the detection and management of duplicate files has become an urgent problem to be solved.
[0003] Duplicate files refer to files with the same or highly similar content. They may be generated for a variety of reasons, such as file backup, version update, misoperation, etc. In the file system, the existence of duplicate files not only wastes valuable storage space, but also may reduce the efficiency of file retrieval and even cause data confusion and security issues. Therefore, how to effectively detect and manage duplicate files has become an important topic in the field of file management and storage.
[0004] Traditional duplicate file detection methods mainly rely on the comparison of basic information such as file name, file size, and file type. However, this method has great limitations and poor accuracy. In order to improve accuracy, some improved methods try to compare the file content word by word. However, when faced with large files or a large number of files to be processed, the amount of calculation will increase exponentially, and the time complexity is extremely high, resulting in an extremely time-consuming processing process, which is almost impractical in practical applications. In addition, non-text files cannot be compared using the above method. Summary of the invention
[0005] In view of this, the present application provides a duplicate file detection method and device, storage medium, and computer equipment, which first classifies files according to file type and groups them according to file size, and then extracts features of specific feature dimensions and common feature dimensions for files in each target file group, and finally obtains the target feature vector of the file. The target feature vector is used to detect the duplication of files, which effectively improves the accuracy and efficiency of duplicate file detection, and can cope with duplicate detection of files of different file types.
[0006] According to one aspect of the present application, a method for detecting duplicate files is provided, comprising:
[0007] In response to a duplicate file detection instruction, a plurality of files to be detected for duplicates are obtained, and the plurality of files are classified according to file types, and the files under each classification are grouped according to file sizes to obtain a plurality of file groups under each classification;
[0008] For each target file group including multiple files, determine the specific feature dimension corresponding to the target file group according to the file type corresponding to the target file group, and perform feature extraction on each file in the target file group according to the specific feature dimension and the general feature dimension to obtain a target feature vector corresponding to each file, and determine whether there are duplicate files in the target file group according to the target feature vector;
[0009] Output the duplicate file detection results under each category.
[0010] According to another aspect of the present application, a duplicate file detection device is provided, comprising:
[0011] A grouping module, configured to obtain a plurality of files to be detected for duplication in response to a duplicate file detection instruction, and classify the plurality of files according to file types, and group the files under each category according to file sizes to obtain a plurality of file groups under each category;
[0012] A duplicate file detection module is used to determine, for each target file group containing multiple files, a specific feature dimension corresponding to the target file group according to the file type corresponding to the target file group, and perform feature extraction on each file in the target file group according to the specific feature dimension and the general feature dimension to obtain a target feature vector corresponding to each file, and determine whether there is a duplicate file in the target file group according to the target feature vector;
[0013] The detection result output module is used to output the duplicate file detection results under each category.
[0014] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned duplicate file detection method is implemented.
[0015] According to another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the above-mentioned duplicate file detection method when executing the program.
[0016] Through the above technical scheme, the present application provides a duplicate file detection method and device, storage medium, and computer equipment, which first classify the files according to the file type and group them according to the file size, and then extract the features of the files in each target file group in terms of specific feature dimensions and general feature dimensions, and finally obtain the target feature vector of the file. The file duplication is detected by the target feature vector, which effectively improves the accuracy and efficiency of duplicate file detection, and can cope with duplicate detection of files of different file types.
[0017] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0019] Figure 1 A schematic diagram of a flow chart of a duplicate file detection method provided in an embodiment of the present application is shown;
[0020] Figure 2 A schematic diagram showing a flow chart of another duplicate file detection method provided in an embodiment of the present application;
[0021] Figure 3 A schematic diagram of the structure of a duplicate file detection device provided in an embodiment of the present application is shown;
[0022] Figure 4 A schematic diagram of the device structure of a computer device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0023] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict.
[0024] In this embodiment, a method for detecting duplicate files is provided. Figure 1 As shown, the method includes:
[0025] Step 101, in response to a duplicate file detection instruction, multiple files to be detected for duplicates are obtained, and the multiple files are classified according to file types, and the files in each classification are grouped according to file sizes to obtain multiple file groups in each classification.
[0026] A method for detecting duplicate files provided by an embodiment of the present application can be used on the client side or on the server side, and can accurately identify duplicate files in a storage space. The following is described using a computer device as an example, and the same technical effect can be achieved by applying the same method and steps to the server side. First, when a user wants to identify duplicate files in a computer device, a detection instruction for duplicate files can be generated by triggering a detection button for duplicate files, after which the computer device can respond to the detection instruction and identify all files that need to be detected repeatedly according to the detection instruction. In another embodiment, the detection instruction for duplicate files can also be regularly generated by a computer device, and specifically, can be implemented by writing an automated script.
[0027] After obtaining the files that need to be detected for duplication, the files can be further grouped according to the file type to perform duplication detection on the premise of the same file type. Here, the file type can be a text file, an image file, an audio file, a video file, etc. Since it is meaningless to compare files of different file types, the files are first classified by file type to obtain files of each file type.
[0028] Within each file type classification, the files are further divided into smaller groups based on file size. This can optimize the detection process because files of similar size may be more similar in content, while files of greatly different sizes are almost never duplicate files, which can greatly reduce the amount of calculation. It should be noted that when grouping by file size, it can be ensured that the size difference between every two files in the group is less than the preset threshold. In addition, in order to improve the accuracy of duplicate file recognition, the same file can be divided into multiple groups. For example, for text files, the preset threshold is 2KB, file 1 is 109KB, file 2 is 110KB, file 3 is 112KB, and file 4 is 111KB. When grouping, file 1, file 2, and file 4 can be grouped into one group, and file 2, file 3, and file 4 can be grouped into another group.
[0029] Step 102, for each target file group containing multiple files, determine the specific feature dimension corresponding to the target file group according to the file type corresponding to the target file group, and perform feature extraction on each file in the target file group according to the specific feature dimension and the general feature dimension to obtain a target feature vector corresponding to each file, and determine whether there are duplicate files in the target file group according to the target feature vector.
[0030] In this embodiment, after the files under each file type are grouped to obtain multiple file groups, further, if there is only one file in a certain file group, then there is no need to perform subsequent operations on the file, and it can be directly determined that the file has no corresponding duplicate files. If there are multiple files in a certain file group, it means that there may be duplicate files in these file groups, and this file group is called a target file group. At this time, the specific feature dimensions corresponding to the group can be determined according to the file type corresponding to this target file group. These specific features are features obtained for the attributes or structures unique to this type of file. For example, for image files, specific features can be image semantic features, image resolution features, etc.; for text files, specific features can be text semantic features, text paragraph features, etc. In addition to the specific feature dimensions corresponding to each file type, some general feature dimensions can also be considered. These features are applicable to all types of files, such as metadata such as file creation time and modification time.
[0031] Afterwards, based on the above feature dimensions, feature extraction is performed on the files in each target file group to generate a target feature vector. This vector is a representation of the file in a specific feature space and is used for subsequent comparison. Subsequently, by comparing the target feature vectors of each file in the same target file group, it is determined whether they are similar enough to be considered as duplicate files. Specifically, the distance or similarity score between different target feature vectors can be calculated, and a judgment can be made based on a preset similarity.
[0032] Step 103: output the duplicate file detection results under each category.
[0033] In this embodiment, finally, the duplicate file detection results under each category are summarized and outputted. The output content may include the specific information of the duplicate files (such as file name, path), the number of duplicate files, etc.
[0034] By applying the technical solution of the present embodiment, the embodiment of the present application first classifies files by file type and groups them by file size, and then extracts features of specific feature dimensions and common feature dimensions for the files in each target file group, and finally obtains the target feature vector of the file. The file duplication detection is performed through the target feature vector, which effectively improves the accuracy and efficiency of duplicate file detection and can cope with duplicate detection of files of different file types.
[0035] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, another duplicate file detection method is provided, such as Figure 2 As shown, the method includes:
[0036] Step 201, in response to a duplicate file detection instruction, obtain multiple files to be detected for duplicates, classify the multiple files according to file types, group the files in each category according to file sizes, and obtain multiple file groups in each category.
[0037] Step 202 : for each target file group including multiple files, determine a specific feature dimension corresponding to the target file group according to a file type corresponding to the target file group.
[0038] Step 203: extract specific features from each file in the target file group according to the specific feature dimension to obtain a first feature vector corresponding to the file.
[0039] Step 204 : extract common features from each file in the target file group according to the common feature dimension to obtain a second feature vector corresponding to the file.
[0040] Step 205: fusing the first feature vector and the second feature vector to obtain a target feature vector corresponding to each file.
[0041] In this embodiment, for each file in the target file group, after selecting the corresponding specific feature dimension according to the file type (such as text, image, audio, etc.), these specific feature dimensions are used to extract features of the file to obtain the feature vector corresponding to each specific feature dimension, and these feature vectors are fused together to obtain the first feature vector corresponding to the file. This first feature vector is the representation of the file in the specific feature space, which captures the attributes or structure unique to the file type.
[0042] Next, common features are extracted for each file. Common feature dimensions are feature dimensions applicable to all types of files, such as file size, creation time, modification time, etc. After extracting the feature vector corresponding to each common feature, these feature vectors are fused together to obtain the second feature vector corresponding to the file. This second feature vector provides some basic information about the file.
[0043] Finally, the target feature vector corresponding to each file is obtained according to the first feature vector and the second feature vector. Specifically, the first feature vector and the second feature vector can be concatenated to obtain the target feature vector. The target feature vector is the final representation of the file in the comprehensive feature space, which combines the information of specific features and general features, and improves the accuracy and robustness of duplicate file detection.
[0044] Step 206: According to the file type corresponding to the target file group, determine a target repetition rate calculation model that matches the file type of the target file group from a preset repetition rate calculation model set; input the target feature vectors corresponding to every two files in the target file group into the target repetition rate calculation model respectively, calculate the feature similarity between every two files through the target repetition rate calculation model, and determine two files whose feature similarity is greater than a preset similarity as duplicate files.
[0045] In this embodiment, different file types generally use different repetition rate calculation models when calculating the repetition rate. Therefore, after obtaining the target feature vector of each file, further, when calculating the repetition rate of the files in each target file group, a target repetition rate calculation model that matches the file type of the target file group can be selected from a preset repetition rate calculation model set according to the file type corresponding to the target file group. For example, text files can use SVM (Support Vector Machine, support vector machine), Naive Bayes, LSTM (Long Short-Term Memory, long short-term memory network), CNN (Convolutional Neural Network, i.e., Convolutional Neural Network) models as target repetition rate calculation models; non-text files can select a suitable model according to the file type, for example, video files can use MLP (Multi-Layer Perceptron) and other multimodal models. In the embodiment of the present application, the target repetition rate calculation model is matched according to the file type, which can ensure that the selected target repetition rate calculation model can accurately process and understand the characteristics of the type of file, thereby improving the accuracy of repetition detection. In addition, when training the repetition rate calculation model, a large number of labeled file samples can be used, with the target feature vector extracted from the file sample as input and the repeated mark as output to train the model. Optimization algorithms such as SGD (Stochastic Gradient Descent) and Adam are used to adjust the parameters, and cross-validation and leave-out validation are used to prevent overfitting. The output of the model is a binary classification result (duplicate / non-duplicate) or feature similarity. When the feature similarity exceeds the preset similarity, it is determined to be a duplicate file.
[0046] Next, the target feature vectors corresponding to every two files in the target file group are input into the selected target repetition rate calculation model. The model can calculate the feature similarity between every two files based on these target feature vectors. Feature similarity is a quantitative indicator used to measure the proximity or similarity of two files in the feature space. The method for calculating feature similarity can include calculating the Euclidean distance, cosine similarity, Manhattan distance, etc. between the target feature vectors, which is not limited here.
[0047] Finally, the calculated feature similarity is compared with the preset similarity. If the feature similarity of two files is greater than the preset similarity, they are determined to be duplicate files. Here, a unified preset similarity can be set for different file types. In addition, different preset similarities can also be set for different file types. For example, for certain types of files (such as text files), a higher preset similarity can be set to ensure the accuracy of detection; while for other types of files (such as audio or video files), since the feature space may be more complex and diverse, a slightly lower preset similarity can be set to capture potential duplications. The embodiment of the present application selects a matching repetition rate calculation model and sets a reasonable preset similarity according to the file type. This method can adapt to different types of files and different detection requirements.
[0048] Step 207, outputting duplicate file detection results under each category.
[0049] In an embodiment of the present application, optionally, when the file type is a text file, the specific feature dimensions include text semantic features, text paragraph features, and text sentence structure features; step 203 includes: for each file in the target file group, obtaining a first text semantic feature vector corresponding to the file through a bag of words model, and obtaining a second text semantic feature vector corresponding to the file through a TF-IDF feature extraction method, and obtaining a third text semantic feature vector corresponding to the file through a word embedding model; performing feature fusion on the first text semantic feature vector, the second text semantic feature vector, and the third text semantic feature vector to obtain a text semantic feature vector corresponding to the file; determining the file bag The number of paragraphs contained in the file and the number of characters contained in each paragraph are taken as vector elements to obtain a text paragraph feature vector corresponding to the file; named entity recognition is performed on each sentence in the file to obtain a target entity contained in the sentence, and a semantic role is assigned to each target entity based on the sentence, and a sentence structure feature vector corresponding to the sentence is obtained according to the semantic role corresponding to each target entity, and a text sentence structure feature vector corresponding to the file is obtained according to the sentence structure feature vector corresponding to each sentence; the text semantic feature vector, the text paragraph feature vector and the text sentence structure feature vector are vector-fused to obtain a first feature vector corresponding to the file.
[0050] In this embodiment, if the file type is a text file, the corresponding specific feature dimensions may include text semantic features, text paragraph features, text sentence structure features, and the like.
[0051] First, text semantic feature extraction. When extracting the text semantic features of a file, the Bag of Words (BOW), TF-IDF feature extraction method, and Word Embedding model can be used to extract features respectively. Specifically, the Bag of Words model can regard the text as an unordered set of words, ignoring the grammar and word order of the text, and only focusing on the frequency of occurrence of words. The Bag of Words model can form a word frequency matrix by counting the number of occurrences of words in each file, and then generate the first text semantic feature vector based on the word frequency matrix. For example, all words are encoded, and for the words that appear in the text file, the encoding of each word can be determined, and the word frequency of the word is followed after the encoding, and finally the first text semantic feature vector is formed in the form of (word encoding 1 word frequency 1 word encoding 2 word frequency 2...).
[0052] In the TF-IDF feature extraction method, TF (Term Frequency) indicates the frequency of occurrence of a word in a file, and IDF (Inverse Document Frequency) indicates the importance of a word in the entire file set (i.e., inverse document frequency). The TF-IDF feature extraction method can calculate the TF-IDF value of each word in each file to form a TF-IDF matrix, and then generate a second text semantic feature vector based on the TF-IDF matrix. For example, the above-mentioned vocabulary encoding can be used. For the words that appear in the text file, the encoding of each word can be determined, and the TF-IDF value of the word is followed by the encoding. Finally, the second text semantic feature vector is formed in the form of (vocabulary encoding 1 TF-IDF value 1 vocabulary encoding 2 TF-IDF value 2...).
[0053] The word embedding model can map words into a continuous, low-dimensional vector space, so that similar words are close to each other in the vector space. Specifically, you can use a pre-trained word embedding model (such as Word2Vec, GloVe, etc.) to convert the words in the file into vectors, and perform operations such as averaging or weighted summation to obtain the third text semantic feature vector.
[0054] Finally, the above three text semantic feature vectors can be fused (such as splicing, etc.) to obtain the final text semantic feature vector. It should be noted that the dimensions of the first text semantic feature vector, the second text semantic feature vector and the third text semantic feature vector between different files are the same, which is convenient for subsequent comparison. When generating these three text semantic feature vectors, the dimension of each text semantic feature vector can be predetermined, and then the corresponding text semantic feature vector is generated according to the predetermined dimension. Specifically, if the dimension of the directly generated text semantic feature vector is smaller than the predetermined dimension, then the dimension can be increased; if the dimension of the directly generated text semantic feature vector is larger than the predetermined dimension, then the dimension can be reduced.
[0055] Second, text paragraph feature extraction. Specifically, the number of paragraphs contained in the file can be counted first, and then the number of characters contained in each paragraph can be counted. Then, the number of paragraphs and the number of characters in each paragraph are used as vector elements to form a text paragraph feature vector. For example, a text paragraph feature vector can be formed in the form of (paragraph 1 number of characters 1 paragraph 2 number of characters 2...). Similarly, the dimensions of text paragraph feature vectors between different files are also the same, and the same dimensions can be ensured by dimensionality increase and dimensionality reduction.
[0056] Third, text sentence structure feature extraction. Specifically, first, the target entity (such as name, place name, organization name, etc.) in each sentence can be identified through Named Entity Recognition (NER) technology. Then, a semantic role (such as agent, patient, etc.) is assigned to each target entity through Semantic Role Labeling (SRL). Subsequently, a sentence structure feature vector is constructed according to the semantic roles corresponding to each target entity. For example, a sentence structure feature vector can be formed in the form of (target entity 1 semantic role 1 target entity 2 semantic role 2...). Finally, the above processing is performed on each sentence in the file, and based on the sentence structure feature vectors corresponding to each sentence, splicing and other operations are performed to obtain the text sentence structure feature vector corresponding to the file. Similarly, the dimensions of the text sentence structure feature vectors between different files are also the same, and the same dimensions can be ensured by dimensionality increase and dimensionality reduction.
[0057] It should be noted that the above processes of text semantic feature extraction, text paragraph feature extraction, and text sentence structure feature extraction can be performed simultaneously.
[0058] Fourth, feature vector fusion. That is, the text semantic feature vector, the text paragraph feature vector and the text sentence structure feature vector are fused (such as splicing, etc.) to obtain the first feature vector corresponding to the file. Since the dimensions of the text semantic feature vectors, text paragraph feature vectors and text sentence structure feature vectors between different files are required to be the same, the dimensions of the final first feature vector are also the same. The embodiment of the present application comprehensively captures the semantics, paragraph and sentence structure information of the text file through multi-dimensional feature extraction and fusion, which helps to improve the accuracy and efficiency of text duplication detection.
[0059] In an embodiment of the present application, optionally, when the file type is an image file, the specific feature dimension includes image semantic features and image resolution features; step 203 includes: for each file in the target file group, inputting the file into the input layer of a pre-trained convolutional neural network, extracting the image semantic features of the file through the convolutional layer and the pooling layer of the convolutional neural network, and performing weighted processing on the image semantic features obtained by the convolutional layer and the pooling layer through a fully connected layer to obtain an image semantic feature vector corresponding to the file; obtaining the horizontal pixel number and the vertical pixel number corresponding to the file, obtaining the resolution corresponding to the file according to the product of the horizontal pixel number and the vertical pixel number, and performing vector conversion on the resolution to obtain the image resolution feature vector corresponding to the file; performing vector fusion on the image semantic feature vector and the image resolution feature vector to obtain a first feature vector corresponding to the file.
[0060] In this embodiment, if the file type is an image file, the corresponding specific feature dimensions may include image semantic features, image resolution features, etc. Image semantic features refer to high-level semantic information such as objects, scenes, actions, etc. contained in the image. This information is crucial for understanding the content of the image and performing tasks such as image classification and recognition; image resolution features refer to the size of the image, usually expressed in terms of horizontal and vertical pixel numbers. Resolution is a basic attribute of an image, which affects the clarity and detail of the image.
[0061] First, image semantic feature extraction. Use a pre-trained convolutional neural network (CNN) to extract image semantic features. CNN is a deep learning model that is particularly suitable for image feature extraction. The pre-trained CNN model has been trained on a large-scale image dataset and can learn rich image features. CNN can include an input layer, a convolution layer, a pooling layer, a fully connected layer, and an output layer. First, the image file is input as input to the input layer of CNN. After that, the image is convolved and pooled through the convolution layer and pooling layer of CNN to extract local and global features of the image. These features are usually represented as high-dimensional vectors. Further, CNN inputs the feature vectors obtained by the convolution layer and the pooling layer into the fully connected layer for weighted processing. The fully connected layer can further combine and transform the features to obtain a more expressive image semantic feature vector. Here, the output of the fully connected layer is the image semantic feature vector.
[0062] Second, image resolution feature extraction. First, obtain the horizontal and vertical pixel counts of the image, and then multiply the horizontal and vertical pixel counts to obtain the image resolution (usually in pixels). Among them, the horizontal and vertical pixel counts of the image are the basic information of the image resolution, which can be obtained through image processing libraries (such as OpenCV, PIL, etc.). Subsequently, the image resolution value is converted into a vector form. Specifically, the resolution value can be put into a one-dimensional vector as an element, or the resolution value can be split into two components, horizontal and vertical, and put into two elements of a two-dimensional vector respectively.
[0063] It should be noted that the above processes of image semantic feature extraction and image resolution feature extraction can be performed simultaneously.
[0064] Third, vector fusion. The image semantic feature vector and the image resolution feature vector are fused (such as splicing, etc.) to obtain a comprehensive feature vector containing image semantics and resolution information. This comprehensive feature vector is the first feature vector corresponding to the image file. Since the dimensions of the image semantic feature vectors and image resolution feature vectors between different files are required to be the same, the dimensions of the final first feature vector are also the same. The embodiment of the present application combines the image semantic features and the image resolution features to obtain the first feature vector corresponding to each image file. The first feature vector integrates the image semantic information and image resolution information of the image file, which helps to improve the accuracy and efficiency of image duplication detection.
[0065] In an embodiment of the present application, optionally, when the file type is an audio file, the specific feature dimensions include spectrum features, duration features, and sampling rate features; step 203 includes: for each file in the target file group, converting the audio in the file from the time domain to the frequency domain by a fast Fourier transform method to obtain converted spectrum data, and inputting the converted frequency domain data into a filter group, calculating the output energy of each filter in the filter group, performing a logarithmic operation on each of the output energies to obtain a logarithmic energy value, performing a discrete cosine transform on the logarithmic energy value to obtain a Mel-frequency cepstral coefficient, and obtaining a spectrum feature vector corresponding to the file based on the Mel-frequency cepstral coefficient; determining the duration of the file, and obtaining a duration feature vector corresponding to the file based on the duration; obtaining the sampling rate of the file, and obtaining a sampling rate feature vector corresponding to the file based on the sampling rate; performing vector fusion on the spectrum feature vector, the duration feature vector, and the sampling rate feature vector to obtain a first feature vector corresponding to the file.
[0066] In this embodiment, if the file type is an audio file, the corresponding specific feature dimensions may include spectrum features, duration features, sampling rate features, etc. Among them, the spectrum feature reflects the energy distribution of the audio signal at different frequencies and is an important feature in audio analysis; the duration feature refers to the duration of the audio file, that is, the duration of the audio signal, which is a basic attribute of the audio; the sampling rate feature refers to the number of samples extracted from the continuous signal per second to form a discrete signal, which determines the resolution of the audio signal and the highest frequency that can be represented.
[0067] First, spectrum feature extraction. First, the audio signal is sampled and quantized to convert it into a digital signal. In order to increase the energy of the high-frequency part, the audio signal is usually pre-emphasized. After that, since the audio signal changes over time, it is divided into multiple short-time frames for processing. Windowing can also be performed. Here, windowing is to reduce the discontinuity between frames. Use fast Fourier transform (FFT) for each frame signal to convert the audio signal from the time domain to the frequency domain to obtain spectrum data. FFT is an efficient method for calculating discrete Fourier transform (DFT) and its inverse transform. Then, the converted frequency domain data is input into the filter bank, and the output energy of each filter can be calculated. Here, the filter bank consists of a series of bandpass filters, each of which corresponds to a specific frequency range. After that, the output energy of each filter is logarithmically operated to obtain the logarithmic energy value, which can reduce the dynamic range of the energy value and make it more suitable for subsequent processing. Subsequently, the logarithmic energy value is discrete cosine transformed to obtain Mel frequency cepstral coefficients (MFCC). MFCC is a feature widely used in audio processing, which can well represent the spectral characteristics of audio signals. Finally, a series of MFCC coefficients can be obtained (that is, each frame corresponds to an MFCC coefficient), which reflect the energy distribution of the audio signal at different frequencies. By arranging these coefficients in a certain order, a high-dimensional vector can be obtained, which is the spectral feature vector corresponding to the audio file.
[0068] Second, duration feature extraction. First, the duration of the audio file can be determined by reading the metadata of the audio file or using a dedicated audio processing library. Then, the duration value is converted into a vector form. Specifically, the duration value can be put into a one-dimensional vector as an element to obtain a duration feature vector.
[0069] Third, sampling rate feature extraction. First, the sampling rate of the audio file can be obtained by reading the metadata of the audio file or using a dedicated audio processing library. Then, the sampling rate is converted into a vector form. This can also be done by putting the sampling rate value as an element into a one-dimensional vector to obtain the sampling rate feature vector.
[0070] It should be noted that the above-mentioned processes of spectrum feature extraction, duration feature extraction, and sampling rate feature extraction can be performed simultaneously.
[0071] Fourth, vector fusion. The spectrum feature vector, duration feature vector and sampling rate feature vector are fused. This usually involves concatenating the three vectors and other operations to obtain a comprehensive feature vector containing all the feature information of the audio file. This comprehensive feature vector is the first feature vector corresponding to the audio file. Among them, since the dimensions of the spectrum feature vectors, duration feature vectors and sampling rate feature vectors between different files are required to be the same, the dimensions of the final first feature vector are also the same. The embodiment of the present application combines spectrum features, duration features and sampling rate features to ultimately obtain the first feature vector corresponding to each audio file. The first feature vector contains various information about the audio file, which helps to improve the accuracy and efficiency of audio duplication detection.
[0072] In addition, the file type can also be a video file. When the file type is a video file, the video file can be split into an audio file and an image file for processing. Specifically, the video can be processed frame by frame to obtain a series of images, and then each image is processed in accordance with the aforementioned processing method of the image file to obtain the first feature vector corresponding to each frame of the image, and the first feature vectors corresponding to each frame of the image are spliced to obtain the image feature vector of the video file; then, the audio therein is processed in accordance with the aforementioned processing method of the audio file to obtain the audio feature vector corresponding to the audio part, and the image feature vector and the audio feature vector are spliced to finally obtain the first feature vector of the video file. It should be noted that in order to reduce the amount of calculation, the feature vector obtained at each stage can be subjected to dimensionality reduction processing.
[0073] In an embodiment of the present application, optionally, step 204 includes: for each file in the target file group, obtaining the file creation time, file modification time and file author corresponding to the file, and obtaining a creation time vector, a modification time vector and an author vector according to the file creation time, the file modification time and the file author, respectively, and performing vector fusion on the creation time vector, the modification time vector and the author vector to obtain a second feature vector corresponding to the file.
[0074] In this embodiment, for each file in the target file group, three key common features of the file are extracted: file creation time, file modification time, and file author. The file creation time refers to the timestamp or date when the file was initially created; the file modification time refers to the timestamp or date when the file was last modified; and the file author refers to the creator or last modifier of the file, which is usually a user name or identifier.
[0075] Next, we can use the file creation time, file modification time, and file author to generate a creation time vector, a modification time vector, and an author vector, respectively. The creation time vector can be obtained by converting the file creation time into a vector. For example, a timestamp or date is converted into a numerical representation, and then converted into a vector using a method such as time serialization or one-hot encoding. Similar to the creation time vector, the modification time of the file is also converted into a vector to obtain the modification time vector. The author of the file is converted into a vector to obtain the author vector. Specifically, the author name can be converted into a numerical representation, such as using word embedding technology (such as Word2Vec) to map the author name into a high-dimensional vector space.
[0076] Afterwards, the above three vectors (creation time vector, modification time vector, author vector) can be fused (such as splicing operation) to generate a comprehensive feature vector, i.e., the second feature vector. For each file in the target file group, the above steps are repeated to finally obtain the second feature vector corresponding to each file. The embodiment of the present application forms the second feature vector through common features such as file creation time, file modification time, and file author, so that the second feature vector can include common feature information under multiple dimensions, which helps to improve the accuracy of subsequent file repeatability recognition.
[0077] Further, as Figure 1 The specific implementation of the method, the embodiment of the present application provides a detection device for duplicate files, such as Figure 3 As shown, the device comprises:
[0078] A grouping module, configured to obtain a plurality of files to be detected for duplication in response to a duplicate file detection instruction, and classify the plurality of files according to file types, and group the files under each category according to file sizes to obtain a plurality of file groups under each category;
[0079] A duplicate file detection module is used to determine, for each target file group containing multiple files, a specific feature dimension corresponding to the target file group according to the file type corresponding to the target file group, and perform feature extraction on each file in the target file group according to the specific feature dimension and the general feature dimension to obtain a target feature vector corresponding to each file, and determine whether there is a duplicate file in the target file group according to the target feature vector;
[0080] The detection result output module is used to output the duplicate file detection results under each category.
[0081] Optionally, the duplicate file detection module is used to:
[0082] Extracting specific features from each file in the target file group according to the specific feature dimension to obtain a first feature vector corresponding to the file;
[0083] Extracting common features from each file in the target file group according to the common feature dimension to obtain a second feature vector corresponding to the file;
[0084] The first feature vector and the second feature vector are fused to obtain a target feature vector corresponding to each file.
[0085] Optionally, when the file type is a text file, the specific feature dimensions include text semantic features, text paragraph features, and text sentence structure features; and the duplicate file detection module is further used to:
[0086] For each file in the target file group, a first text semantic feature vector corresponding to the file is obtained by using a bag-of-words model, a second text semantic feature vector corresponding to the file is obtained by using a TF-IDF feature extraction method, and a third text semantic feature vector corresponding to the file is obtained by using a word embedding model; the first text semantic feature vector, the second text semantic feature vector and the third text semantic feature vector are subjected to feature fusion to obtain a text semantic feature vector corresponding to the file;
[0087] Determine the number of paragraphs contained in the file and the number of characters contained in each paragraph, use the number of paragraphs and the number of characters contained in each paragraph as vector elements, and obtain a text paragraph feature vector corresponding to the file;
[0088] Perform named entity recognition on each sentence in the file to obtain target entities contained in the sentence, and assign a semantic role to each target entity based on the sentence, obtain a sentence structure feature vector corresponding to the sentence according to the semantic role corresponding to each target entity, and obtain a text sentence structure feature vector corresponding to the file according to the sentence structure feature vector corresponding to each sentence;
[0089] The text semantic feature vector, the text paragraph feature vector and the text sentence structure feature vector are vector-fused to obtain a first feature vector corresponding to the file.
[0090] Optionally, when the file type is an image file, the specific feature dimension includes an image semantic feature and an image resolution feature; and the duplicate file detection module is further used to:
[0091] For each file in the target file group, the file is input into the input layer of the pre-trained convolutional neural network, the image semantic features of the file are extracted through the convolutional layer and the pooling layer of the convolutional neural network, and the image semantic features obtained by the convolutional layer and the pooling layer are weighted through the fully connected layer to obtain the image semantic feature vector corresponding to the file;
[0092] Obtaining the number of horizontal pixels and the number of vertical pixels corresponding to the file, obtaining the resolution corresponding to the file according to the product of the number of horizontal pixels and the number of vertical pixels, and performing vector conversion on the resolution to obtain an image resolution feature vector corresponding to the file;
[0093] The image semantic feature vector and the image resolution feature vector are fused to obtain a first feature vector corresponding to the file.
[0094] Optionally, when the file type is an audio file, the specific feature dimensions include spectrum features, duration features, and sampling rate features; and the duplicate file detection module is further used to:
[0095] For each file in the target file group, convert the audio in the file from the time domain to the frequency domain by a fast Fourier transform method to obtain converted spectrum data, input the converted frequency domain data into a filter bank, calculate the output energy of each filter in the filter bank, perform a logarithmic operation on each output energy to obtain a logarithmic energy value, perform a discrete cosine transform on the logarithmic energy value to obtain a Mel-frequency cepstral coefficient, and obtain a spectrum feature vector corresponding to the file based on the Mel-frequency cepstral coefficient;
[0096] Determine the duration of the file, and obtain a duration feature vector corresponding to the file based on the duration;
[0097] Obtaining a sampling rate of the file, and obtaining a sampling rate feature vector corresponding to the file according to the sampling rate;
[0098] The frequency spectrum feature vector, the duration feature vector and the sampling rate feature vector are fused to obtain a first feature vector corresponding to the file.
[0099] Optionally, the duplicate file detection module is further used to:
[0100] For each file in the target file group, the file creation time, file modification time and file author corresponding to the file are obtained, and a creation time vector, a modification time vector and an author vector are obtained respectively according to the file creation time, the file modification time and the file author, and the creation time vector, the modification time vector and the author vector are vector-fused to obtain a second feature vector corresponding to the file.
[0101] Optionally, the duplicate file detection module is further used to:
[0102] According to the file type corresponding to the target file group, determining a target repetition rate calculation model matching the file type of the target file group from a preset repetition rate calculation model set;
[0103] The target feature vectors corresponding to every two files in the target file group are respectively input into the target repetition rate calculation model, the feature similarity between every two files is calculated by the target repetition rate calculation model, and two files whose feature similarity is greater than a preset similarity are determined as duplicate files.
[0104] It should be noted that for other corresponding descriptions of the functional units involved in the duplicate file detection device provided in the embodiment of the present application, reference can be made to Figure 1 to Figure 2 The corresponding description in the method will not be repeated here.
[0105] The present application also provides a computer device, which may be a personal computer, a server, a network device, etc. Figure 4 As shown, the computer device includes a bus, a processor, a memory and a communication interface, and may also include an input and output interface and a display device. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps in each method embodiment are implemented.
[0106] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0107] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium may be non-volatile or volatile, and stores a computer program thereon. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0108] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0109] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0110] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.
[0111] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0112] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for detecting duplicate files, characterized in that: include: In response to a duplicate file detection instruction, a plurality of files to be detected for duplicates are obtained, and the plurality of files are classified according to file types, and the files under each classification are grouped according to file sizes to obtain a plurality of file groups under each classification; For each target file group including multiple files, determine the specific feature dimension corresponding to the target file group according to the file type corresponding to the target file group, and perform feature extraction on each file in the target file group according to the specific feature dimension and the general feature dimension to obtain a target feature vector corresponding to each file, and determine whether there are duplicate files in the target file group according to the target feature vector; Output the duplicate file detection results under each category.
2. The method according to claim 1, characterized in that The step of extracting features from each file in the target file group according to the specific feature dimension and the general feature dimension to obtain a target feature vector corresponding to each file includes: Extracting specific features from each file in the target file group according to the specific feature dimension to obtain a first feature vector corresponding to the file; Extracting common features from each file in the target file group according to the common feature dimension to obtain a second feature vector corresponding to the file; The first feature vector and the second feature vector are fused to obtain a target feature vector corresponding to each file.
3. The method according to claim 2, characterized in that When the file type is a text file, the specific feature dimension includes text semantic features, text paragraph features, and text sentence structure features; and the specific feature extraction is performed on each file in the target file group according to the specific feature dimension to obtain a first feature vector corresponding to the file, including: For each file in the target file group, a first text semantic feature vector corresponding to the file is obtained by using a bag-of-words model, a second text semantic feature vector corresponding to the file is obtained by using a TF-IDF feature extraction method, and a third text semantic feature vector corresponding to the file is obtained by using a word embedding model; the first text semantic feature vector, the second text semantic feature vector and the third text semantic feature vector are subjected to feature fusion to obtain a text semantic feature vector corresponding to the file; Determine the number of paragraphs contained in the file and the number of characters contained in each paragraph, use the number of paragraphs and the number of characters contained in each paragraph as vector elements, and obtain a text paragraph feature vector corresponding to the file; Perform named entity recognition on each sentence in the file to obtain target entities contained in the sentence, and assign a semantic role to each target entity based on the sentence, obtain a sentence structure feature vector corresponding to the sentence according to the semantic role corresponding to each target entity, and obtain a text sentence structure feature vector corresponding to the file according to the sentence structure feature vector corresponding to each sentence; The text semantic feature vector, the text paragraph feature vector and the text sentence structure feature vector are vector-fused to obtain a first feature vector corresponding to the file.
4. The method according to claim 2, characterized in that: When the file type is an image file, the specific feature dimension includes an image semantic feature and an image resolution feature; and extracting specific features from each file in the target file group according to the specific feature dimension to obtain a first feature vector corresponding to the file includes: For each file in the target file group, the file is input into the input layer of the pre-trained convolutional neural network, the image semantic features of the file are extracted through the convolutional layer and the pooling layer of the convolutional neural network, and the image semantic features obtained by the convolutional layer and the pooling layer are weighted through the fully connected layer to obtain the image semantic feature vector corresponding to the file; Obtaining the number of horizontal pixels and the number of vertical pixels corresponding to the file, obtaining the resolution corresponding to the file according to the product of the number of horizontal pixels and the number of vertical pixels, and performing vector conversion on the resolution to obtain an image resolution feature vector corresponding to the file; The image semantic feature vector and the image resolution feature vector are fused to obtain a first feature vector corresponding to the file.
5. The method according to claim 2, characterized in that: When the file type is an audio file, the specific feature dimensions include spectrum features, duration features, and sampling rate features; and the specific feature extraction is performed on each file in the target file group according to the specific feature dimensions to obtain a first feature vector corresponding to the file, including: For each file in the target file group, convert the audio in the file from the time domain to the frequency domain by a fast Fourier transform method to obtain converted spectrum data, input the converted frequency domain data into a filter bank, calculate the output energy of each filter in the filter bank, perform a logarithmic operation on each output energy to obtain a logarithmic energy value, perform a discrete cosine transform on the logarithmic energy value to obtain a Mel-frequency cepstral coefficient, and obtain a spectrum feature vector corresponding to the file based on the Mel-frequency cepstral coefficient; Determine the duration of the file, and obtain a duration feature vector corresponding to the file based on the duration; Obtaining a sampling rate of the file, and obtaining a sampling rate feature vector corresponding to the file according to the sampling rate; The frequency spectrum feature vector, the duration feature vector and the sampling rate feature vector are fused to obtain a first feature vector corresponding to the file.
6. The method according to claim 2, characterized in that The step of extracting common features from each file in the target file group according to the common feature dimension to obtain a second feature vector corresponding to the file includes: For each file in the target file group, the file creation time, file modification time and file author corresponding to the file are obtained, and a creation time vector, a modification time vector and an author vector are obtained respectively according to the file creation time, the file modification time and the file author, and the creation time vector, the modification time vector and the author vector are vector-fused to obtain a second feature vector corresponding to the file.
7. The method according to claim 1, characterized in that The determining, according to the target feature vector, whether there are duplicate files in the target file group comprises: According to the file type corresponding to the target file group, determining a target repetition rate calculation model matching the file type of the target file group from a preset repetition rate calculation model set; The target feature vectors corresponding to every two files in the target file group are respectively input into the target repetition rate calculation model, the feature similarity between every two files is calculated by the target repetition rate calculation model, and two files whose feature similarity is greater than a preset similarity are determined as duplicate files.
8. A duplicate file detection device, characterized in that: include: A grouping module, configured to obtain a plurality of files to be detected for duplication in response to a duplicate file detection instruction, and classify the plurality of files according to file types, and group the files under each category according to file sizes to obtain a plurality of file groups under each category; A duplicate file detection module is used to determine, for each target file group containing multiple files, a specific feature dimension corresponding to the target file group according to the file type corresponding to the target file group, and perform feature extraction on each file in the target file group according to the specific feature dimension and the general feature dimension to obtain a target feature vector corresponding to each file, and determine whether there is a duplicate file in the target file group according to the target feature vector; The detection result output module is used to output the duplicate file detection results under each category.
9. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.