File classification method and device, electronic equipment, storage medium and program product
Through file feature clustering and preset proportion recognition, the problems of high resource consumption and long time in massive file classification are solved, and efficient and accurate file classification is achieved.
Patent Information
- Application Number
- CN202311870942.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
AI Technical Summary
When classifying massive files, the prior art consumes a high resource consumption and the identification task is completed for a long time, so it is impossible to quickly and effectively classify files.
By obtaining file characteristics for clustering, selecting files with preset proportions for content recognition, and determining category information when the similarity exceeds the threshold, reducing the identification steps of files by file.
It improves the efficiency of file classification, reduces resource consumption and identification time, and ensures the accuracy of identification.
Smart Images

Figure CN120277032A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a file classification method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] With the promulgation and implementation of the Data Security Law and the Personal Information Protection Law, there is a need to classify the stored data, and different security levels are set for different categories of data to ensure data security.
[0003] When classifying files, the current method is to identify each file one by one. However, in a storage product or system, there may be 100,000-level, 1,000,000-level, or even 100,000,000-level files. The method of identifying and classifying each file one by one will result in high resource consumption and a long time to complete the identification task. Therefore, a solution that can quickly classify a large number of files is needed. Summary of the Invention
[0004] Embodiments of this application provide a file classification method, apparatus, electronic device, storage medium, and program product to improve the classification efficiency of a large number of files.
[0005] In a first aspect, embodiments of this application provide a file classification method, which includes: obtaining first file features respectively corresponding to a plurality of files to be identified; clustering the plurality of files to be identified based on the first file features to obtain files of at least one category; selecting a first preset proportion of files in the files of the same category for content identification to obtain a plurality of identification results; and if the similarity of the plurality of identification results exceeds a first preset threshold, determining the plurality of identification results as the category information of the corresponding category of files.
[0006] In a second aspect, embodiments of this application provide a file classification apparatus, which includes: an obtaining module, configured to obtain first file features respectively corresponding to a plurality of files to be identified; a clustering module, configured to cluster the plurality of files to be identified based on the first file features to obtain files of at least one category; an identification module, configured to select a first preset proportion of files in the files of the same category for content identification to obtain a plurality of identification results; and a determining module, configured to if the similarity of the plurality of identification results exceeds a first preset threshold, determine the plurality of identification results as the category information of the corresponding category of files.
[0007] In a third aspect, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor implements the method according to any one of the above when executing the computer program.
[0008] Fourthly, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the above is implemented.
[0009] Fifthly, an embodiment of the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the method described in any one of the above is implemented.
[0010] Compared with the prior art, the present application has the following advantages:
[0011] The present application provides a file classification method, device, electronic device, storage medium and program product. First, obtain the first file features respectively corresponding to a plurality of files to be recognized; secondly, cluster the plurality of files to be recognized based on the first file features to obtain at least one category of files; then, select a first preset proportion of files in the files of the same category for content recognition to obtain a plurality of recognition results; if the similarity of the plurality of recognition results exceeds a first preset threshold, determine the plurality of recognition results as the category information of the corresponding category of files. In the embodiment of the present application, cluster the plurality of files to be recognized based on the first file features, select a first preset proportion of files in the files of the same category for content recognition, and if the similarity of the plurality of recognition results exceeds a first preset threshold, determine the plurality of recognition results as the category information of the corresponding category of files, so that it is not necessary to perform content recognition on each file to be recognized one by one, which can reduce the time and resources consumed for file classification; moreover, based on the recognized files after clustering, the recognition accuracy can be guaranteed, thereby improving the efficiency of file classification.
[0012] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. Description of the Drawings
[0013] In the drawings, unless otherwise specified, the same reference numerals throughout the drawings denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.
[0014] Figure 1 It is a schematic diagram of a file directory in the related art.
[0015] Figure 2 It is a schematic diagram of an application scenario of the file classification method provided by the present application.
[0016] Figure 3 Flowchart of a file classification method according to an embodiment of the present application.
[0017] Figure 4 Flowchart of a file classification method according to an embodiment of the present application.
[0018] Figure 5 Structural block diagram of a file classification device according to an embodiment of the present application.
[0019] Figure 6 Block diagram of an electronic device for implementing the embodiment of the present application. Detailed implementation manners
[0020] In the following, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature and not restrictive.
[0021] To facilitate the understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.
[0022] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0023] Figure 1 Schematic diagram of a file directory in the related art. Files are stored in a host or a storage device, and the files are organized in a multi-level directory plus file manner. For example Figure 1 shown, the top layer of the file directory is the root directory. There are a large number of subdirectories and files under the root directory, presenting a tree structure. The topmost is " / "ROOT", which is the entrance of the file system. All subdirectories (such as Figure 1 shown / BIN, / BOOT, / ETC, / USR...), files (such as Figure 1 shown "ABC"a...b...c..."DEF"d...e...f...) are all under " / ". " / "ROOT" is the organizer of the file system and also the top-level leader. In the data security scenario, when classifying files, all files are traversed and identified one by one according to the directory structure of the file system. Identifying a large amount of data is a very time-consuming and resource-consuming task.
[0024] Figure 2 This is a schematic diagram of an application scenario of the file classification method provided for this application. First, obtain files. The files can be stored in any storage space, such as a host, a storage device, etc. The files can be of various types, such as picture files, document files, audio files, video files, source code files, etc. Each file type includes specific file formats. The picture file formats include bmp, jpg, png, tif, gif, pcx, tga, exif, fpx, svg, psd, cdr, pcd, dxf, ufo, eps, ai, raw, WMF, webp, avif, apng, etc.
[0025] Secondly, extract file features and cluster them. File features can be extracted from any of the following dimensions: file size, file naming, file extension, file storage time, file directory, file creator, file read and write permissions, or file meta-information, etc. Files with the same file features are clustered into one category. For example, when traversing the / id / directory, it is found that the number of files in this directory is 1 million, and the file naming rule is the pinyin of the name + storage time + file extension. For example, zhangsan_2023-12-01_11-06-23.jpg; lisi_2023-12-02_11-06-23.jpg; wangwu_2023-12-03_11-06-23.jpg, etc. These 1 million files are clustered into one category.
[0026] Thirdly, select a preset proportion of files in the same category for content recognition. For example, recognize the text, pictures, audio, video, etc. in the files, and find out the sensitive personal information contained therein, such as ID card numbers, email addresses, phone numbers, ID card pictures, passport pictures, etc. Specifically, different recognition methods can be adopted according to the different file types. For example, for text-type files, after generally parsing them into text, keyword matching, regular expressions, dictionaries, machine learning algorithms, deep learning algorithms, etc. are used for content recognition. For picture-type files, an image recognition model or Optical Character Recognition (OCR) technology can be used to extract text for recognition.
[0027] Then, compare the recognition results to determine the category information. If the proportion of files with the same selected recognition result reaches a predetermined value, all files with the same characteristics can be classified into this category. If the proportion of files with the same recognition result does not reach the predetermined value, extract the file characteristics of the current category of files, perform clustering to obtain multiple subcategories, then select a preset proportion of files in each subcategory for content recognition, and compare the recognition results to determine the category information of the subcategories... until all file classifications are completed.
[0028] Finally, set the security level. Set the security level for each category of files after classification. For example, among 1 million pictures, randomly select 10,000 of them for content recognition. If it is finally found that all such files are ID card pictures, then all files in this category can be marked as ID card pictures and the security level can be set to S4.
[0029] The embodiment of the present application provides a file classification method. The method in this embodiment can be applied to a computing device, and the computing device may include: a server, a user terminal, etc. As Figure 3 shown is a flowchart of the file classification method according to an embodiment of the present application, including:
[0030] Step S301, obtain the first file characteristics respectively corresponding to multiple files to be recognized.
[0031] Step S302, perform clustering on the multiple files to be recognized based on the first file characteristics to obtain at least one category of files.
[0032] Step S303, select a first preset proportion of files in the same category of files for content recognition to obtain multiple recognition results.
[0033] Step S304, if the similarity of the multiple recognition results exceeds a first preset threshold, determine the multiple recognition results as the category information of the corresponding category of files.
[0034] Among them, the files to be recognized can be stored in any storage space, for example, a host, a storage device, etc. Files can be of various types, for example, picture files, document files, audio files, video files, source code files, etc. When performing content recognition on the files to be recognized, different recognition methods can be specifically adopted according to the different file types.
[0035] Exemplarily, multiple files to be recognized belong to the same subdirectory under the same root directory or different subdirectories under the same root directory; the similarity of the naming rules respectively corresponding to different subdirectories exceeds a third preset threshold.
[0036] Among them, the first file feature is a feature extracted based on information of at least one of the following dimensions: file size, file naming, file extension, file storage time, file directory, file creator, file read and write permissions, or file meta information.
[0037] Among them, the category information may include a security level, or a corresponding security level is set according to the category information.
[0038] The recognition results may include, but are not limited to, ID card numbers, email addresses, phone numbers, ID card pictures, passport pictures, etc.
[0039] The embodiment of the present application provides a file classification method. First, obtain the first file features respectively corresponding to multiple files to be recognized; secondly, cluster the multiple files to be recognized based on the first file features to obtain files of at least one category; then, select files with a first preset ratio in the files of the same category for content recognition to obtain multiple recognition results; if the similarity of the multiple recognition results exceeds a first preset threshold, determine the multiple recognition results as the category information of the corresponding category files. In the embodiment of the present application, cluster the multiple files to be recognized based on the first file features, select files with a first preset ratio in the files of the same category for content recognition, and if the similarity of the multiple recognition results exceeds a first preset threshold, determine the multiple recognition results as the category information of the corresponding category files, so that it is not necessary to perform content recognition on each file to be recognized, which can reduce the time and resources consumed for file recognition; moreover, based on the clustered files for recognition, the recognition accuracy can be guaranteed, thereby improving the efficiency of file classification.
[0040] The following introduces the specific implementation process of each of the above steps through multiple implementation manners:
[0041] In one implementation manner, the file classification method further includes: if the similarity of the multiple recognition results does not exceed a first preset threshold, obtain the second file features of the files of the same category; the dimensions of the second file features are different from those of the first file features; cluster the files of the same category based on the second file features to obtain files of at least one subcategory; select files with a second preset ratio in the files of the same subcategory for content recognition, and determine the category information of the corresponding subcategory files based on the recognition results of the files of the same subcategory.
[0042] In practical applications, after selecting a first preset proportion of files from the same category for recognition, if the similarity of multiple recognition results does not exceed the first preset threshold, then a second file feature different from the dimension of the first file feature is extracted from the files in this category. The second file feature is a feature extracted based on the information of at least one of the following dimensions: file size, file name, file extension, file storage time, file directory, file creator, read / write permission of the file, or meta-information of the file. Cluster the files in the same category again based on multiple second file features to obtain at least one sub-category of files; select a second preset proportion of files from the files in the same sub-category for content recognition. If the similarity of multiple recognition results of the files in this sub-category exceeds the first preset threshold, for example, the recognition results are the same, then determine the recognition result as the category information of the corresponding sub-category of files; otherwise, cluster and recognize the files in this sub-category again until all file classifications are completed.
[0043] In one implementation, clustering multiple files to be recognized based on the first file feature to obtain at least one category of files includes: determining multiple files to be recognized with the similarity of the first file feature exceeding a second preset threshold as one category to obtain at least one category of files.
[0044] For example, if the file name naming rules of multiple files are the same, then cluster these files into one category. In this embodiment, first extract file features for clustering, and then select a certain proportion of files for content recognition. While ensuring the classification accuracy, it avoids full-scale recognition of all files, reduces the recognition time and cost, and improves the classification efficiency.
[0045] In one implementation, selecting a first preset proportion of files from the files in the same category for content recognition to obtain multiple recognition results includes: if the selected file type is an audio file, convert the audio file into a text file and perform content recognition on the text file to obtain the recognition result; if the selected file type is a video file, convert the video file into an image file and perform content recognition on the image file to obtain the recognition result.
[0046] Among them, if the selected file type is a video file, image frames can be collected at preset time intervals and image recognition can be performed using an image recognition model. The image recognition model can be determined according to specific needs. For example, a face recognition model, etc.
[0047] In one implementation, selecting a first preset proportion of files from the files in the same category for content recognition to obtain multiple recognition results includes: if the selected file type is an image file, perform content recognition on the image file to obtain the recognition result; or convert the image file into a text file and perform content recognition on the text file to obtain the recognition result.
[0048] Among them, if the selected file type is an image file, an image recognition model can be used for image recognition, or text can be extracted and recognized by means of OCR. For text recognition, keyword matching, regular expressions, dictionaries, machine learning algorithms, deep learning algorithms, etc. can be used.
[0049] The embodiments of the present application provide a file classification method. The method in this embodiment can be applied to a computing device, and the computing device may include: a server, a user terminal, etc. As Figure 4 shown in the flowchart of the file classification method according to an embodiment of the present application, it includes:
[0050] Step S401, obtain first file features respectively corresponding to a plurality of files to be recognized.
[0051] Among them, the plurality of files to be recognized belong to the same sub-directory under the same root directory or different sub-directories under the same root directory; the similarity of the naming rules respectively corresponding to different sub-directories exceeds a third preset threshold.
[0052] Among them, the files to be recognized can be of various types. For example, picture files, document files, audio files, video files, source code files, etc.
[0053] Among them, the first file feature is a feature extracted based on information in at least one of the following dimensions: file size, file naming, file extension, file storage time, file directory, file creator, file read and write permissions, or file meta-information.
[0054] Step S402, determine multiple files to be recognized whose similarity of the first file features exceeds a second preset threshold as one category, and obtain files of at least one category.
[0055] For example, if the naming rules of the file names of multiple files are the same, then these files are grouped into one category.
[0056] Step S403, select a first preset proportion of files in the files of the same category for content recognition to obtain multiple recognition results.
[0057] Step S404, determine whether the similarity of the multiple recognition results exceeds a first preset threshold.
[0058] Step S405, if the similarity of the multiple recognition results exceeds the first preset threshold, then determine the multiple recognition results as the category information of the corresponding category of files.
[0059] Step S406, if the similarity of the multiple recognition results does not exceed the first preset threshold, then obtain second file features of the files of the same category.
[0060] The dimensions of the second file feature are different from those of the first file feature. The second file feature is a feature extracted based on information of at least one of the following dimensions: file size, file naming, file extension, file storage time, file directory, file creator, file read-write permission, or file meta-information.
[0061] Step S407: Cluster files of the same category based on the second file feature to obtain at least one sub-category of files.
[0062] Step S408: Select files with a second preset ratio in the files of the same sub-category for content recognition, and determine the category information of the corresponding sub-category of files based on the recognition results of the files in the same sub-category.
[0063] Select files with a second preset ratio in the files of the same sub-category for content recognition. If the similarity of multiple recognition results of the files in this sub-category exceeds a first preset threshold, for example, the recognition results are the same, then determine the recognition results as the category information of the corresponding sub-category of files; otherwise, cluster and recognize the files of this sub-category again until all file classifications are completed.
[0064] Corresponding to the application scenario and method of the method provided in the embodiments of the present application, the embodiments of the present application also provide a file classification device. As Figure 5 shown is a structural block diagram of a file classification device according to an embodiment of the present application. The device includes:
[0065] An acquisition module 501, configured to acquire first file features respectively corresponding to multiple files to be recognized.
[0066] A clustering module 502, configured to cluster multiple files to be recognized based on the first file feature to obtain at least one category of files.
[0067] An identification module 503, configured to select files with a first preset ratio in the files of the same category for content recognition to obtain multiple recognition results.
[0068] A determination module 504, configured to, if the similarity of multiple recognition results exceeds a first preset threshold, determine the multiple recognition results as the category information of the corresponding category of files.
[0069] An embodiment of the present application provides a file classification device. First, obtain first file features corresponding to multiple files to be recognized respectively; secondly, cluster the multiple files to be recognized based on the first file features to obtain files of at least one category; then, select files with a first preset ratio in the files of the same category for content recognition to obtain multiple recognition results; if the similarity of the multiple recognition results exceeds a first preset threshold, determine the multiple recognition results as the category information of the corresponding category of files. In the embodiment of the present application, cluster the multiple files to be recognized based on the first file features, select files with a first preset ratio in the files of the same category for content recognition, and if the similarity of the multiple recognition results exceeds a first preset threshold, determine the multiple recognition results as the category information of the corresponding category of files, so that it is not necessary to perform content recognition on each file to be recognized one by one, which can reduce the time and resources consumed for file recognition; moreover, based on the clustered files for recognition, the recognition accuracy can be guaranteed, thereby improving the efficiency of file classification.
[0070] In one implementation, the determining module 504 is further configured to: if the similarity of the multiple recognition results does not exceed the first preset threshold, obtain second file features of the files of the same category; the second file features are different from the first file features in dimension, cluster the files of the same category based on the second file features to obtain files of at least one subcategory; select files with a second preset ratio in the files of the same subcategory for content recognition, and determine the category information of the corresponding subcategory of files based on the recognition results of the files of the same subcategory.
[0071] In one implementation, the clustering module 502 is configured to: determine multiple files to be recognized with a similarity of the first file features exceeding a second preset threshold as one category to obtain files of at least one category.
[0072] In one implementation, the recognition module 503 is configured to: if the selected file type is an audio file, convert the audio file into a text file, perform content recognition on the text file to obtain a recognition result; if the selected file type is a video file, convert the video file into an image file, perform content recognition on the image file to obtain a recognition result.
[0073] In one implementation, the recognition module 503 is configured to: if the selected file type is an image file, perform content recognition on the image file to obtain a recognition result; or convert the image file into a text file, perform content recognition on the text file to obtain a recognition result.
[0074] In one implementation, the first file features or the second file features are features extracted based on information of at least one of the following dimensions: file size, file name, file extension, file storage time, file directory, file creator, file read and write permissions, or file meta information.
[0075] In one implementation, multiple files to be recognized belong to the same sub-directory under the same root directory or different sub-directories under the same root directory; the similarity of the naming rules corresponding to different sub-directories exceeds a third preset threshold.
[0076] For the functions of each module in the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, and they have corresponding beneficial effects, which will not be elaborated here.
[0077] Figure 6 It is a block diagram of an electronic device for implementing the embodiments of the present application. As Figure 6 shown, the electronic device includes: a memory 610 and a processor 620. The memory 610 stores a computer program that can run on the processor 620. When the processor 620 executes the computer program, the method in the above embodiments is implemented. The number of the memory 610 and the processor 620 can be one or more.
[0078] The electronic device further includes:
[0079] A communication interface 630, configured to communicate with external devices and perform data interaction and transmission.
[0080] If the memory 610, the processor 620, and the communication interface 630 are implemented independently, the memory 610, the processor 620, and the communication interface 630 can be interconnected through a bus and complete communication with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0081] Optionally, in a specific implementation, if the memory 610, the processor 620, and the communication interface 630 are integrated on a chip, the memory 610, the processor 620, and the communication interface 630 can complete communication with each other through an internal interface.
[0082] The embodiments of the present application provide a computer-readable storage medium, which stores a computer program that, when executed by a processor, implements the method provided in the embodiments of the present application.
[0083] An embodiment of the present application further provides a chip, which includes a processor for calling and running instructions stored in a memory, so that a communication device installed with the chip executes the method provided by the embodiment of the present application.
[0084] An embodiment of the present application further provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is configured to execute code in the memory. When the code is executed, the processor is configured to execute the method provided by the embodiment of the application.
[0085] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the advanced risc machines (ARM) architecture.
[0086] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0087] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0088] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0089] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.
[0090] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed.
[0091] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in connection with these instruction execution systems, apparatus, or devices.
[0092] It should be understood that the various parts of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware. This program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiment.
[0093] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.
[0094] As described above, only the exemplary embodiments of the present application are provided, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope recorded in the present application can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for classifying documents, characterized in that, The method includes: Obtaining first file features respectively corresponding to multiple files to be recognized; Clustering the multiple files to be recognized based on the first file features to obtain files of at least one category; Selecting files with a first preset proportion in files of the same category for content recognition to obtain multiple recognition results; If the similarity of the multiple recognition results exceeds a first preset threshold, determining the multiple recognition results as the category information of the corresponding category files.
2. The method according to claim 1, wherein The method further includes: If the similarity of the multiple recognition results does not exceed the first preset threshold, obtaining second file features of the files of the same category; the second file features are different from the first file features in dimension; Clustering the files of the same category based on the second file features to obtain files of at least one sub-category; Selecting files with a second preset proportion in files of the same sub-category for content recognition, and determining the category information of the corresponding sub-category files based on the recognition results of the files of the same sub-category.
3. The method according to claim 1, wherein The clustering the multiple files to be recognized based on the first file features to obtain files of at least one category includes: Determining multiple files to be recognized with the similarity of the first file features exceeding a second preset threshold as one category to obtain files of at least one category.
4. The method according to any one of claims 1 to 3, characterized in that, The selecting files with a first preset proportion in files of the same category for content recognition to obtain multiple recognition results includes: If the selected file type is an audio file, converting the audio file into a text file, and performing content recognition on the text file to obtain a recognition result; If the selected file type is a video file, converting the video file into an image file, and performing content recognition on the image file to obtain a recognition result.
5. The method according to any one of claims 1 to 3, characterized in that, The selecting files with a first preset proportion in files of the same category for content recognition to obtain multiple recognition results includes: If the selected file type is an image file, performing content recognition on the image file to obtain a recognition result; or converting the image file into a text file, and performing content recognition on the text file to obtain a recognition result.
6. The method according to any one of claims 1-3, characterized in that The first file features or the second file features are features extracted based on information of at least one of the following dimensions: File size, file naming, file extension, file storage time, file directory, file creator, read and write permissions of the file, or meta information of the file.
7. The method according to any one of claims 1 to 3, characterized in that, The multiple files to be recognized belong to the same sub-directory under the same root directory or different sub-directories under the same root directory; the similarity of the naming rules respectively corresponding to the different sub-directories exceeds a third preset threshold.
8. A file classification device, characterized in that, The device includes: An obtaining module, configured to obtain first file features respectively corresponding to multiple files to be recognized; A clustering module, configured to cluster the multiple files to be recognized based on the first file features to obtain files of at least one category; An identifying module, configured to select files with a first preset proportion in files of the same category for content recognition to obtain multiple recognition results; A determining module, configured to, if the similarity of the multiple recognition results exceeds a first preset threshold, determine the multiple recognition results as the category information of the corresponding category files.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program, it implements the method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.
11. A computer program product, characterized in that, The computer program product includes a computer program. When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.