File classification method and apparatus, electronic device, storage medium, and program product

Through file feature clustering and preset proportion recognition, the problems of high resource consumption and long time in massive file classification are solved, and efficient and accurate file classification is achieved.

WO2025138958A1PCT designated stage expired Publication Date: 2025-07-03HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/114918
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-08-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

When classifying massive files, the prior art consumes high resources and the identification task is completed for a long time, making it difficult to efficiently classify files.

Method used

By obtaining file characteristics for clustering, selecting files with preset proportions for content recognition, and determining category information when the similarity exceeds the threshold, reducing the need for file-by-file identification.

Benefits of technology

It improves the efficiency of file classification, reduces resource consumption and identification time, and ensures the accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024114918_03072025_PF_FP_ABST
    Figure CN2024114918_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of cloud computing. Provided are a file classification method and apparatus, an electronic device, a storage medium and a program product. The method comprises: acquiring first file features respectively corresponding to a plurality of files to be identified; on the basis of the first file features, clustering the plurality of said files to obtain files of at least one category; and selecting from the files of the same category files at a first preset proportion for content identification, so as to obtain a plurality of identification results; and if the similarity of the plurality of identification results exceeds a first preset threshold value, determining the plurality of identification results as category information of the files of the corresponding category. In the embodiments of the present disclosure, files at the first preset proportion are selected from files of the same category for content identification, if the similarity of the plurality of identification results exceeds the first preset threshold, the plurality of identification results are determined as the category information of the files of the corresponding category, such that there is no need to perform content identification on files to be identified one by one, thereby improving the efficiency of classifying files.
Need to check novelty before this filing date? Find Prior Art

Description

File classification method, device, electronic device, storage medium and program product

[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311870942.0 and application name “File classification method, device, electronic device, storage medium and program product”, the entire contents of which are incorporated by reference in this disclosure. Technical Field

[0002] The present disclosure relates to the field of data processing technology, and in particular to a file classification method, device, electronic device, storage medium, and program product. Background Art

[0003] With the promulgation and implementation of the "Data Security Law" and the "Personal Information Protection Law", there is a need to classify stored data and set different security levels for different categories of data to ensure data security.

[0004] The current approach to file classification is to identify each file individually. However, a storage product or system may store 100,000, 1,000,000, or even 100 million files. Identifying and reclassifying each file individually results in high resource consumption and a long recognition time. Therefore, a solution is needed that can quickly classify massive amounts of files.

[0005] Summary of the Invention

[0006] The embodiments of the present disclosure provide a file classification method, apparatus, electronic device, storage medium, and program product to improve the classification efficiency of massive files.

[0007] In a first aspect, an embodiment of the present disclosure provides a file classification method, the method comprising: obtaining first file features corresponding to a plurality of files to be identified respectively; clustering the plurality of files to be identified based on the first file features to obtain files of at least one category; selecting a first preset proportion of files from files of the same category for content recognition to obtain a plurality of recognition results; if the similarity of the plurality of recognition results exceeds a first preset threshold, determining the plurality of recognition results as category information of files of the corresponding category.

[0008] In a second aspect, an embodiment of the present disclosure provides a file classification device, which includes: an acquisition module for acquiring first file features corresponding to multiple files to be identified; a clustering module for clustering multiple files to be identified based on the first file features to obtain files of at least one category; an identification module for selecting a first preset proportion of files from files of the same category for content identification to obtain multiple identification results; and a determination module for determining multiple identification results as category information of files of corresponding categories if the similarity of the multiple identification results exceeds a first preset threshold.

[0009] In a third aspect, an embodiment of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the above methods when executing the computer program.

[0010] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above-mentioned methods is implemented.

[0011] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements any of the methods described above.

[0012] Compared with the prior art, the present disclosure has the following advantages:

[0013] The present disclosure provides a file classification method, device, electronic device, storage medium and program product. First, first file features corresponding to a plurality of files to be identified are obtained; second, based on the first file features, the plurality of files to be identified are clustered to obtain at least one category of files; then, a first preset proportion of files are selected from the files in the same category for content recognition to obtain a plurality of recognition results; if the similarity of the plurality of recognition results exceeds a first preset threshold, the plurality of recognition results are determined as category information of files of the corresponding category. In the embodiment of the present disclosure, a plurality of files to be identified are clustered based on the first file features, a first preset proportion of files are selected from the files in the same category for content recognition, and if the similarity of the plurality of recognition results exceeds a first preset threshold, the plurality of recognition results are determined as category information of files of the corresponding category, thereby eliminating the need to perform content recognition on the files to be identified one by one, which can reduce the time and resource consumption of file classification; moreover, recognition based on the clustered files can ensure recognition accuracy, thereby improving the efficiency of file classification.

[0014] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the specific implementation methods of the present disclosure are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present disclosure and should not be regarded as limiting the scope of the present disclosure.

[0016] FIG1 is a schematic diagram of a file directory in the related art.

[0017] FIG2 is a schematic diagram of an application scenario of the file classification method provided by the present disclosure.

[0018] FIG3 is a flowchart of a file classification method according to an embodiment of the present disclosure.

[0019] FIG4 is a flowchart of a file classification method according to an embodiment of the present disclosure.

[0020] FIG5 is a structural block diagram of a file classification device according to an embodiment of the present disclosure.

[0021] FIG6 is a block diagram of an electronic device for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present disclosure. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.

[0023] To facilitate understanding of the technical solutions of the embodiments of the present disclosure, the following describes the related technologies of the embodiments of the present disclosure. The following related technologies are optional solutions that can be combined with the technical solutions of the embodiments of the present disclosure in any way, and all of them fall within the scope of protection of the embodiments of the present disclosure.

[0024] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0025] Figure 1 is a schematic diagram of a file directory in the related art. Files are stored in a host or storage device, and the files are organized in the form of multi-level directories plus files. As shown in Figure 1, the top level of the file directory is the root directory, and there are a large number of subdirectories and files under the root directory, presenting a tree structure. The top is / "ROOT", which is the entrance to the file system. All subdirectories ( / BIN, / BOOT, / ETC, / USR... as shown in Figure 1) and files ("ABC" a...b...c... "DEF" d...e...f... as shown in Figure 1) are under / . / "ROOT" is the organizer of the file system and the top leader. In a data security scenario, when classifying files, all files are traversed and identified one by one according to the directory structure of the file system. Identifying massive amounts of data is a very time-consuming and resource-consuming task.

[0026] FIG2 is a schematic diagram of an application scenario of the file classification method provided by the present disclosure. First, a file is obtained. The file can be stored in any storage space, such as a host, a storage device, etc. The file can be of various types, such as image files, document files, audio files, video files, source code files, etc. Each file type includes a specific file format. Image file formats include bmp, jpg, png, tif, gif, pcx, tga, exif, fpx, svg, psd, cdr, pcd, dxf, ufo, eps, ai, raw, wmf, webp, avif, apng, etc.

[0027] Secondly, extract file features and cluster them. File features can include those extracted from any of the following dimensions: file size, file name, file extension, file storage time, file directory, file creator, file read and write permissions, or file metadata, etc. Files with the same file features are clustered into one category. For example, traversing to the / id / directory, it is found that there are 1 million files in this directory, and the file naming rules are all the name pinyin + storage time + file extension. For example, zhangsan_2023-12-01_11-06-23.jpg; lisi_2023-12-02_11-06-23.jpg; wangwu_2023-12-03_11-06-23.jpg, etc. These 1 million files are clustered into one category.

[0028] Again, select a preset proportion of files from the same category for content recognition. For example, identify the text, pictures, audio, video, etc. in the file to find out the sensitive personal information contained therein, such as ID number, email address, phone number, ID card picture, passport picture, etc. Different recognition methods can be used according to the different file types. For example, text-type files are generally parsed into text, and then keyword matching, regular expressions, dictionaries, machine learning algorithms, deep learning algorithms, etc. are used for content recognition. For picture-type files, image recognition models or optical character recognition (OCR) technology can be used to extract text for recognition.

[0029] The recognition results are then compared to determine the category. If the proportion of files with identical recognition results reaches a predetermined threshold, all files with the same characteristics are grouped into that category. If the proportion of files with identical recognition results does not reach the predetermined threshold, file features are extracted from the files in the current category and clustered to form multiple subcategories. Within each subcategory, a predetermined proportion of files are selected for content recognition, and the recognition results are compared to determine the subcategory's category information. This continues until all files have been classified.

[0030] Finally, set the security level. Set a security level for each categorized file. For example, randomly select 10,000 of 1 million images and perform content recognition on them. If all of them are ID card images, then all files in that category can be marked as such and set the security level to S4.

[0031] The present disclosure provides a method for file classification. The method in this embodiment can be applied to a computing device, which may include a server, a user terminal, etc. FIG3 is a flowchart of a method for file classification according to an embodiment of the present disclosure, including:

[0032] Step S301: obtaining first file features corresponding to a plurality of files to be identified.

[0033] Step S302: clustering the plurality of files to be identified based on the first file feature to obtain files of at least one category.

[0034] Step S303 : Selecting files of a first preset proportion from the files of the same category for content recognition to obtain a plurality of recognition results.

[0035] Step S304: If the similarity of the multiple recognition results exceeds a first preset threshold, the multiple recognition results are determined as category information of files of corresponding categories.

[0036] The file to be identified can be stored in any storage space, such as a host computer, a storage device, etc. The file can be of various types, such as image files, document files, audio files, video files, source code files, etc. When performing content identification on the file to be identified, different identification methods can be used according to different file types.

[0037] Exemplarily, the multiple files to be identified belong to the same subdirectory under the same root directory or different subdirectories under the same root directory; and the similarity of the naming rules corresponding to different subdirectories exceeds a third preset threshold.

[0038] The first file feature is a feature extracted based on information in at least one of the following dimensions: file size, file naming, file extension, file storage time, file directory, file creator, file read and write permissions, or file meta-information.

[0039] The category information may include a security level, or a corresponding security level may be set according to the category information.

[0040] The recognition results may include but are not limited to ID number, email address, phone number, ID card picture, passport picture, etc.

[0041] The disclosed embodiment provides a file classification method, first, obtaining first file features corresponding to a plurality of files to be identified respectively; second, clustering the plurality of files to be identified based on the first file features to obtain at least one category of files; then, selecting a first preset proportion of files from the same category for content recognition to obtain a plurality of recognition results; if the similarity of the plurality of recognition results exceeds a first preset threshold, the plurality of recognition results are determined as category information of files of the corresponding category. In the disclosed embodiment, clustering the plurality of files to be identified based on the first file features, selecting a first preset proportion of files from the same category for content recognition, if the similarity of the plurality of recognition results exceeds a first preset threshold, the plurality of recognition results are determined as category information of files of the corresponding category, thereby eliminating the need to perform content recognition on each of the files to be identified, which can reduce the time and resources consumed for file recognition; moreover, recognition based on the clustered files can ensure recognition accuracy, thereby improving the efficiency of file classification.

[0042] The following describes the specific implementation process of each of the above steps through various implementation methods:

[0043] In one implementation, the file classification method further includes: if the similarity of multiple recognition results does not exceed a first preset threshold, obtaining a second file feature of files in the same category; the second file feature has a different dimension from the first file feature; clustering the files in the same category based on the second file feature to obtain at least one subcategory of files; selecting a second preset proportion of files in the same subcategory for content recognition, and determining the category information of the corresponding subcategory files based on the recognition results of the files in the same subcategory.

[0044] In actual applications, after selecting a first preset ratio from files of the same category for identification, if the similarity of multiple identification results does not exceed a first preset threshold, a second file feature having a dimension different from the first file feature is extracted from the files of that category. The second file feature is a feature extracted based on information from at least one of the following dimensions: file size, file naming, file extension, file storage time, file directory, file creator, file read / write permissions, or file metadata. Based on the multiple second file features, the files of the same category are clustered again to obtain at least one subcategory of files; a second preset ratio of files are selected from the same subcategory for content identification. If the similarity of multiple identification results for files in the subcategory exceeds the first preset threshold, for example, the identification results are the same, the identification result is determined as the category information of the corresponding subcategory file. Otherwise, the files of the subcategory are clustered and identified again until all files are classified.

[0045] In one implementation, clustering multiple files to be identified based on a first file feature to obtain files of at least one category includes: determining multiple files to be identified whose similarity of the first file feature exceeds a second preset threshold as one category to obtain files of at least one category.

[0046] For example, if multiple files share the same naming convention, these files are clustered into one category. In this embodiment, file features are first extracted for clustering, and then a certain percentage of files are selected for content recognition. This avoids the need to fully recognize all files while ensuring classification accuracy, reducing recognition time and cost and improving classification efficiency.

[0047] In one implementation, a first preset proportion of files are selected from files of the same category for content recognition to obtain multiple recognition results, including: if the selected file type is an audio file, the audio file is converted into a text file, and content recognition is performed on the text file to obtain a recognition result; if the selected file type is a video file, the video file is converted into an image file, and content recognition is performed on the image file to obtain a recognition result.

[0048] Among them, if the selected file type is a video file, image frames can be collected at preset intervals and image recognition can be performed using an image recognition model. The image recognition model can be determined according to specific needs, for example, a face recognition model.

[0049] In one implementation, a first preset proportion of files are selected from files of the same category for content recognition to obtain multiple recognition results, including: if the selected file type is an image file, content recognition is performed on the image file to obtain a recognition result; or, the image file is converted into a text file, content recognition is performed on the text file to obtain a recognition result.

[0050] Among them, if the selected file type is an image file, you can use the image recognition model for image recognition, or use the OCR method to extract text for recognition. Text recognition can use keyword matching, regular expressions, dictionaries, machine learning algorithms, deep learning algorithms, etc.

[0051] The present disclosure provides a method for classifying files. The method in this embodiment can be applied to a computing device, which may include a server, a user terminal, etc. FIG4 is a flowchart of the method for classifying files according to an embodiment of the present disclosure, including:

[0052] Step S401: obtaining first file features corresponding to a plurality of files to be identified.

[0053] The multiple files to be identified belong to the same subdirectory under the same root directory or different subdirectories under the same root directory; and the similarity of the naming rules corresponding to different subdirectories exceeds a third preset threshold.

[0054] The files to be identified may be of various types, such as image files, document files, audio files, video files, source code files, etc.

[0055] The first file feature is a feature extracted based on information in at least one of the following dimensions: file size, file naming, file extension, file storage time, file directory, file creator, file read and write permissions, or file meta-information.

[0056] Step S402 : multiple files to be identified whose similarity of the first file feature exceeds a second preset threshold are determined as one category, thereby obtaining files of at least one category.

[0057] For example, if multiple files have the same naming rules, these files are grouped into one category.

[0058] Step S403 : Selecting a first preset proportion of files from the files of the same category to perform content recognition, and obtaining a plurality of recognition results.

[0059] Step S404: Determine whether the similarity of the multiple recognition results exceeds a first preset threshold.

[0060] Step S405 : If the similarity of the multiple recognition results exceeds a first preset threshold, the multiple recognition results are determined as category information of files of corresponding categories.

[0061] Step S406: If the similarity of the multiple recognition results does not exceed the first preset threshold, the second file feature of the file in the same category is obtained.

[0062] The second file feature has a different dimension from the first file feature. The second file feature is a feature extracted based on information in at least one of the following dimensions: file size, file name, file extension, file storage time, file directory, file creator, file read / write permissions, or file meta-information.

[0063] Step S407: cluster the files of the same category based on the second file feature to obtain files of at least one subcategory.

[0064] Step S408 : Select files of a second preset proportion from the files of the same subcategory to perform content recognition, and determine category information of the corresponding subcategory files based on the recognition results of the files of the same subcategory.

[0065] A second preset proportion of files are selected from the same subcategory files for content recognition. If the similarity of multiple recognition results of files in the subcategory exceeds a first preset threshold, for example, the recognition results are the same, the recognition results are determined as the category information of the corresponding subcategory files. Otherwise, the files of the subcategory are clustered and recognized again until all files are classified.

[0066] Corresponding to the application scenario and method of the method provided in the embodiment of the present disclosure, the embodiment of the present disclosure also provides a file classification device. As shown in Figure 5, a structural block diagram of the file classification device of an embodiment of the present disclosure is shown. The device includes:

[0067] The acquisition module 501 is configured to acquire first file features corresponding to a plurality of files to be identified.

[0068] The clustering module 502 is configured to cluster the plurality of files to be identified based on the first file feature to obtain files of at least one category.

[0069] The recognition module 503 is configured to select files of a first preset proportion from the files of the same category for content recognition to obtain a plurality of recognition results.

[0070] The determination module 504 is configured to determine the multiple recognition results as category information of files of corresponding categories if the similarity of the multiple recognition results exceeds a first preset threshold.

[0071] The disclosed embodiment provides a file classification device, which first obtains first file features corresponding to a plurality of files to be identified; secondly, clusters the plurality of files to be identified based on the first file features to obtain files of at least one category; then, selects a first preset proportion of files from the files of the same category to perform content recognition to obtain a plurality of recognition results; if the similarity of the plurality of recognition results exceeds a first preset threshold, the plurality of recognition results are determined as category information of files of the corresponding category. In the disclosed embodiment, the plurality of files to be identified are clustered based on the first file features, and selects a first preset proportion of files from the files of the same category to perform content recognition; if the similarity of the plurality of recognition results exceeds a first preset threshold, the plurality of recognition results are determined as category information of files of the corresponding category, thereby eliminating the need to perform content recognition on each of the files to be identified, which can reduce the time and resources consumed for file recognition; moreover, recognition based on the clustered files can ensure recognition accuracy, thereby improving the efficiency of file classification.

[0072] In one implementation, the determination module 504 is further used to: if the similarity of multiple recognition results does not exceed a first preset threshold, obtain a second file feature of files in the same category; the second file feature has a different dimension from the first file feature, clustering the files in the same category based on the second file feature to obtain at least one subcategory of files; select a second preset proportion of files in the same subcategory for content recognition, and determine the category information of the corresponding subcategory files based on the recognition results of the files in the same subcategory.

[0073] In one implementation, the clustering module 502 is configured to: determine a plurality of to-be-identified files whose similarity of the first file feature exceeds a second preset threshold as one category, and obtain files of at least one category.

[0074] In one implementation, the recognition module 503 is used to: if the selected file type is an audio file, convert the audio file into a text file, perform content recognition on the text file, and obtain a recognition result; if the selected file type is a video file, convert the video file into an image file, perform content recognition on the image file, and obtain a recognition result.

[0075] In one implementation, the recognition module 503 is configured to: if the selected file type is an image file, perform content recognition on the image file to obtain a recognition result; or convert the image file into a text file, perform content recognition on the text file to obtain a recognition result.

[0076] In one implementation, the first file feature or the second file feature is a feature extracted based on information of at least one of the following dimensions: file size, file name, file extension, file storage time, file directory, file creator, file read and write permissions, or file metadata.

[0077] In one implementation, the multiple files to be identified belong to the same subdirectory under the same root directory or different subdirectories under the same root directory; and the similarity of the naming rules corresponding to the different subdirectories exceeds a third preset threshold.

[0078] The functions of each module in the embodiment of the present disclosure can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0079] Figure 6 is a block diagram of an electronic device used to implement embodiments of the present disclosure. As shown in Figure 6 , the electronic device includes a memory 610 and a processor 620. The memory 610 stores a computer program that can be executed on the processor 620. When the processor 620 executes the computer program, the method described in the above embodiments is implemented. The number of memory 610 and processor 620 can be one or more.

[0080] The electronic device also includes:

[0081] The communication interface 630 is used to communicate with external devices and perform data exchange transmission.

[0082] If the memory 610, processor 620, and communication interface 630 are implemented independently, the memory 610, processor 620, and communication interface 630 can be interconnected via a bus and communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, FIG6 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.

[0083] Optionally, in a specific implementation, if the memory 610, the processor 620 and the communication interface 630 are integrated on a chip, the memory 610, the processor 620 and the communication interface 630 can communicate with each other through an internal interface.

[0084] An embodiment of the present disclosure provides a computer-readable storage medium storing a computer program, which implements the method provided in the embodiment of the present disclosure when the program is executed by a processor.

[0085] An embodiment of the present disclosure further provides a chip, which includes a processor for calling and executing instructions stored in a memory, so that a communication device equipped with the chip executes the method provided by the embodiment of the present disclosure.

[0086] An embodiment of the present disclosure also provides a chip, including: an input interface, an output interface, a processor and a memory. The input interface, the output interface, the processor and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.

[0087] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0088] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM) and direct RAM bus random access memory (DR RAM).

[0089] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present disclosure are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0090] In the description of this specification, reference to the terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and integrate different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0091] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.

[0092] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process. The scope of the preferred embodiments of the present disclosure includes additional implementations in which the functions may be performed in a different order than shown or discussed, including in a substantially simultaneous manner or in a reverse order depending on the functions involved.

[0093] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other system that can fetch instructions from an instruction execution system, apparatus or device and execute instructions), or used in combination with such instruction execution systems, apparatuses or devices.

[0094] It should be understood that various parts of the present disclosure may be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above-described method embodiments may be performed by a program instructing the relevant hardware. The program may be stored in a computer-readable storage medium. When executed, the program includes one or a combination of the steps of the method embodiments.

[0095] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the aforementioned integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0096] The above description is merely an exemplary embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope of the present disclosure, and such modifications or substitutions should be included within the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A method for classifying documents, wherein, The method includes: Obtaining first file features respectively corresponding to multiple files to be recognized; Clustering the multiple files to be recognized based on the first file features to obtain files of at least one category; Selecting files with a first preset ratio in the files of the same category for content recognition to obtain multiple recognition results; If the similarity of the multiple recognition results exceeds a first preset threshold, determining the multiple recognition results as the category information of the corresponding category files.

2. The method according to claim 1, wherein The method further includes: If the similarity of the multiple recognition results does not exceed the first preset threshold, obtaining second file features of the files of the same category; the second file features are different from the first file features in dimension; Clustering the files of the same category based on the second file features to obtain files of at least one sub-category; Selecting files with a second preset ratio in the files of the same sub-category for content recognition, and determining the category information of the corresponding sub-category files based on the recognition results of the files of the same sub-category.

3. The method according to claim 1 or 2, wherein The clustering the multiple files to be recognized based on the first file features to obtain files of at least one category includes: Determining multiple files to be recognized with the similarity of the first file features exceeding a second preset threshold as one category to obtain files of at least one category.

4. The method according to any one of claims 1-3, wherein, The selecting files with a first preset ratio in the files of the same category for content recognition to obtain multiple recognition results includes: If the selected file type is an audio file, converting the audio file into a text file, and performing content recognition on the text file to obtain a recognition result; If the selected file type is a video file, converting the video file into an image file, and performing content recognition on the image file to obtain a recognition result.

5. The method according to any one of claims 1-4, wherein, The selecting files with a first preset ratio in the files of the same category for content recognition to obtain multiple recognition results includes: If the selected file type is an image file, performing content recognition on the image file to obtain a recognition result; or converting the image file into a text file, and performing content recognition on the text file to obtain a recognition result.

6. The method according to any one of claims 1-5, wherein The first file features or the second file features are features extracted based on information of at least one of the following dimensions: File size, file naming, file extension, file storage time, file directory, file creator, file's Read and write permissions or file meta-information.

7. The method according to any one of claims 1-6, wherein, The multiple files to be recognized belong to the same sub-directory under the same root directory or different sub-directories under the same root directory; the similarity of the naming rules respectively corresponding to the different sub-directories exceeds a third preset threshold.

8. A document classification device, wherein, The device includes: An obtaining module, configured to obtain first file features respectively corresponding to multiple files to be recognized; A clustering module, configured to cluster the multiple files to be recognized based on the first file features to obtain files of at least one category; A recognition module, configured to select files with a first preset ratio in the files of the same category for content recognition to obtain multiple recognition results; A determining module, configured to, if the similarity of the multiple recognition results exceeds the first preset threshold, determine the multiple recognition results as the category information of the corresponding category files.

9. An electronic device, wherein, It includes a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, it implements the method according to any one of claims 1-7.

10. A computer-readable storage medium, wherein, A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, it implements the method according to any one of claims 1-7.

11. A computer program product, wherein, The computer program product includes a computer program. When the computer program is executed by a processor, it implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • File processing method and device, mobile terminal, and computer readable storage medium

    CN108363817A

  • Interactive voiceprint clustering method and system, electronic equipment and storage medium

    CN114596863A

  • Image file processing method and device, electronic equipment and storage medium

    CN116740732A

  • Information processing device, information processing method and program

    JP2018151705A