File classification method and device, storage medium and program product

By extracting the feature information of cloud storage files and using a classification model to generate multi-level hierarchical labels, the problem of cumbersome cloud storage file management operations is solved, and refined classification and efficient search are achieved.

CN122019770APending Publication Date: 2026-05-12UC MOBILE CHINA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UC MOBILE CHINA CO LTD
Filing Date
2025-12-08
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing cloud storage products suffer from cumbersome and inefficient file management, especially when faced with diverse data types, making it difficult to achieve refined classification.

Method used

By extracting the feature information of files stored in the cloud drive, a classification model is used to generate classification labels with at least two levels, including type information, image information, path information, metadata information, etc., to achieve fine-grained classification and search of files.

Benefits of technology

It simplifies the operation of user file management, improves the efficiency of file management, and enables fine-grained classification and fast searching of files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019770A_ABST
    Figure CN122019770A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a file classification method and device, a storage medium and a program product. The method comprises the following steps: extracting feature information of a first file stored in a network disk; determining a classification label of the first file by utilizing a classification model according to the feature information of the first file; the classification labels are hierarchical structure labels and comprise hierarchical labels with at least two levels of hierarchical structure relations, and the classification labels are used for displaying the first file in a classified mode and / or searching the first file. According to the method, refined classification of the network disk files is realized, so that the file management operation of a user is simplified, and the operation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a file classification method, device, storage medium, and program product. Background Technology

[0002] Cloud storage, also known as online USB drive or network hard drive, is a network-based online storage service. Essentially, it involves cloud storage service providers allocating their server hardware resources to users. Cloud storage offers users file storage, sharing, access, backup, and other document management functions. Users can manage and edit files in the cloud storage via the network.

[0003] Users are storing increasingly diverse types of data in their cloud storage, including audio, video, images, documents, and other file formats. The growing scale of these data assets presents challenges for users in managing them. Users face difficulties in finding, editing, categorizing, and cleaning up files, resulting in cumbersome operations and low efficiency.

[0004] While most mainstream cloud storage products currently have basic file categorization functions, they mostly categorize files based on simple attributes such as time and location, which still cannot solve the problems of cumbersome operation and low efficiency for users. Summary of the Invention

[0005] This application provides a file classification method, device, storage medium, and program product, which enables fine-grained classification of files on cloud storage, thereby simplifying user file management operations and improving operational efficiency.

[0006] In a first aspect, embodiments of this application provide a document classification method, including:

[0007] Extract the feature information of the first file stored in the cloud drive;

[0008] Based on the feature information of the first file, a classification model is used to determine the classification label of the first file. The classification label includes a hierarchical label with at least two levels of hierarchical structure. The classification label is used to classify and display the first file and / or to search for the first file.

[0009] In some implementations, the feature information includes type information and at least one of the following: image information, path information, at least part of the file content, and metadata information.

[0010] In some implementations, the feature information includes the type information and the image information, wherein the image information includes the original image data obtained by image decoding of the first file;

[0011] The step of determining the classification label of the first file using a classification model based on the feature information of the first file includes:

[0012] When the type information is an image type, based on the original image data, the classification model is used to classify and identify objects in the first file to obtain object classification labels with at least two levels of hierarchical structure.

[0013] The category tags of the first document include the category tags of the items.

[0014] The step of determining the classification label of the first file using a classification model based on the feature information of the first file further includes:

[0015] When the type information is an image type, based on the original image data, the classification model is used to perform face classification and recognition on the first file to obtain a face classification label. The face classification label is used to indicate whether a face is included and the face information when a face is included.

[0016] The classification tags of the first document include the face classification tags.

[0017] In some implementations, the feature information further includes path information, and the image information further includes character information obtained by optical character recognition of the first file;

[0018] The step of determining the classification label of the first file using a classification model based on the feature information of the first file further includes:

[0019] When the type information is an image type and the object classification label includes a data label, the first file is classified and identified using the classification model based on the character information, the original image data, and the path information to obtain a data classification label with at least two levels of hierarchical structure.

[0020] The category tags of the first document also include the data category tags.

[0021] In some implementations, the feature information further includes metadata information, which includes at least one of shooting time, shooting device, and shooting location, and the classification tag of the first file also includes the metadata information.

[0022] In some implementations, the feature information includes the type information and the path information. The path information includes the path of the first file and the filenames of other files in the same directory as the first file. The type information of the first file is document information, video information, or audio information.

[0023] The step of determining the classification label of the first file using a classification model based on the feature information of the first file includes:

[0024] If the type information is a document type, video type, or audio type, the classification label of the first file is determined based on the path information and the classification model.

[0025] In some implementations, the feature information also includes the at least part of the file content;

[0026] The step of determining the classification label of the first file based on the path information and using the classification model includes:

[0027] Based on the path information and at least part of the file content, the classification model is used to determine the classification label of the first file.

[0028] In some implementations, at least a portion of the file content is extracted in the following manner:

[0029] If the type information is a document type, extract the text of the first file with the first preset number of words, or extract the directory or summary of the first file to obtain the at least part of the file content;

[0030] If the type information is video type, at least some subtitles of the first file are extracted, or auxiliary enhancement information in at least some frames of the first file is extracted, or at least some speech of the first file is converted into text to obtain the at least some file content;

[0031] If the type information is audio, at least a portion of the speech of the first file is converted into text to obtain the content of the at least a portion of the file.

[0032] In some implementations, the step of classifying and recognizing objects in the first file based on the original image data using the classification model to obtain at least two levels of object classification labels includes:

[0033] The original image data is input into the object classification model to obtain at least two levels of object classification labels. The object classification model is trained in batches using multiple image samples. The sample labels of each batch of image samples include object classification labels with at least two levels of hierarchical structure corresponding to the same highest-level object classification.

[0034] Some implementations also include:

[0035] Obtain multiple second files containing category tags, including target tags, where the target tags are human faces or pets;

[0036] Clustering is performed on the multiple second files to obtain at least one target label cluster.

[0037] Some implementations also include:

[0038] Obtain a third file containing the target label, and perform a similarity comparison between the third file and the at least one target label cluster.

[0039] The third file is added to the target label cluster with the highest similarity among the at least one target label clusters, or a new target label cluster including the third file is generated.

[0040] Some implementations also include:

[0041] In response to a file search command, the search text in the file search command is converted into structured search tags;

[0042] Based on the category tags of each file in the cloud drive, files that match the structured search tags are retrieved from the cloud drive.

[0043] Secondly, embodiments of this application provide a document classification device, comprising:

[0044] The extraction module is used to extract the feature information of the first file stored in the cloud drive;

[0045] The classification module is used to determine the classification label of the first file based on the feature information of the first file using a classification model. The classification label includes hierarchical labels with at least two levels of hierarchical structure. The classification label is used to classify and display the first file and / or to search for the first file.

[0046] In some implementations, the feature information includes type information and at least one of the following: image information, path information, at least part of the file content, and metadata information.

[0047] In some implementations, the feature information includes the type information and the image information, wherein the image information includes the original image data obtained by image decoding of the first file;

[0048] The classification module is used for:

[0049] When the type information is an image type, based on the original image data, the classification model is used to classify and identify objects in the first file to obtain object classification labels with at least two levels of hierarchical structure.

[0050] The classification tags of the first document include the face classification tags.

[0051] In some implementations, the classification module is used for:

[0052] Based on the original image data, the first file is classified and identified using the classification model to obtain a face classification label. The face classification label is used to indicate whether a face is included and the face information when a face is included.

[0053] The classification tags of the first document include the face classification tags.

[0054] In some implementations, the feature information further includes path information, and the image information further includes character information obtained by optical character recognition of the first file;

[0055] The classification module is used for:

[0056] When the type information is an image type and the object classification label includes a data label, the first file is classified and identified using the classification model based on the character information, the original image data, and the path information to obtain a data classification label with at least two levels of hierarchical structure.

[0057] The category tags of the first document also include the data category tags.

[0058] In some implementations, the feature information further includes metadata information, which includes at least one of shooting time, shooting device, and shooting location, and the classification tag of the first file also includes the metadata information.

[0059] In some implementations, the feature information includes the type information and the path information. The path information includes the path of the first file and the filenames of other files in the same directory as the first file. The type information of the first file is document information, video information, or audio information.

[0060] The classification module is used for:

[0061] If the type information is a document type, video type, or audio type, the classification label of the first file is determined based on the path information and the classification model.

[0062] In some implementations, the feature information also includes the at least part of the file content;

[0063] The classification module is used for:

[0064] Based on the path information and at least part of the file content, the classification model is used to determine the classification label of the first file.

[0065] In some implementations, the extraction module is used for:

[0066] If the type information is a document type, extract the text of the first file with the first preset number of words, or extract the directory or summary of the first file to obtain the at least part of the file content;

[0067] If the type information is video type, at least some subtitles of the first file are extracted, or auxiliary enhancement information in at least some frames of the first file is extracted, or at least some speech of the first file is converted into text to obtain the at least some file content;

[0068] If the type information is audio, at least a portion of the speech of the first file is converted into text to obtain the content of the at least a portion of the file.

[0069] In some implementations, the classification module is used for:

[0070] The original image data is input into the object classification model to obtain at least two levels of object classification labels. The object classification model is trained in batches using multiple image samples. The sample labels of each batch of image samples include object classification labels with at least two levels of hierarchical structure corresponding to the same highest-level object classification.

[0071] In some implementations, the classification module is also used for:

[0072] Obtain multiple second files containing category tags, including target tags, where the target tags are human faces or pets;

[0073] Clustering is performed on the multiple second files to obtain at least one target label cluster.

[0074] In some implementations, the classification module is also used for:

[0075] Obtain a third file containing the target label, and perform a similarity comparison between the third file and the at least one target label cluster.

[0076] The third file is added to the target label cluster with the highest similarity among the at least one target label clusters, or a new target label cluster including the third file is generated.

[0077] Some implementations also include a search module, used for:

[0078] In response to a file search command, the search text in the file search command is converted into structured search tags;

[0079] Based on the category tags of each file in the cloud drive, files that match the structured search tags are retrieved from the cloud drive.

[0080] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;

[0081] The memory stores computer-executed instructions;

[0082] The processor executes computer execution instructions stored in the memory, causing the processor to perform the method described in any of the first aspects.

[0083] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in any of the first aspects.

[0084] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method shown in any of the first aspects.

[0085] This application provides a file classification method, device, storage medium, and program product. The method, based on the feature information of a first file stored on a cloud drive, uses a classification model to determine the classification tags of the first file. The classification tags include hierarchical tags with at least two levels of hierarchical structure, achieving refined classification of the first file. Based on these hierarchical tags with at least two levels of hierarchical structure, the first file can be categorized and displayed, and / or searched for based on these tags, thereby simplifying user operations for managing cloud drive files and improving operational efficiency. Attached Figure Description

[0086] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0087] Figure 1 A flowchart illustrating a document classification method provided for an exemplary embodiment of this application;

[0088] Figure 2 A schematic diagram of a cloud storage page provided for an exemplary embodiment of this application. Figure 1 ;

[0089] Figure 3 A schematic diagram of a cloud storage page provided for an exemplary embodiment of this application. Figure 2 ;

[0090] Figure 4A schematic diagram of a cloud storage page provided for an exemplary embodiment of this application. Figure 3 ;

[0091] Figure 5 A schematic diagram of a cloud storage page provided for an exemplary embodiment of this application. Figure 4 ;

[0092] Figure 6 A schematic diagram of a cloud storage page provided for an exemplary embodiment of this application. Figure 5 ;

[0093] Figure 7 A schematic diagram of the structure of a document classification device provided for an exemplary embodiment of this application;

[0094] Figure 8 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. Detailed Implementation

[0095] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0096] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0097] In this embodiment of the application, for the first file stored in the cloud drive, by extracting the feature information of the first file, at least two levels of tags for the first file are generated using a classification model, thereby achieving fine classification of the file. File management based on these at least two levels of tags can simplify user operations and improve operational efficiency.

[0098] The technical solutions shown in this application will now be described in detail through specific embodiments. It should be noted that the following embodiments may exist independently or in combination with each other; for identical or similar content, the description will not be repeated in different embodiments.

[0099] The execution subject in this application embodiment can be an electronic device or a file classification device installed in an electronic device. The electronic device can be a cloud storage server, and the file classification device can be implemented through software or a combination of software and hardware. The file classification device can be a processor in the electronic device. For ease of understanding, the following description will use an electronic device as the execution subject.

[0100] Figure 1 This is a flowchart illustrating a document classification method provided for an exemplary embodiment of this application. Please refer to [link / reference]. Figure 1 The methods may include:

[0101] S201. Extract the feature information of the first file stored in the cloud drive.

[0102] The first file can be any file stored on the cloud drive, including images, documents, videos, audio files, etc.

[0103] The feature information of the first file may include one or more of the following: type information, image information, path information, at least part of the file content, and metadata information. In some preferred embodiments, the feature information includes type information and at least one of the following: image information, path information, at least part of the file content, and metadata information.

[0104] For example, when a user saves their first file to a cloud drive, the cloud drive can determine the file type information based on the file format, i.e., the file extension. For instance, the file type information for first files in doc, docx, and pdf formats is "document type"; for jpg, png, and gif formats, it's "image type"; for mp4, avi, and mov formats, it's "video type"; and for mp3, wav, and aac formats, it's "audio type".

[0105] Image information includes information extracted from images or video frames.

[0106] The path information includes the storage path of the first file. The path information may include the path of the first file, or it may include the path of the first file and the filenames of other files in the same directory as the first file.

[0107] At least part of the file content refers to information representing at least part of the content of a file extracted from text, images, audio, or video frames.

[0108] Metadata information describes the attributes of the first file.

[0109] S202. Based on the feature information of the first document, determine the classification label of the first document using a classification model; the classification label is a hierarchical structure label, including hierarchical labels with at least two levels of hierarchical structure relationship, and the classification label is used to classify and display the first document and / or to search for the first document.

[0110] Category tags are hierarchical tags, and each category tag can include multiple levels. For example, the hierarchical structure of category tags can include multiple first-level tags, where each first-level tag can include one or more second-level tags, each second-level tag can include one or more third-level tags, and so on.

[0111] When classifying the first document, the classification tags for the first document include at least two levels of tags, that is, they can include a first-level tag and sub-tags of the first-level tags. It should be noted that the classification tags for the first document can include one or more first-level tags, and correspondingly, one or more sub-tags under each first-level tag.

[0112] The classification labels of the first document may include labels extracted from the semantic feature information of the first document. For example, the classification model may include natural language model, large language model, multimodal large language model, neural network model, etc. The classification model can extract the corresponding semantic feature information based on the feature information of the first document and obtain the corresponding classification labels.

[0113] For the first file of an image type, category tags are used to identify the subject or background of the image, the behavior or action in the image, the time and location of the photo, etc. Taking a category tag with a three-level hierarchical structure as an example, the category tags are described in the format of "first-level tag - second-level tag - third-level tag". For example, they can be "plant - flower - carnation", "plant - tree - sycamore tree", "animal - pet - cat", "animal - mammal - horse", "scene - activity place - taking the subway", "scene - activity place - taking a boat", "scene - natural scenery - rain", "scene - natural scenery - sunrise and sunset", "document - media detection - screen detection", "document - media detection - paper detection", etc. For example, if the first file is a photo of a sycamore tree, the category tag can include "plant - tree - sycamore tree". This category tag includes a three-level hierarchical structure, namely, a first-level tag identified by "plant", a second-level tag identified by "tree", and a third-level tag identified by "sycamore tree". For example, the first file is a photo of a sycamore tree in the rain. The category tags for the first file can be two hierarchical tags with a three-level hierarchical structure: "plant-tree-sycamore tree" and "scene-natural landscape-rain".

[0114] For the first file of a document type, category tags are used to identify the field of content or purpose of the document. Taking a hierarchical structure of two-level tags as an example, category tags are described in the format of "first-level tag - second-level tag". For example, they could be "Exam Type - Certified Public Accountant", "Exam Subject - English", "Content Type - Exam Paper Exercises", "Content Type - Knowledge Point Summary", "Fiction Type - Literary Fiction", "Large Scene - Job Resume", etc. For example, if the first file is a PDF English exam paper, the category tags could include the two hierarchical tags "Exam Subject - English" and "Content Type - Exam Paper Exercises". Similarly, if the first file is a PDF English note, the category tags could include the two hierarchical tags "Exam Subject - English" and "Content Type - Knowledge Point Summary".

[0115] For the first file of video or audio type, category tags are used to identify the field or purpose of the video or audio. Taking a hierarchical tag structure with a two-level relationship as an example, the category tags are described in the format of "first-level tag - second-level tag". For example, it could be "Movie Genre - Comedy", "Movie Genre - Suspense", "Movie Region - Europe and America", "Exam Type - Legal Professional Qualification Examination", "Exam Subject - English", etc. For example, if the first file is a training video of a legal professional qualification examination institution in MP4 format, the category tag could include "Exam Type - Legal Professional Qualification Examination". For example, if the first file is a trailer of a European and American comedy movie in MP4 format, the category tags could include the two hierarchical tags "Movie Genre - Comedy" and "Movie Region - Europe and America", which have a two-level hierarchical relationship.

[0116] In addition to the category tags in the examples above, the category tags of the first file can also include tags extracted from the metadata information of the first file. For example, the metadata information can include tags such as the time and location of the first file.

[0117] In some scenarios, the category tags of the first file are used to categorize and display the first file. For example, cloud storage stores and displays photos or videos uploaded by users in categories. Photos with the same tag are placed in the same folder, so users can view all photos with the same tag in a folder. The same tag can refer to the same first-level tag, the same second-level tag, or the same third-level tag, etc.

[0118] In some scenarios, at least two levels of tags for the first file are used to search for the first file. For example, when a user searches for a file in a cloud storage service, they can enter search text or voice to search. The search text or voice can be tags or other natural language. The cloud storage service matches the user's search text or voice with at least two levels of tags for each of the stored files. If the tags of the first file match the user's search text or voice, then the first file is returned as the search result.

[0119] The file classification method of this application determines the classification tag of the first file based on the type of the first file stored in the cloud drive. The classification tag includes at least two levels of tags, realizing the fine classification of the first file. Based on the at least two levels of tags, the first file can be classified and displayed and / or the first file can be searched based on the at least two levels of tags, thereby simplifying the user's operation of managing cloud drive files and improving operation efficiency.

[0120] Based on the above embodiments, the method for determining the classification label of the first document will be explained.

[0121] In some embodiments, the feature information of the first file includes type information and image information.

[0122] For example, the type information of the first file is image type, and the image information of the first file includes the original image data obtained by image decoding of the first file.

[0123] If the type information of the first file is an image, based on the original image data, a classification model is used to classify and identify objects in the first file, resulting in a hierarchical structure label based on object classification, that is, an object classification label with at least two levels of hierarchical structure relationship; based on the original image data, a classification model is used to classify and identify faces in the first file, resulting in a face classification label; the face classification label is used to indicate whether a face is included and the face information if a face is included; the classification labels of the first file include object classification labels and face classification labels.

[0124] When the type information of the first file is an image, classifying the first file is equivalent to classifying the image. In this embodiment, the image is classified into object classification and face classification. The image decoding method in this embodiment is not limited and can employ image decoding algorithms from related technologies. Object classification can identify the scene, objects, background, environment, animals, plants, etc., of the image. Face classification can identify whether there is a face in the image and the corresponding facial information. Facial information can include the position of the face in the image, the emotion of the face, whether the face image contains multiple people or a single person, and whether the face image is a selfie or a group photo, etc. The classification model for image recognition can, for example, be the ConvNext model.

[0125] For example, if the first file is a photo of a snow-capped mountain, then performing object classification on the first file will yield object classification labels including "scene-natural scenery-snow-mountain". Performing object classification on the first file will yield a face classification label that excludes faces. Accordingly, the classification labels for the first file will include "scene-natural scenery-snow-mountain" and exclude faces.

[0126] For example, if the first file is a selfie of a person in front of a snow-capped mountain, then performing object classification on the first file will yield object classification labels including "scene-natural scenery-snow-mountain". Performing face classification on the first file will yield face classification labels including "face-single person-selfie". Accordingly, the classification labels for the first file will include "scene-natural scenery-snow-mountain" and "face-single person-selfie".

[0127] In some embodiments, object classification and recognition can be performed by training an object classification model. The original image data is input into the object classification model to obtain object classification labels with at least two levels of hierarchical structure. The object classification model is trained in batches using multiple image samples. The sample labels of each batch of image samples include object classification labels with at least two levels of hierarchical structure under the same highest-level object classification.

[0128] During the training of the object classification model, training can be performed in batches. When labeling sample data, only the hierarchical structure labels of the object classification corresponding to the current batch of training can be labeled. For example, if the current batch of training corresponds to the plant classification, meaning the highest-level object classification is plant, then the image samples can be labeled with object classification labels that have a three-level hierarchical structure, such as plant-flower-carnation or plant-tree-sycamore. If there are other objects besides plants in the image, then those other objects do not need to be labeled. During training, a sampling mask strategy is used to update only the model parameters corresponding to the plant classification. The model's output results for other classifications do not participate in the model parameter update, reducing data labeling costs and improving training efficiency.

[0129] In some embodiments, the feature information of the first file includes path information in addition to type information and image information.

[0130] In some embodiments, the image information also includes character information obtained by optical character recognition of the first document.

[0131] For example, the object classification labels obtained by classifying and recognizing the first document include a "document" label. For instance, if the first document is a photograph of a resume on a computer screen, the object classification label would be "document - media detection - screen detection." Similarly, if the first document is a photograph of a paper resume, the object classification label would be "document - media detection - paper detection." In the above scenarios, the first document is identified as "document" during object classification and recognition, and the corresponding media can also be identified. Based on this, the embodiments of this application can further classify the document.

[0132] If the type information of the first file is an image type, and the object classification label obtained by object classification recognition includes a data label, the first file is classified and recognized using a classification model based on character information, original image data, and path information to obtain a data classification label with at least two levels of hierarchical structure; the classification label of the first file also includes a data classification label.

[0133] For example, the document classification labels with at least two levels of hierarchical structure can be "Documents - Certificates - ID Card", "Documents - Certificates - Passport", or "Documents - Contracts - Rental Contract". These document classification labels can clearly define the detailed classification of documents, making it easier for users to manage their files.

[0134] When the category label includes "Document," meaning the first document is document, optical character recognition (OCR) is performed on the first document to identify character information, such as Chinese text, foreign language text, etc. This character information is then used as input to further subdivide the document. In addition to the character information in the first document, its path information can also be extracted. Typically, when users save files, file names, folder names, etc., often contain descriptive information. Extracting the path information of the first document allows us to obtain this descriptive information, which is also used as input to further subdivide the document.

[0135] For example, a classification model can be pre-trained for this type of document. Character information, raw image data, and path information are input into the classification model to obtain document classification labels with at least two levels of hierarchical structure. During the classification model training process, the labels for document samples include document classification labels with at least two levels of hierarchical structure. By training the classification model, the accuracy of document classification is improved. The classification model can be a multimodal large language model.

[0136] Based on the above embodiments, the feature information of the first file may further include metadata information. For a first file of image type, the metadata information includes at least one of shooting time, shooting device, and shooting location. The classification tag of the first file also includes metadata information. For example, the first file is a photo taken by a user at a scenic spot using a certain model of mobile phone. The metadata information of the first file includes the shooting time as "January 1, 2025", the shooting device as a certain model of mobile phone, and the shooting location as a certain scenic spot. All of the above metadata information can be used as the classification tag of the first file.

[0137] In some embodiments, the characteristic information of the first file includes type information and path information. The path information of the first file includes the path of the first file and the filenames of other files in the same directory as the first file.

[0138] Optionally, if the type information of the first file is document, video, or audio, the classification label of the first file is determined based on the path information using a classification model. This classification model can be a natural language model.

[0139] When users store files, file names, folder names, etc., often include descriptive information about the files. This descriptive information can be obtained by extracting the path information of the first file. For example, a user backs up a folder to a cloud drive. The folder is named "Postgraduate Entrance Examination". This folder contains two subfolders: "English" and "Advanced Mathematics". The "English" subfolder contains three videos: one named "Teacher Zhang's Training Video 001", one named "002", and another named "003".

[0140] Taking the video file "Teacher Zhang's Training Video 001" as an example, when classifying this first file, the path information of the first file is extracted, including the path "Postgraduate Entrance Examination / English / Teacher Zhang's Training Video 001", and the filenames of other files in the same directory as "Teacher Zhang's Training Video 001", namely "002" and "003". Based on the above information, two classification labels with a two-level hierarchical structure are obtained, one is "Exam Subject - English" and the other is "Content Type - Training Video".

[0141] Taking a video file named "003" as an example, when classifying this first file, the extracted path information includes the path "Postgraduate Entrance Exam / English / 003" and the filenames of other files in the same directory as "003", namely "Teacher Zhang's Training Video 001" and "002". Thus, when classifying the first file named "003", in addition to obtaining the category tag "Exam Subject - English", we can also obtain the category tag "Content Type - Training Video" based on the information "Teacher Zhang's Training Video 001". In this way, we can achieve fine-grained classification of the first file by utilizing its own path and the filenames of other files in the same directory.

[0142] In some embodiments, the first file feature information may include at least a portion of the file content in addition to type information and path information. Based on the path information and at least a portion of the file content, a classification model is used to determine the classification label of the first file. The classification model may be a natural language model.

[0143] In some embodiments, if the type information of the first file is a document type, the text of the first file with the first preset number of words is extracted, or the directory or summary of the first file is extracted to obtain at least a portion of the file content of the first file.

[0144] Taking a paper with a docx format as the first file as an example, the table of contents of the first file reflects part of the paper's content. By extracting the table of contents, at least part of the file's content can be obtained. Alternatively, the first 500 words of the first file can be extracted. The beginning of a document usually provides an overview of the document's content. By extracting the first 500 words of the first file, at least part of the file's content can be obtained. Or, a text content abstracting algorithm can be used to extract an abstract of the first file, thereby obtaining at least part of the file's content. Combining the path information of the first file with at least part of the file's content for classification improves the accuracy of classification.

[0145] In some embodiments, if the type information of the first file is video, at least some of the subtitles of the first file are extracted, or auxiliary enhancement information (such as hard caption text, interaction and annotation information in video frames) in at least some of the frames of the first file are extracted, or at least some of the speech in the first file is converted into text to obtain at least some of the file content of the first file.

[0146] For the first file of a video type, its content can be represented by video subtitles, audio, or auxiliary enhancement information in some frames. For example, for a movie video, key information such as the movie title can be extracted by extracting the hard-coded text (image text) from the opening frames. For some educational videos, information such as the subject being taught can be extracted by extracting audio or text from some frames. This application does not limit the methods for extracting video subtitles, auxiliary enhancement information in video frames, and speech-to-text conversion; methods from related technologies can be used.

[0147] For example, the first file is a movie video, and the path information of the first file is "comedy / 001". By extracting the hard caption text from the first few frames of the video, the movie name was extracted as at least part of the file content of the first file. Based on the path information and at least part of the file content, the category tag is determined to be "movie type - comedy".

[0148] In some embodiments, if the type information of the first file is audio, at least a portion of the speech of the first file is converted into text to obtain at least a portion of the file content of the first file.

[0149] Similar to the first file of the aforementioned video type, the method for speech-to-text conversion in this application embodiment is not limited, and methods from related technologies can be used. For example, if the first file is the audio of a song, by converting the speech to text and extracting the lyrics, accurate classification can be performed based on the lyrics.

[0150] In the above embodiments, the classification process for the first file can be performed online when the user stores the file to the cloud drive. Building upon the above embodiments, for image files, if their category tags include faces or pets, images of the same person or pet can also be grouped. The process of grouping faces or pets can be performed offline; for example, this process can be performed once a day, adding images of faces or pets uploaded to the cloud drive the previous day to historical groups or generating new groups.

[0151] In some embodiments, multiple second files containing classification labels including target labels are obtained; the multiple second files are clustered to obtain at least one target label cluster. The target label is either a face or a pet.

[0152] When grouping the target label for the first time, since there are no historical groups yet, multiple second files including the target label are clustered to obtain at least one target label cluster, and each cluster is a group. The clustering algorithm is not limited in this embodiment; for example, k-means clustering, density-based clustering, etc., can be used, and the same applies in subsequent embodiments.

[0153] In some embodiments, a third file containing the target label is obtained, and the third file is compared with at least one target label cluster for similarity; the third file is added to the target label cluster with the highest similarity among at least one target label cluster, or a new target label cluster containing the third file is generated.

[0154] After the initial grouping of target tags, subsequent groupings are performed. Since historical groupings already exist (i.e., already generated target tag clusters), the third file containing the target tag is compared for similarity to the previously generated clusters. If the similarity is greater than a preset threshold, it is added to the cluster with the highest similarity. If the similarity is below the preset threshold, it indicates that the target tag in the third file does not belong to the same person as the target tag in the current cluster, and a new cluster is generated. Each subsequent grouping of target tags serves as the historical target tag cluster for the next execution.

[0155] The similarity comparison between the third file and at least one target label cluster can be performed by comparing the third file with any file in each target label cluster, or by comparing the third file with the cluster centers of each target label cluster. This application is not limited to this. It should be noted that in the foregoing embodiments, when classifying image-type files, feature information of each image can be generated and stored. In this embodiment, clustering or similarity comparison can be performed directly based on the image feature information.

[0156] The category tags in this embodiment can be used to categorize and display files uploaded to the cloud drive. The categorization tags can be any type of tag or any level of tag from the aforementioned embodiments. For example, such as... Figure 2 As shown, photos of the same person are displayed in categories. Figure 2 The image shows photos of singer "LMN" and actor "XYZ" categorized and displayed in the cloud drive's photo album. Figure 3 The image shows photos of "snow-capped mountains" and "contracts" categorized in the cloud drive's photo album.

[0157] Optionally, for document, video, and audio files, category tags can also be displayed on the preview pages for those file types, for example, such as... Figure 4 As shown, clicking on a PDF document named "Vocabulary" in the cloud drive displays a "CET-4 / 6 Exam" tab. Clicking the "CET-4 / 6 Exam" tab will take you to an aggregated page of all "CET-4 / 6 Exam" documents.

[0158] For example, such as Figure 5 As shown, clicking on a postgraduate entrance exam English training video named "Video 001" in the cloud drive displays "Postgraduate Entrance Exam" and "English" tags. Clicking the "Postgraduate Entrance Exam" tag will take the user to an aggregated page of all "Postgraduate Entrance Exam" videos. Clicking the "English" tag will take the user to an aggregated page of all "English" videos.

[0159] Furthermore, when users search for files in the cloud storage, they can quickly find the files they need by using category tags. In some embodiments, in response to a file search command, the search text in the file search command is converted into structured search tags; based on the category tags of each file in the cloud storage, files that match the structured search tags are retrieved from the cloud storage.

[0160] It should be noted that when a user enters text to search, the search text in the search command is the text entered by the user. When a user enters voice to search, the search text in the search command refers to the search text obtained after converting the voice into text.

[0161] Converting search text into structured search tags can be achieved by converting the text into synonyms or near-synonyms, thus transforming it into predefined structured search tags. These tags can be any type of category tag or any level of category tag. For example, if a user searches for "kitten" or "cat," it will be converted to "cat." Alternatively, converting search text into structured search tags can involve understanding the user's natural language search text and extracting structured search tags. For instance, the search text can be input into an intent understanding model to obtain structured search tags. These structured search tags are then compared with the category tags of files in the cloud drive, such as through vector similarity comparison, to determine matching files. Searching based on file category tags improves the accuracy and efficiency of search results, facilitating efficient file management for users.

[0162] Optionally, on the page where users search for files in the cloud drive, frequently used category tags and category tags with a large number of files can be displayed below the search box, such as... Figure 6 As shown, category tags such as "snow mountain" and "contract" are displayed below the search box, allowing users to select category tags for searching and improve efficiency.

[0163] Figure 7 This is a schematic diagram of a document sorting device provided for an exemplary embodiment of this application. Please refer to [link / reference]. Figure 7 Document classification device 700, including:

[0164] Extraction module 701 is used to extract feature information of the first file stored in the cloud drive;

[0165] The classification module 702 is used to determine the classification label of the first file based on the feature information of the first file using a classification model. The classification label includes a hierarchical label with at least two levels of hierarchical structure. The classification label is used to classify and display the first file and / or to search for the first file.

[0166] In some implementations, the feature information includes type information and at least one of the following: image information, path information, at least part of the file content, and metadata information.

[0167] In some implementations, the feature information includes type information and image information, and the image information includes the original image data obtained by image decoding of the first file;

[0168] Classification module 702 is used for:

[0169] When the type information is an image, the first file is classified and identified based on the original image data using a classification model to obtain object classification labels with at least two levels of hierarchical structure.

[0170] The first document's category tags include item category tags.

[0171] In some implementations, the classification module 702 is used for:

[0172] Based on the original image data, a classification model is used to perform face classification and recognition on the first file to obtain face classification labels. The face classification labels are used to indicate whether a face is included and the face information when a face is included.

[0173] The first document's category tags include a face category tag.

[0174] In some implementations, the feature information also includes path information, and the image information also includes character information obtained by optical character recognition of the first file;

[0175] Classification module 702 is used for:

[0176] When the type information is image type and the object classification label includes data label, the first file is classified and identified using a classification model based on character information, original image data and path information, and data classification labels with at least two levels of hierarchical structure are obtained.

[0177] The first document's category tags also include data category tags.

[0178] In some implementations, the feature information also includes metadata information, which includes at least one of the following: shooting time, shooting device, and shooting location. The classification label of the first file also includes metadata information.

[0179] In some implementations, the feature information includes type information and path information. The path information includes the path of the first file and the filenames of other files in the same directory as the first file. The type information of the first file is document information, video information, or audio information.

[0180] Classification module 702 is used for:

[0181] When the type information is document type, video type, or audio type, the classification label of the first file is determined based on the path information using a classification model.

[0182] In some implementations, the feature information also includes at least a portion of the file content;

[0183] Classification module 702 is used for:

[0184] Based on the path information and at least part of the file content, the classification label of the first file is determined using a classification model.

[0185] In some implementations, the extraction module 701 is used for:

[0186] If the type information is document type, extract the text of the first file with the first preset number of words, or extract the table of contents or summary of the first file to obtain at least part of the file content;

[0187] If the type information is video, extract at least some of the subtitles from the first file, or extract auxiliary enhancement information from at least some frames of the first file, or convert at least some of the speech of the first file into text to obtain at least some of the file content;

[0188] If the type information is audio, at least a portion of the speech in the first file is converted into text to obtain at least a portion of the file content.

[0189] In some implementations, the classification module 702 is used for:

[0190] The original image data is input into the object classification model to obtain at least two levels of object classification labels. The object classification model is trained in batches using multiple image samples. The sample labels of each batch of image samples include object classification labels with at least two levels of hierarchical structure corresponding to the same highest-level object classification.

[0191] In some implementations, the classification module 702 is also used for:

[0192] Obtain multiple second files containing category tags, including target tags such as human face or pet;

[0193] Clustering multiple second files yields at least one target label cluster.

[0194] In some implementations, the classification module 702 is also used for:

[0195] Obtain a third file containing the target label and cluster it with at least one target label;

[0196] The third file is added to the target label cluster with the highest similarity in at least one target label cluster, or a new target label cluster is generated that includes the third file.

[0197] Some implementations also include a search module, used for:

[0198] In response to a file search command, the search text in the file search command is converted into structured search tags;

[0199] Based on the category tags of each file in the cloud drive, retrieve files that match the structured search tags from the cloud drive.

[0200] The document classification device provided in this application embodiment can execute the technical solution shown in the above method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0201] Figure 8 This is a schematic diagram of the structure of an electronic device provided for an exemplary embodiment of this application. Please refer to... Figure 8 The electronic device 800 may include at least one processor 801 for implementing the file classification method provided in the embodiments of this application.

[0202] Optionally, the electronic device 800 further includes at least one memory 802 for storing program instructions and / or data. The memory 802 is coupled to the processor 801. The coupling in this embodiment is an indirect coupling or communication connection between devices, units, or modules, and can be electrical, mechanical, or other forms, used for information exchange between devices, units, or modules. The processor 801 may operate in conjunction with the memory 802. The processor 801 may execute program instructions stored in the memory 802. At least one of the at least one memory may be included in the processor.

[0203] Optionally, the electronic device 800 further includes a communication interface 803 for communicating with other devices via a transmission medium, thereby enabling the electronic device 800 to communicate with other devices. The communication interface 803 may be, for example, a transceiver, interface, bus, circuit, or a device capable of transmitting and receiving functions. The processor 801 can utilize the communication interface 803 to transmit and receive data and / or information, and to implement the methods provided in the embodiments of this application. For details, please refer to the detailed descriptions in the preceding embodiments; further elaboration is not repeated here.

[0204] This application embodiment does not limit the specific connection medium between the processor 801, memory 802, and communication interface 803. This application embodiment... Figure 8 The processor 801, memory 802, and communication interface 803 are connected via bus 804. Bus 804 is... Figure 8 The connections between other components are shown in thick lines only and are not intended to be limiting. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 8 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0205] It should be understood that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method embodiments can be completed by the integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0206] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0207] Accordingly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods described in the above-described method embodiments.

[0208] Accordingly, embodiments of this application may also provide a computer program product, including a computer program, which, when executed by a processor, can implement the methods shown in the above-described method embodiments.

[0209] The terms “unit”, “module”, etc., used in this specification may be used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution.

[0210] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. In the several embodiments provided in this application, it should be understood that the disclosed apparatus, devices, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0211] The unit described as a separate component may or may not be physically separate. The component shown as a unit may or may not be a physical unit; that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0212] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0213] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. This computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, magnetic tape), optical media (e.g., digital video disc (DVD)), or semiconductor media (e.g., solid-state disk (SSD)).

[0214] If this function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0215] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0216] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A document classification method, characterized in that, include: Extract the feature information of the first file stored in the cloud drive; Based on the feature information of the first file, a classification model is used to determine the classification label of the first file. The classification label includes a hierarchical label with at least two levels of hierarchical structure. The classification label is used to classify and display the first file and / or to search for the first file.

2. The method according to claim 1, characterized in that, The feature information includes type information and at least one of the following: image information, path information, at least part of the file content, and metadata information.

3. The method according to claim 2, characterized in that, The feature information includes the type information and the image information, wherein the image information includes the original image data obtained by image decoding of the first file; The step of determining the classification label of the first file using a classification model based on the feature information of the first file includes: When the type information is an image type, based on the original image data, the classification model is used to classify and identify objects in the first file to obtain object classification labels with at least two levels of hierarchical structure. The category tags of the first document include the category tags of the items.

4. The method according to claim 3, characterized in that, The step of determining the classification label of the first file using a classification model based on the feature information of the first file further includes: When the type information is an image type, based on the original image data, the classification model is used to perform face classification and recognition on the first file to obtain a face classification label. The face classification label is used to indicate whether a face is included and the face information when a face is included. The classification tags of the first document include the face classification tags.

5. The method according to claim 4, characterized in that, The feature information also includes path information, and the image information also includes character information obtained by optical character recognition of the first file; The step of determining the classification label of the first file using a classification model based on the feature information of the first file further includes: When the type information is an image type and the object classification label includes a data label, the first file is classified and identified using the classification model based on the character information, the original image data, and the path information to obtain a data classification label with at least two levels of hierarchical structure. The category tags of the first document also include the data category tags.

6. The method according to claim 3, characterized in that, The feature information also includes metadata information, which includes at least one of shooting time, shooting device, and shooting location. The classification tag of the first file also includes the metadata information.

7. The method according to claim 2, characterized in that, The feature information includes the type information and the path information. The path information includes the path of the first file and the filenames of other files in the same directory as the first file. The type information of the first file is document information, video information, or audio information. The step of determining the classification label of the first file using a classification model based on the feature information of the first file includes: If the type information is a document type, video type, or audio type, the classification label of the first file is determined based on the path information and the classification model.

8. The method according to claim 7, characterized in that, The feature information also includes at least a portion of the file content; The step of determining the classification label of the first file based on the path information and using the classification model includes: Based on the path information and at least part of the file content, the classification model is used to determine the classification label of the first file; At least a portion of the file content was extracted using the following method: If the type information is a document type, extract the text of the first file with the first preset number of words, or extract the directory or summary of the first file to obtain the at least part of the file content; If the type information is video type, at least some subtitles of the first file are extracted, or auxiliary enhancement information in at least some frames of the first file is extracted, or at least some speech of the first file is converted into text to obtain the at least some file content; If the type information is audio, at least a portion of the speech of the first file is converted into text to obtain the content of the at least a portion of the file.

9. The method according to claim 4 or 5, characterized in that, Based on the original image data, the classification model is used to classify and identify objects in the first file, resulting in at least two levels of object classification labels, including: The original image data is input into the object classification model to obtain at least two levels of object classification labels. The object classification model is trained in batches using multiple image samples. The sample labels of each batch of image samples include object classification labels with at least two levels of hierarchical structure corresponding to the same highest-level object classification.

10. The method according to any one of claims 1-8, characterized in that, Also includes: Obtain multiple second files containing category tags, including target tags, where the target tags are human faces or pets; Clustering is performed on the multiple second files to obtain at least one target label cluster.

11. The method according to claim 10, characterized in that, Also includes: Obtain a third file containing the target label, and perform a similarity comparison between the third file and the at least one target label cluster. The third file is added to the target label cluster with the highest similarity among the at least one target label clusters, or a new target label cluster including the third file is generated.

12. The method according to any one of claims 1-8, characterized in that, Also includes: In response to a file search command, the search text in the file search command is converted into structured search tags; Based on the category tags of each file in the cloud drive, files that match the structured search tags are retrieved from the cloud drive.

13. An electronic device, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-12.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-12.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-12.