File classification method, device and equipment and readable storage medium

By adding detailed classification path description information to each category node in the multi-level classification structure and using large models for reasoning, the problem of degradation of file classification accuracy and efficiency in the existing technology is solved, and efficient and stable file classification is achieved.

CN120256636APending Publication Date: 2025-07-04ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510321767.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing file classification methods have cumulative errors when dealing with multi-level categories and dynamically changing scenarios, resulting in a decrease in classification accuracy and efficiency. Traditional methods require frequent updates of models, affecting business continuity and stability.

Method used

Convert file classification tasks into large-scale generation tasks. By adding detailed classification path description information to each category node in the multi-level classification structure, and using the large-scale model to combine these description information for inference, the hierarchical classification results of the file are generated to reduce the cumulative error in the hierarchical classification process.

Benefits of technology

It improves the accuracy and efficiency of file classification, can adapt to changes in the category system in real time, without the need to reconstruct and train models, and ensures the continuity and stability of the classification system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256636A_ABST
    Figure CN120256636A_ABST
Patent Text Reader

Abstract

The invention provides a file classification method, device and equipment and a readable storage medium, and the method comprises the steps: converting a file classification task into a generation task of a large model, determining the classification path description information of each class node in each hierarchy through a preset multi-hierarchy classification structure, and carrying out the classification path description information of each class node in each hierarchy; the information comprises all nodes passing from a hierarchical root node to a category node and description information of classification categories of the nodes, and a hierarchical category result to which the file belongs is generated through reasoning by combining a large model with the classification path description information to represent a hierarchical classification category of the file, so that accumulative errors in a layer-by-layer classification process are reduced; the problem of classification performance reduction caused by information loss is avoided, and the file classification accuracy and classification efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a file classification method, device, equipment, and readable storage medium. Background Art

[0002] In the context of big data and informatization, efficient classification of file materials plays an important role in improving the management efficiency of digital assets and the business processing ability. However, the existing file classification methods have obvious deficiencies in dealing with multi-level categories and dynamically changing file classification scenarios.

[0003] Currently, the hierarchical classification method can classify layer by layer, but cumulative errors are likely to occur during the hierarchical classification process, thus affecting the classification accuracy; while the classification method that directly uses the leaf nodes of the multi-level classification system as the basis for file classification ignores the important classification information contained in the parent nodes and ancestor nodes of the leaf nodes, resulting in a decline in classification performance. Summary of the Invention

[0004] In view of this, to solve the above technical problems, this application provides a file classification method, device, equipment, and readable storage medium.

[0005] Specifically, this application is implemented through the following technical solutions:

[0006] According to the first aspect of the embodiments of this application, a file classification method is provided, and the method includes:

[0007] For a preset multi-level classification structure, determine the classification path description information of each category node in each level; wherein, the classification path description information is determined based on all the nodes passed from the root node of the level to this category node and the description information of the classification category represented by each node;

[0008] Obtain the text sequence obtained by parsing the file to be classified;

[0009] Associate and input the text sequence with the classification path description information of each category node into a large model, and obtain the hierarchical classification result output by the large model after reasoning based on the classification path description information of each category node, and this result indicates the hierarchical classification category to which the file to be classified belongs in the multi-level classification structure.

[0010] Optionally, the method further includes:

[0011] For each level of the multi-level classification structure, construct and generate prompt words; wherein, the generated prompt words are used to prompt the large language model to enhance the discrimination between the description information of each category node at the same level;

[0012] Input each category node at the same level in the multi-level classification structure and the generated prompt words into the large language model, and obtain the description information of each category node at the same level output by the large language model.

[0013] Optionally, when there are multiple category nodes at the same level in the multi-level classification structure, the method further includes:

[0014] Obtain multiple batches of description information generated multiple times for the multiple category nodes at the same level; each batch includes the description information of all category nodes at this level;

[0015] For the description information in each batch, based on the text similarity between the description information of pairwise category nodes at the same level, determine the text description discrimination degree of this batch;

[0016] Use the batch of description information with the highest text description discrimination degree among the multiple batches of description information as the description information of the multiple category nodes at the same level.

[0017] Optionally, the determining the text description discrimination degree of this batch based on the text similarity between the description information of pairwise category nodes at the same level includes:

[0018] For the multiple description information corresponding one by one to each category node in the same batch, respectively determine the complement of the text similarity between every two description information with respect to 1 as the text description discrimination degree between these two description information;

[0019] Sum up the text description discrimination degrees between all different pairwise description information, and determine the average discrimination degree according to the summation result and the logarithm of all different pairwise description information as the text description discrimination degree of this batch.

[0020] Optionally, the determining the classification path description information of each category node in each level includes:

[0021] For each node passed from the root node of the level to this category node, organize the node and the description information of the classification category it represents into an information item;

[0022] In the path order of the passed nodes, integrate each information item through a set symbol link to form the classification path description information of this category node.

[0023] Optionally, the large model determines the category node to which the text sequence corresponding to the file to be classified belongs according to the classification path description information of each category node, and generates a category sequence formed by the nodes passed from the root node of the level to the belonging category node as the hierarchical classification result of the text sequence corresponding to the file to be classified; the method further includes:

[0024] Calculate the character matching similarity of the category sequence output by the large model with the category nodes at each level of the multi-level classification structure, and determine the node path in the multi-level classification structure with the highest character similarity to the category sequence;

[0025] Use the category sequence represented by all nodes in the node path as the updated hierarchical classification result of the text sequence corresponding to the file to be classified.

[0026] Optionally, the associating and inputting the text sequence with the classification path description information of each category node into the large model includes:

[0027] Construct a classification prompt word for document classification according to the text sequence and the classification path description information of each category node; the classification prompt word at least further includes the description information of the classification task executed by the large model and the corresponding classification output requirements;

[0028] Input the classification prompt word into the large model so that the large model generates the hierarchical classification result of the text sequence corresponding to the file to be classified according to the classification prompt word.

[0029] Optionally, the obtaining the text sequence obtained by parsing the file to be classified includes:

[0030] When the file to be classified includes a picture, obtain at least one of the text information included in the picture, the file name, and the text description of the content included in the picture as the text sequence of the picture;

[0031] When the file to be classified includes audio or video, obtain the text information obtained by converting the speech included in the audio or video as the text sequence of the audio or video.

[0032] According to the second aspect of the embodiments of the present application, there is provided a file classification device, and the device includes:

[0033] A classification preprocessing module, configured to determine the classification path description information of each category node in each level for a preset multi-level classification structure; wherein, the classification path description information is determined based on all nodes passed from the hierarchical root node to this category node and the description information of the classification category represented by each node;

[0034] A file parsing module, configured to obtain the text sequence obtained by parsing the file to be classified;

[0035] A large model classification module is used to input the text sequence and the classification path description information of each category node into a large model in an associated manner, and obtain the hierarchical classification result output by the large model after reasoning based on the classification path description information of each category node. This result indicates the hierarchical classification category to which the file to be classified belongs in the multi-level classification structure.

[0036] Optionally, the device further includes:

[0037] For each level of the multi-level classification structure, generate prompting words; wherein, the generated prompting words are used to prompt the large language model to enhance the discrimination between the description information of each category node at the same level.

[0038] Input the category nodes at the same level in the multi-level classification structure and the generated prompting words into the large language model together, and obtain the description information of each category node at the same level output by the large language model.

[0039] Optionally, in the case where there are multiple category nodes at the same level in the multi-level classification structure, the device further includes:

[0040] A batch generation module is used to obtain multiple batches of description information generated multiple times for the multiple category nodes at the same level; each batch includes the description information of all category nodes at this level.

[0041] A discrimination calculation module is used to, for the description information in each batch, determine the text description discrimination of this batch based on the text similarity between the description information of two category nodes at the same level.

[0042] A screening module is used to select the batch of description information with the highest text description discrimination among the multiple batches of description information as the description information of the multiple category nodes at the same level.

[0043] Optionally, the discrimination calculation module is specifically used for:

[0044] For the multiple description information corresponding one by one to each category node in the same batch, respectively determine the complement of the text similarity between every two pieces of description information relative to 1 as the text description discrimination between these two pieces of description information.

[0045] Sum up the text description discriminations between all different pairs of description information, and determine the average discrimination based on the sum result and the logarithm of all different pairs of description information as the text description discrimination of this batch.

[0046] Optionally, the classification preprocessing module is specifically used for:

[0047] For each node from the root node of the hierarchy to the category node, organize the description information of the node and the classification category it represents into an information item;

[0048] According to the path sequence of the passed nodes, each information item is integrated by setting symbolic links to form the classification path description information of the category node.

[0049] Optionally, the large model determines the category node to which the text sequence corresponding to the to-be-classified file belongs based on the classification path description information of each category node, and generates a category sequence formed by the nodes passed from the hierarchical root node to the category node as a hierarchical classification result of the text sequence corresponding to the to-be-classified file; the device further includes:

[0050] Calculate the character matching similarity of the category nodes of the category sequence output by the large model and the multi-level classification structure layer by layer, and determine the node path in the multi-level classification structure with the highest character similarity to the category sequence;

[0051] The category sequence represented by all nodes in the node path is used as the hierarchical classification result after the text sequence corresponding to the file to be classified is updated.

[0052] Optionally, the large model classification module is specifically used for:

[0053] According to the text sequence and the classification path description information of each category node, a classification prompt word for document classification is constructed; the classification prompt word at least includes description information of the classification task performed by the large model and the corresponding classification output requirements;

[0054] The classification prompt word is input into the large model so that the large model generates a hierarchical classification result of the text sequence corresponding to the file to be classified according to the classification prompt word.

[0055] Optionally, the file parsing module is specifically used for:

[0056] In the case where the file to be classified includes a picture, obtaining at least one of text information included in the picture, a file name, and a text description of the content included in the picture as a text sequence of the picture;

[0057] In the case that the file to be classified includes audio or video, text information obtained by converting the speech included in the audio or video is obtained as a text sequence of the audio or video.

[0058] According to a third aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the above file classification method by calling the computer program.

[0059] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above file classification method is implemented.

[0060] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:

[0061] In the above technical solution provided by the present application, the file classification task is converted into a generation task of a large model. By determining the classification path description information of each category node in each level of a preset multi-level classification structure, this information includes the description information of all nodes and their classification categories passed from the root node of the level to this category node, and the large model combines this classification path description information to infer and generate the hierarchical category result to which the file belongs to represent the hierarchical classification category of the file, reducing the cumulative error in the process of layer-by-layer classification, avoiding the problem of decline in classification performance caused by information loss, and improving the accuracy and classification efficiency of file classification.

[0062] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. In addition, any one of the embodiments of the present application does not need to achieve all of the above effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0064] Figure 1A is a flowchart of the steps of a file classification method shown in an exemplary embodiment of the present application;

[0065] Figure 1B is a schematic diagram of a multi-level classification structure of files shown in an exemplary embodiment of the present application;

[0066] Figure 2 is a schematic diagram of generating the description information of the classification category represented by a category node shown in an exemplary embodiment of the present application;

[0067] Figure 3A is a flowchart of the steps of generating description information and evaluation in multiple batches shown in an exemplary embodiment of the present application;

[0068] Figure 3BIt is a schematic diagram for calculating the text description discrimination degree of a batch shown in an exemplary embodiment of the present application;

[0069] Figure 4 It is an example diagram of a multi-level classification structure shown in an exemplary embodiment of the present application;

[0070] Figure 5 It is a schematic diagram of the structure of a document classification device shown in an exemplary embodiment of the present application;

[0071] Figure 6 It is a schematic diagram of the hardware of an electronic device shown in an exemplary embodiment of the present application. Detailed implementation manners

[0072] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0073] It should be understood that although terms such as first, second, and third may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first classification threshold may also be referred to as the second classification threshold, and similarly, the second classification threshold may also be referred to as the first classification threshold. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0074] In the context of big data and informatization, efficient classification of document materials plays an important role in improving the management efficiency of digital assets and the business processing ability. For example, for a large number of document materials generated within an enterprise, they constitute rich digital assets. To make full use of these digital assets and assist the subsequent development and evolution of the enterprise, it is necessary to classify the document materials efficiently to improve the efficiency of business query and use and provide strong support for enterprise decision-making. However, the existing document classification methods have obvious deficiencies in dealing with multi-level categories and dynamically changing document classification scenarios.

[0075] At present, the hierarchical classification method can classify layer by layer, but it is necessary to perform detailed data annotation for each level, which increases the workload of data preparation. At the same time, cumulative errors are likely to occur during the multi-level classification process, thus affecting the final classification accuracy. The classification method that directly uses the leaf nodes of the multi-level classification system as the basis for file classification ignores the important classification information contained in the parent nodes and ancestor nodes of the leaf nodes, resulting in a decline in classification performance. In addition, most of the existing methods are based on fixed classification categories. When the category system changes, it is necessary to reconstruct and train the model, which not only takes time and effort, but may also affect the continuity and stability of the business due to frequent model updates. The limitations of the existing classification methods seriously affect the efficiency and accuracy of document classification in practical applications.

[0076] In view of this, in response to the limitations of the existing file classification methods in dealing with multi-level categories and dynamic change scenarios, the present application proposes a zero-shot multi-level file classification method based on a large model. This method converts the document classification task into a generation task of the large model, making it more in line with the training paradigm of the large model and fully utilizing the powerful capabilities of the large model in text understanding and generation. Before classifying using the large model, an extended description of each classification category at each level is performed and a description related to the complete path is generated to ensure that the key information of the classification categories at different levels is retained and is more easily understood and processed by the large language model, avoiding the problem of decline in classification performance caused by information loss in traditional methods.

[0077] The file classification method provided by the present application can be widely applicable to file classification scenarios of multi-level classifications in different fields, such as enterprise document management; academic papers are classified and managed according to subject fields, research directions, publication years, etc.; policy documents, regulations, announcements and other documents are classified and managed according to policy fields, issuing agencies, publication years, etc.; legal cases, contracts, legal documents, etc. of law firms or legal institutions are classified according to legal fields, case types, parties, etc.; teaching resources of educational institutions or online learning platforms are classified and managed according to subjects, grades, course types, etc.; personal documents such as photos, documents, videos, etc. are classified and managed according to time, themes, types, etc.

[0078] See Figure 1A The step flowchart of the exemplary file classification method, the file classification method provided by the present application can at least include the following steps:

[0079] S101, for a preset multi-level classification structure, determine the classification path description information of each category node in each level; wherein, the classification path description information is determined based on all the nodes passed from the root node of the level to the category node and the description information of each node representing the classification category;

[0080] This multi-level classification structure represents a system for classifying documents, which includes multiple classification levels where the classification categories are refined level by level according to the document setting logic, attribute relationships, or other classification criteria, and there is an inclusion or being-included relationship between adjacent levels. Each level includes one or more category nodes, representing different classification categories. The classification categories represented by the category nodes at the same level belong to the same level of sub-classification categories, which are in a parallel relationship at the same level and are all further subdivisions of the classification categories of the adjacent upper level.

[0081] In the multi-level classification structure, the top-level node is called the level root node, which is the starting point or basic classification for all other nodes; each specific classification category is regarded as a category node, and these nodes are connected to each other according to the hierarchical relationship to form this multi-level classification structure.

[0082] For example, referring to Figure 1B A schematic diagram of a multi-level classification structure of documents shown exemplarily. The first level includes the node "Design Department Documents", which serves as the level root node. Next, it is divided into two types of nodes at the second level, namely office documents and design documents, according to the document type. Among them, the node "Office Documents" is further subdivided into two types of nodes at the third level, namely "Reports" and "Plans". The nodes "Reports" and "Plans" can be further classified according to the cycle into three types of nodes at the fourth level, namely "Annual Reports / Plans", "Quarterly Reports / Plans", and "Monthly Reports / Plans". Similarly, the node "Design Documents" is further subdivided into two types of nodes at the third level, namely "Drawings" and "Models". The node "Drawings" can be classified according to the design field into four types of nodes at the fourth level, namely "Architectural Drawings", "Mechanical Drawings", "Electrical Drawings", and "Other Drawings", while the node "Models" can be divided into two types of nodes at the fourth level, namely "3D Models" and "2D Models".

[0083] The description information of the classification category represented by each node is an extended description of the classification category represented by the node, which may include but is not limited to: the specific definition of the classification category, such as the covered scope, main features, or attributes, to be used to establish the boundary of the classification category; examples of the classification category, such as providing specific instances or application scenarios belonging to the classification category. The instances can be specific objects, concepts, or events, which together reflect the core features of the classification category.

[0084] For example, taking Figure 1BTaking the category node "Mechanical Drawings" in [[]] as an example, its explanatory information can include, but is not limited to: drawings drawn to express information such as the structure, dimensions, materials, and technical requirements of mechanical components, using specific graphical symbols, lines, dimension markings, and written descriptions, etc., to accurately describe the shape, size, assembly relationship, and manufacturing process requirements of mechanical components, including detailed information such as the geometric shape, dimension markings, material identification, tolerance requirements, surface roughness, and heat treatment requirements of mechanical components.

[0085] In this embodiment, when the large model processes the file classification task, it relies on an accurate understanding of the classification categories. By adding explanatory information to each category node at each level of the multi-level classification structure, rich context is provided for the large model. At the same time, there may be subtle differences or overlaps between different classification categories at the same level. By providing detailed explanatory information, the uniqueness and distinguishing points of each category can be highlighted, helping to distinguish similar but different classification categories, so that the large model can more accurately grasp the essential characteristics of the classification categories, thereby improving the accuracy and efficiency of classification.

[0086] The classification path description information of each category node is used to clearly characterize the position of the classification category represented by this category node in the entire multi-level classification structure, as well as the connections and differences between this classification category and other classification categories, including the node path information from the hierarchical root node to this category node, which can be specifically determined based on all the nodes passed from the root node to this category node and the explanatory information of the classification category represented by each node.

[0087] Based on the fact that explanatory information is added to the classification category represented by each category node in the multi-level classification structure, for any category node in this multi-level classification structure, according to the node order of each node included in the node path from the hierarchical root node to this category node, the classification categories represented by these nodes and the explanatory information of each node in these nodes can be organized into ordered data to form the classification path description information of this category node. Among them, data formats such as key-value pairs, sets, or JSON can be used to organize the classification path description information.

[0088] Still taking Figure 1B the category node "Mechanical Drawings" in [[]] as an example, the classification path description information of this category node at least includes {root node "Design Department Documents" + explanatory information}, {second-level node "Design Documents" + explanatory information}, {third-level node "Drawings" + explanatory information}, and {fourth-level node "Mechanical Drawings" + explanatory information}.

[0089] S102, obtain the text sequence obtained by parsing the file to be classified;

[0090] The file to be classified indicates a file that needs to be classified and can be a file in any format, including but not limited to documents (such as PDF, Word, Excel, PPT, etc.), pictures, audio, and video. The text sequence includes a data sequence presented in text form extracted from the file to be classified, and the text sequence is the basis for subsequent classification tasks.

[0091] Regarding the acquisition of this text sequence, for the text information included in the file to be classified, this type of text information can be directly extracted as a component of the text sequence of the file to be classified. In the case where the file to be classified includes pictures, at least one of the text information included in the pictures, the file name, and the text description of the content included in the pictures can be obtained as the text sequence of the pictures. That is, for pictures that themselves contain text information, the text information in the pictures can be directly extracted as the text sequence of the pictures; for pictures that do not contain text information, text descriptions can also be generated for the image content included in the pictures through large language models or other image processing tools, and thus the text descriptions can be used as the text sequences of the pictures.

[0092] In the case where the file to be classified includes audio or video, the text information obtained by converting the speech included in the audio or video is used as the text sequence of the audio or video. That is, for audio or video with subtitle information, the subtitle information can be directly stored in the order of the audio or video playback progress to obtain a text sequence composed of subtitle information; for video or audio without subtitles, the speech can be converted into text information through existing speech-to-text technologies, and this text information is arranged in the order of speech playback to form the text sequence of the audio or video.

[0093] In the process of parsing the file to be classified into a text sequence, various files such as PDF, Office documents (Word, Excel, PPT), and pictures can be automatically converted into a text sequence in Markdown format. The Markdown format, with its simplicity, good readability, and support for structured content, enables large language models to understand and generate text more efficiently.

[0094] To further improve the quality of the obtained text sequence to enhance the classification accuracy, data cleaning can also be performed on the obtained text sequence to remove noise and useless information in the text, such as extra spaces, line breaks, special characters, HTML tags, etc., and retain the content that can better represent the data characteristics and is beneficial to the document classification task. At the same time, further processing of the text sequence can be carried out through cleaning operations such as spelling correction, grammar correction, missing word filling, misspelled word correction, duplicate item deletion, special symbol removal, stop word removal, and regularization of text format, so as to ensure the quality of the text data input into the subsequent module, making it cleaner, more accurate, and easier to understand. In this embodiment, the above data cleaning operations can be flexibly selected to adapt to different real-world scenarios and achieve a comprehensive and clear extraction of the content of the file to be classified.

[0095] S103, associate and input the text sequence with the classification path description information of each category node into the large model, and obtain the hierarchical classification result output by the large model after reasoning based on the classification path description information of each category node. This result indicates the hierarchical classification category to which the file to be classified belongs in the multi-level classification structure.

[0096] The large model described in this embodiment includes a large neural network model in the field of deep learning. This type of model has a large number of parameters and a complex network structure, and can capture and process complex features and patterns in the data. For example, the large model can include, but is not limited to, Transformer models such as BERT, GPT series, convolutional neural network models, etc. Through its complex network structure and a large number of parameters, the large model can automatically extract useful feature representations from the input text data and capture the context information in the text data to accurately understand the meaning of the text data and perform classification.

[0097] By inputting the text sequence of the file to be classified and the classification path description information of each category node in the multi-level classification structure into the large model together, based on the fact that the classification path description information of the category node includes all the nodes passed from the root node to this category node and the description information of the classification category represented by each node, the large model can fully understand the hierarchy and logical relationship of the multi-level classification structure after receiving this classification path description information.

[0098] Based on the detailed classification path description information, the large model can conduct in-depth reasoning and analysis on the file to be classified. According to the content of the text sequence and in combination with the classification path description information of each category node, it gradually infers the most likely classification category of the file to be classified at each classification level. This process not only considers the direct relevance of the text content but also incorporates the hierarchy and logical relationships of the classification structure, thereby improving the accuracy and rationality of classification. The hierarchical classification result output by the large model clearly indicates the level and specific classification category to which the file to be classified belongs in the multi-level classification structure. This hierarchical classification result is based on the large model's in-depth understanding of the text content and accurate grasp of the classification structure, and thus has high reliability and practicality.

[0099] Regarding the associated input of the text sequence and the classification path description information of each category node into the large model, in order to better guide the classification logic of the large model, a classification prompt word for instructing the large model to classify files can be constructed according to the text sequence and the classification path description information of each category node. This classification prompt word can provide a structured information framework to help the large model better understand the input content and perform accurate and efficient classification reasoning. After obtaining the classification prompt word, it can be input into the large model so that the large model generates the hierarchical classification result of the text sequence corresponding to the file to be classified according to this classification prompt word.

[0100] The classification prompt word integrates the text sequence and the classification path description information to clarify the classification goal, path, and possible classification options for the large model, thereby enhancing the large model's ability to interpret the input information and improving the accuracy and efficiency of classification. The classification prompt word can at least also include the description information of the classification task executed by the large model and the corresponding classification output requirements, so as to obtain the hierarchical classification result that meets the classification output requirements output by the large model. In addition, in addition to the basic requirements for the language, style, etc. of the text generation, customized requirements can also be added to the classification prompt word, making the content generated by the large model based on this prompt word more in line with the requirements of document classification. For example, examples of customized requirements:

[0101] Requirement 1: "Please select the most likely category from the given ##category nodes##"

[0102] Requirement 2: "Please give the reason for classification based on the content of ##text sequence## and the classified category"

[0103] Requirement 1 limits the scope of the content generated by the large model for the document category, preventing it from generating some information that is completely irrelevant to the document category in its generative output. Requirement 2 combines the large model's powerful understanding, summarization, and question-and-answer capabilities, enabling it to provide interpretability for its classification results.

[0104] The construction of classification prompts can be achieved using a classification prompt template, which includes a fixed prompt structure and replaceable variable parameters. In this embodiment, the text sequence and the classification path description belong to replaceable variable parameters. By inserting the text sequence of the file to be classified and the classification path description information of each category node as variable parameters into the classification prompt template, the template filling can be completed, thereby generating a classification prompt for a specific input.

[0105] For example, an exemplary classification prompt may include the following content:

[0106] [You are a file classification assistant. You can reason based on the classification path description information of each category node provided by the following multi-level classification structure module, identify the category node to which the input text sequence belongs, and output the hierarchical classification result for this text sequence.

[0107] Please fully understand and learn the classification path description information of each category node provided by this multi-level classification structure module:

[0108]

Multi-level classification structure module (input all category nodes and their classification path description information)

[0109] Please identify the category node to which the following text sequence belongs based on the learned classification categories, and output the hierarchical classification result according to the output format requirements below:

[0110]

Text sequence

[0111]

Output format requirements: (such as outputting the position information of the category node to which this text sequence belongs in the multi-level classification structure, including the complete node path from the root node to the category node to which it belongs)

[0113] After inputting the constructed classification prompt into the large model, it can more effectively guide the classification process of the large model, enabling the large model to perform more accurate classification reasoning based on the structured prompt information, thereby generating the hierarchical classification result of the text sequence corresponding to the file to be classified and improving the accuracy and efficiency of classification.

[0114] The hierarchical classification result refers to the result of assigning the file to be classified to the most appropriate classification level and classification category according to the preset multi-level classification structure of the file, indicating the specific position of the file to be classified in this multi-level classification structure.

[0115] For example, taking the file to be classified as a mechanical drawing, generating a text description of the content included in the mechanical drawing as the parsed text sequence, and comparing this text sequence with Figure 1B ​The classification path description information of each category node in the multi-level classification structure shown is input into the large model. In an ideal state, the hierarchical classification result generated by the large model is like "Design Department Documents - Design Documents - Drawings - Mechanical Drawings".

[0116] Based on the fact that the large model is a generative model and is affected by different model parameters and different settings, the large model may not necessarily generate a hierarchical classification result that conforms to the specification. For example, for the above hierarchical classification result of mechanical drawings, the generation of the large model may be "Design Department Documents - Design Documents - Design-Type Drawings - Design-Type Mechanical Drawings". Therefore, the hierarchical classification result generated by the large model can be further processed to obtain a hierarchical classification result that conforms to the multi-level classification structure. For specific details, please refer to the subsequent embodiments.

[0117] Regarding the hierarchical classification result, it can also include the classification basis for which the large model determines the text sequence as the category of this hierarchical classification. This classification basis represents the information or logic relied on by the large model when determining the text sequence as a certain hierarchical classification category. It can include analysis information such as keywords, phrases, sentence structures, and context relationships in the text sequence, as well as knowledge, rules, and patterns learned within the large model. By comprehensively considering these factors, the large model can judge the matching degree between the text sequence and the classification category represented by each category node and give the corresponding classification result.

[0118] This classification basis can be used as the basis for explaining the classification result, enhancing the interpretability and traceability of the hierarchical classification result output by the large model, enabling users to understand how the large model makes classification decisions, helping users understand why the text is classified into a certain category, and thus increasing users' trust in the hierarchical classification result.

[0119] In the embodiments of the present disclosure, converting the file classification task into a large model generation task and generating the category path to which the document belongs through the large model can capture the semantic information of the document more accurately, reduce the cumulative error in the hierarchical classification process, and thus significantly improve the classification accuracy. Specifically, the method first determines the classification path description information of each category node in each level for a preset multi-level classification structure. This information includes the description information of all nodes and their classification categories passed from the root node of the level to this category node. Then, the text sequence of the file to be classified is obtained, and this text sequence is associated with the classification path description information of each category node and input into the large model, so that the large model can more accurately understand and process the key information of different hierarchical categories, and infer and realize the accurate positioning of the file to be classified in the multi-level classification structure according to the input information, thereby determining the hierarchical classification category to which the file to be classified belongs in the multi-level classification structure. In this way, it is ensured that the description information of each category is completely retained, avoiding the problem of decreased classification performance caused by information loss in the traditional method, and improving the accuracy and classification efficiency of file classification.

[0120] In addition, based on the above method, in the case of changes in the multi-level classification structure, the classification path description information can be reconstructed for the newly added, deleted, or replaced category nodes according to the multi-level classification structure, and then input into the large model for file classification, realizing the rapid classification of files under the change of the classification category system. It can adapt to the change of the category system in real time without reconstructing and training the model again, which not only reduces the maintenance cost, but also ensures the continuity and stability of the classification system.

[0121] In some embodiments, in order to enhance the accuracy and efficiency of the description information of each category node in the multi-level classification structure, this embodiment provides an implementation method for automatically constructing the description information. For the description information of each category node in each level of the multi-level classification structure described in the foregoing embodiments, the text generation and understanding capabilities of the large language model can be utilized, combined with the algorithm advantages of deep learning, to automatically generate the description information for the classification category represented by this category node. This description information can clearly and accurately convey the core features and meanings of each category node, thus providing a more intuitive and understandable classification navigation for users.

[0122] Based on the possible overlapping parts among the classification categories represented by each category node in the same level of the multi-level classification structure, it may be difficult to fully ensure the distinguishability between the description information of each category node in the same level only relying on the automatic generation ability of the large language model. Therefore, in order to further improve the accuracy and practicality of the description information, this embodiment further introduces the construction process of generating prompt words. Through carefully designed prompt words, the large language model is guided to pay more attention to the differences between different category nodes in the same level when generating description information, so as to enhance the distinguishability and practicality of the description information.

[0123] That is, referring to Figure 2 the schematic diagram of generating the description information of the classification category represented by an exemplary category node, this embodiment can be implemented through the following steps:

[0124] S201. For each level of the multi-level classification structure, first construct a generation prompt word; wherein, the generation prompt word is used to prompt the large language model to enhance the distinguishability between the description information of each category node in the same level;

[0125] The construction of the generation prompt word can be achieved through a prompt word template. The prompt word template includes fixed prompt text such as "Please generate description information for these classification categories of {replaceable parameter variable} in the same level, and require that the description information between each category node in the same level is distinguishable from each other", and use the classification categories represented by each category node in the same level as the replaceable parameter variable. When constructing the generation prompt word for the current level, fill the classification categories of the category nodes included in the current level into the prompt word template to generate the generation prompt word for the current level.

[0126] Alternatively, in order to more precisely guide the large language model, a generation prompt word can be constructed for each category node in each level of the multi-level classification structure. For example, the prompt word template includes fixed prompt text such as "Given the classification category {X1} represented by the category node {M1}, expand the description of it, and require it to be different from other classification categories {X2...} in the same level", then the generation prompt word specific to each category node can be constructed for each category node by filling the template.

[0127] S202. Input each category node in the same level of the multi-level classification structure and the generation prompt word into the large language model, and obtain the description information of each category node in the same level output by the large language model.

[0128] The large language model used in this embodiment and the large model used for document classification mentioned above are two independent models. This large language model can fully understand the input category nodes and automatically generate explanatory information for the classification categories represented by these category nodes. This large language model can be implemented using models such as BERT or GPT series, or it can be of the same model type as the large model.

[0129] Input each category node at the same level in the multi-level classification structure and the generated prompt words into the large language model. Through this step, the large language model can fully understand and integrate the information of the category nodes and the generated prompt words, and then output explanatory information for each category node at the same level. These explanatory information not only accurately reflect the core features of each category node, but also effectively improve the discrimination and recognition between the explanatory information of different category nodes at the same level through the guidance of the generated prompt words.

[0130] In the embodiments of the present disclosure, by constructing generated prompt words, it clearly guides the large language model to pay attention to the discrimination between information when generating explanatory information for category nodes at the same level, avoiding the generated explanatory information from being too similar or ambiguous, reducing the blindness and trial-and-error cost of the large language model when generating explanatory information, enabling the model to generate information that meets the requirements faster after receiving clear prompt words, thereby improving the overall generation efficiency and the accuracy and readability of the generated explanatory information. By guiding the large model in this way, the generated explanatory information for different categories can be made as different as possible, thereby expanding the classification interval between different classification categories and reducing the difficulty of classification by the large language model.

[0131] In some embodiments, due to the uncertainty in the generation of the large language model, in order to further improve the accuracy and readability of the explanatory information for each category node in the multi-level classification structure, and at the same time ensure sufficient discrimination between the explanatory information of multiple multi-category nodes at the same level, this embodiment proposes a mechanism for generating explanatory information and evaluation in multiple batches, which is applicable to the case where there are multiple category nodes at the same level in the multi-level classification structure. This mechanism finds the batch of explanatory information that best meets the requirements through multiple generations of explanatory information and batch screening, thereby ensuring the discrimination between the explanatory information of multiple multi-category nodes at the same level. See Figure 3A The flowchart of the steps for generating explanatory information and evaluation in multiple batches shown exemplarily. This embodiment can be implemented through the following steps:

[0132] S301, obtain multiple batches of explanatory information generated multiple times for the multiple category nodes at the same level; each batch includes the explanatory information for all category nodes at this level;

[0133] Multi-batch description information representation: In the multi-level classification structure, when multiple category nodes are included in the same level, the classification categories represented by the multiple category nodes at the same level are regarded as one batch, and multiple batches of description information are generated in multiple times. Each batch contains the description information of all category nodes at this level, so as to ensure that each batch comprehensively covers the category nodes at this level. Among them, the description information of a single category node can be implemented by the method described in the foregoing embodiments.

[0134] For example, assume that a certain third level includes category nodes A, B, C, and D. Then, the category nodes A, B, C, and D can be regarded as one batch to generate multiple batches of description information. The description information of each batch includes E_A, E_B, E_C, and E_D, corresponding to the category nodes A, B, C, and D respectively.

[0135] S302. For the description information in each batch, based on the text similarity between the description information of two-by-two category nodes in the same level, determine the text description discrimination degree of this batch.

[0136] Text similarity is an index used to measure the similarity degree of two text paragraphs or sentences in terms of content, vocabulary, structure, etc. The higher the similarity, the closer the two texts are in expression, and they may contain similar information or viewpoints.

[0137] Text description discrimination is used to represent the degree to which the description information of different category nodes in the same level can clearly distinguish each other in terms of content, expression method, etc. The higher the discrimination degree, the more unique the description information of different category nodes, and the easier it is for the large model to recognize and understand the unique features and differences of each category.

[0138] Based on this, if the average text similarity between all the description information in the same batch is higher, it indicates that the text description discrimination degree of this batch is lower. Therefore, when determining the text description discrimination degree of each batch, for the description information generated in each batch, the method of calculating text similarity can be used first to evaluate the semantic similarity degree between the description information of two-by-two category nodes in the same level. Among them, the text similarity can be determined by calculation methods such as cosine similarity and Jaccard similarity, and the most suitable algorithm can be specifically selected according to actual needs. Next, based on the calculated text similarity, the text description discrimination degree can be determined based on the reciprocal, difference or other calculation methods of indicators that can reflect information difference.

[0139] Regarding the calculation of this text description discrimination degree, it can be determined by calculating the text description discrimination degree between every two description information in the same batch, so as to obtain the average value of the text description discrimination degree of the entire batch. That is, see Figure 3BSchematic diagram for calculating the text description discrimination of an exemplary batch. The text description discrimination can be determined in the following manner:

[0140] S3021. For multiple description messages corresponding one by one to each category node in the same batch, respectively determine the complement of the text similarity between every two description messages with respect to 1 as the text description discrimination between these two description messages;

[0141] Use text similarity algorithms such as cosine similarity, Jaccard similarity, edit distance, etc. to calculate the text similarity between any two description messages in the same batch. The similarity value usually ranges from 0 to 1, where 0 indicates completely dissimilar and 1 indicates completely identical.

[0142] For each pair of description messages, calculate the complement of their text similarity with respect to 1. That is, if the text similarity of two description messages is S (0 ≤ S ≤ 1), then the text description discrimination between them is 1 - S. This complement reflects the degree of difference in content between the two description messages, and store the calculated text description discrimination between each pair of description messages.

[0143] For example, for Figure 3B the N = 4 description messages in batch 2 exemplified in, form n*(n - 1) / 2 = 6 pairs of description messages, and respectively calculate the complement of the text similarity between each pair of description messages with respect to 1 to obtain the text description discrimination Fi of the i-th pair of description messages (i takes values of 1 ≤ i ≤ n*(n - 1) / 2).

[0144] S3022. Sum up the text description discriminations between all different pairs of description messages, and determine the average discrimination based on the sum result and the number of pairs of all different pairs of description messages as the text description discrimination of this batch.

[0145] The number of pairs of all different pairs of description messages is the number of all different pairs of description messages when multiple description messages corresponding one by one to each category node in the same batch are combined pairwise, indicating the total number of pairs of description messages for which the discrimination needs to be calculated.

[0146] Based on the above description, the process of determining the text description discrimination of the entire batch based on the text description discrimination between pairwise description messages provided in this embodiment can be expressed by the following formula:

[0147]

[0148] Among them, text(i) represents the extended description information of the i-th category, and sim(.) represents the text similarity calculation function. That is, for the description information of n different category nodes at a certain level, the similarity is calculated pairwise, and the complement of the similarity with respect to 1 is taken as the text description distinguishability of the description information (text(i), text(j)) of the node pair (i, j). Finally, the average of the text description distinguishabilities of n(n - 1) / 2 description information pairs is taken as the text description distinguishability of this batch.

[0149] In this step, the text description distinguishabilities between all different pairwise description information calculated in step S3021 are summed up, and the obtained summation result reflects the overall difference degree between the description information in the entire batch. Next, the summation result is divided by the logarithm to obtain the average distinguishability as the text description distinguishability of this batch. The average distinguishability reflects the average difference degree between the description information in this batch. The higher the average distinguishability, the more unique the description information in this batch is in terms of content and the easier it is to distinguish.

[0150] Regarding the calculation of this text description distinguishability, it is also possible to calculate the text similarity between every two description information in the same batch, and then sum up and calculate the average value of the calculated text similarities between all different pairwise description information, so as to obtain the average value of the text similarities between the description information in the same batch. Then, the text description distinguishability of this batch is evaluated based on the complement of this average value with respect to 1. It can be understood that other applicable implementation manners can also be adopted for the calculation method of this text description distinguishability, and the present application does not limit this.

[0151] S303, Use the batch of description information with the highest text description distinguishability among the multi-batch description information as the description information of multiple category nodes at the same level.

[0152] After obtaining the multi-batch description information of multiple category nodes belonging to the same level, compare the text description distinguishabilities of each batch, and select the batch of description information with the highest distinguishability as the final description information of multiple category nodes at the same level, thereby ensuring that the finally output description information not only accurately reflects the characteristics of the category nodes, but also has sufficient distinguishability from each other, avoiding information confusion and misunderstanding.

[0153] For example, for Figure 3AThe X category nodes in the same level shown in the example generate 4 batches of description information respectively, and each batch includes description information corresponding to the X category nodes one by one. Taking batch 1 as an example, by calculating the text similarity between the X description information in batch 1, the text description discrimination of batch 1 is determined, and the text description discrimination of batches 2 to 4 is obtained in the same way. From batches 1 to 4, a batch with the highest text description discrimination is selected, such as batch 2, and the X description information included in batch 2 is used as the description information of the X category nodes in this level, and batches 1, 3, and 4 will be discarded.

[0154] In the disclosed embodiment, by obtaining multiple batches of description information generated multiple times for multiple category nodes in the same level, it is ensured that each category node has multiple possible description options, avoiding the one-sidedness or limitations that may be caused by a single generation, determining the text description discrimination of each batch for the description information in the batch, quantitatively evaluating the degree of difference in the description information of different category nodes in the same batch, and using the batch of description information with the highest text description discrimination among the multiple batches of description information as the description information of multiple category nodes in the same level, optimizing the description of each category node, ensuring the maximum difference between the extended descriptions of different categories, and effectively expanding the classification interval between different categories. This optimization not only improves the discriminability of information, but also reduces the difficulty of large language models in subsequent classification tasks.

[0155] In some embodiments, regarding the determination of the classification path description information of each category node in each level as described in the above-mentioned embodiment, in order to more effectively convey the position of the category node in the multi-level classification structure and the path to which it belongs, this embodiment provides a method for generating the classification path description information based on setting symbolic links, which can be implemented by the following steps:

[0156] First, for each node from the root node of the hierarchy to the category node, the description information of the node and the classification category it represents is organized into an information item. That is, the information item includes both the description information of the node and the classification category of the node. If there are Y nodes from the root node to the category node, Y information items can be formed.

[0157] Next, in the order of the paths of the passed nodes, each information item is integrated through a set symbol link to form the classification path description information of the category node. This set symbol is used to connect information items, and at the same time endows the classification path description information with structure and logic. In this way, each node and its classification information can be connected in an orderly manner to form a clear and coherent classification path. The path formed by integrating each information item through the set symbol link constitutes the classification path description information of the category node, which can not only intuitively show the position of the category node in the classification structure, but also provide strong support for understanding and identifying different categories through detailed classification information.

[0158] For example, for Figure 4 an exemplary multi-level classification structure example diagram, the classification path description information of each category node in this classification structure can be generated by organizing with the set symbol "|". For example, the classification path description information of each category node can be expressed as:

[0159] First-level category {one}: {description information};

[0160] ……

[0161] First-level category {two}: {description information}|Second-level category {1}: {description information}|Third-level category {B}: {description information}

[0162] ……

[0163] First-level category {two}: {description information}|Second-level category {3}: {description information}|Third-level category {C}: {description information}|Fourth-level category {d}: {description information}

[0164] Such as Figure 4 There are a total of 13 category nodes, constructing 13 classification path description information. The classification path description information of all category nodes in the multi-level classification structure forms a set, which is used as the description of this multi-level classification structure and is input into the large model so that the large model can accurately understand the file classification system and logic.

[0165] In the embodiments of the present disclosure, through the above method, it is ensured that each category node can be defined by its complete classification path, thus avoiding category confusion or misunderstanding caused by information loss. Through the accumulation of the description information of each node and its classification category in the path, the integrity of the information related to the classification category represented by any category node is fully guaranteed. The obtained classification path description information not only contains rich semantic information, but also significantly improves the distinguishability between different categories, enabling the classification path description information to convey the classification attributes and semantic features of the node while conveying the position of the node.

[0166] In some embodiments, the large model generates, based on the classification path description information for each category node, the category node to which the text sequence corresponding to the file to be classified belongs, and generates a category sequence formed by the nodes passed from the hierarchical root node to the belonging category node as the hierarchical classification result of the text sequence corresponding to the file to be classified.

[0167] Since the large model belongs to a generative model and is affected by different model parameters and different settings, the hierarchical classification result generated by the large model may not exactly match the character descriptions of each classification category in the multi-level classification structure. For example, assuming that the hierarchical classification result in the multi-level classification structure is "Automobile|New Energy|Three-Electric System", the generation of the large model may be "Automobile|New Energy Vehicle|New Energy Three-Electric System". Therefore, in order to ensure that the final hierarchical classification structure of the file to be classified is consistent with the category path in the preset multi-level classification structure, improve the accuracy of classification, and avoid classification errors caused by inaccurate output of the large model, after obtaining the hierarchical classification result output by the large model, the hierarchical classification result can be further optimized by the following steps through a method of performing character similarity matching of classification categories at each level between the hierarchical classification result and the multi-level classification structure:

[0168] a1. Calculate the character matching similarity between the category sequence output by the large model and the nodes of each level in the multi-level classification structure to determine the node path in the multi-level classification structure with the highest character similarity to the category sequence; a2. Use the category sequence represented by all nodes in the node path as the updated hierarchical classification result of the text sequence corresponding to the file to be classified.

[0169] In this embodiment, by comparing the category sequence output by the large model with the preset multi-level classification structure level by level, and calculating the character matching similarity between the classification category represented by the category node and each level category in the category sequence, the node path in the multi-level classification structure with the highest character similarity can be found to correct the classification deviation that may be generated by the large model and ensure that the final classification result conforms to the preset multi-level classification structure.

[0170] To calculate the character matching similarity between two strings (i.e., each level category in the category sequence generated by the large model and the classification category represented by each level category node in the multi-level classification structure), this embodiment uses the shortest edit distance formula to calculate the character matching similarity. The shortest edit distance (also known as the Levenshtein distance) refers to the minimum number of edit operations required to convert one string to another, where the edit operations include inserting, deleting, and replacing characters. Using the method of dynamic programming, the shortest edit distance calculation method is as follows:

[0171]

[0172] Through the above formula, the final D[m][n] can be calculated by dynamic programming, where m and n are the lengths of the two strings s1 and s2 respectively, which is the required minimum edit distance. Then the character matching similarity is 2D[m][n] / (m + n).

[0173] Taking the previous hierarchical classification result in the multi-level classification structure as "Automobile|New Energy|Three-Electric System" and the generation of the large model as "Automobile|New Energy Vehicle|New Energy Three-Electric System" as an example, assuming that for the two strings s1 = "Three-Electric System" and s2 = "New Energy Three-Electric System" at the third level, an example of calculating the minimum edit distance between them is as follows:

[0174] Initialize the matrix D with a size of (len(s1)+1)x(len(s2)+1) and initialize all elements to 0 (note that in fact, the first row and the first column need to be initialized to values from 0 to the lengths of their respective strings, representing the edit distances for converting an empty string to another string). For example, the initialized matrix D is as follows, where () represents the value to be filled:

[0175] [0,1,2,3,4,5,6]

[0176] [1,(),(),(),(),(),()]

[0177] [2,(),(),(),(),(),()]

[0178] [3,(),(),(),(),(),()]

[0179] [4,(),(),(),(),(),()]

[0180] According to the recurrence formula of the minimum edit distance, fill the remaining part of the matrix D. For each position (i,j), we compare s1[i - 1] and s2[j - 1]:

[0181] If they are equal, then D[i][j] = D[i - 1][j - 1] (no editing is required).

[0182] If they are not equal, then D[i][j] = min(D[i - 1][j]+1, D[i][j - 1]+1, D[i - 1][j - 1]+1) (corresponding to deletion, insertion, and replacement operations respectively).

[0183] The filled matrix D is as follows:

[0184] [0,1,2,3,4, 5,6]

[0185] [1, 1, 2, 3, 4, X1, X2] # X1 and X2 are values to be calculated [2, 2, X3, X4, X5, X6, X7] # X3, X4, X5, X6, X7 are values to be calculated [3, 3, (), (), (), X8, X9] # X8 and X9 are values to be calculated, depending on previous values [4, 4, (), (), (), (4 + X), (5 + X)] # When i = 4, j = 6, due to the same "system", only consider the insertion of "new energy". X is the minimum edit distance calculated previously

[0186] Regarding the above X1 - X9 to be calculated:

[0187] X1 = min(2 + 1, 3 + 1, 4 + 1) = min(3, 4, 5) = 3 (Delete s1[1] = 'electric' or insert 'new' before s1 or replace 'three' with 'new')

[0188] X2 = min(3 + 1, X1 + 1, 4 + 1) = min(4, 4, 5) = 4 (Continue the above operations)

[0189] X3 = min(1 + 1, 2 + 1, 3 + 0) = min(2, 3, 3) = 2 ('three' = 'three', no editing required, but consider the previous edit distance)

[0190] X4 = min(2 + 1, X3 + 1, 3 + 1) = min(3, 3, 4) = 3

[0191] X5 = min(3 + 1, X4 + 1, 4 + 1) = min(4, 4, 5) = 4

[0192] X6 = min(4 + 1, X5 + 1, 4 + 0) = min(5, 5, 4) = 4 ('system' ='system', no editing required)

[0193] X7 = min(4 + 1, X6 + 1, 5 + 1) = min(5, 5, 6) = 5

[0194] X8 = min(3 + 1, X7 + 1, 5 + 0) = min(4, 6, 5) = 4 ('system' ='system', no editing required)

[0195] X9 (i.e., D[4][6]) = min(4 + 1, X8 + 1, 4 + X) = min(5, 5, 4 + X), but here X is actually the minimum edit distance when calculating up to D[3][5]. Since the prefixes of "Three Electric Systems" and "New Energy Three Electric Systems" do not match, we need to backtrack to D[0][3] (i.e., 3, indicating that it takes 3 edits to convert an empty string to "New Energy") and then add the edit distance to convert "System" to "Three Electric Systems" (which is 0 here because they are the same). However, in this specific case, it is known that the minimum edit distance from D[3][4] to D[4][6] considers the insertion of "New Energy", so X can be regarded as 3 (to "New Energy") plus 0 (because "System" already matches). But through recursion, it is known that X8 is 4 (considering the previous edits), so X9 = min(5, 5, 4 + 3) = 4 (but actually, since we directly jump from D[3][5] to D[4][6] considering the insertion of "New Energy" and know that "System" is the same, we can directly get D[4][6] = 5. The explanation here is to show the recursive process). However, a more accurate recursion should be calculated step by step based on the previous values rather than directly jumping. But in this example, for simplicity, it can be directly stated that D[4][6] = 5 is the correct result obtained through recursion.

[0196] Finally, the value of D[len(s1)][len(s2)] is the minimum edit distance required to convert s1 to s2. In this example, the value of D[4][6] is 5, indicating that it takes 5 edit operations to convert s1 to s2 (specifically, inserting "New Energy"), and thus the character matching similarity can be calculated using the minimum edit distance.

[0197] Through this method, the character matching similarity between each node in the category sequence output by the large model and the corresponding node in the multi-level classification structure can be dynamically calculated, so as to find the most matching node path and update the hierarchical classification result, in order to select the category with the highest character matching similarity to the document classification category generated by the large model from the multi-level classification mechanism as the hierarchical classification result of the final model.

[0198] Corresponding to the embodiment of the foregoing document classification method, see Figure 5 As shown, the present application also provides an embodiment of a document classification device, and the device includes:

[0199] A classification preprocessing module 501, configured to determine the classification path description information of each category node in each level for a preset multi-level classification structure; wherein, the classification path description information is determined based on all the nodes passed from the hierarchical root node to this category node and the description information of the classification category represented by each node;

[0200] A file parsing module 502, configured to obtain a text sequence obtained by parsing a file to be classified;

[0201] A large model classification module 503, configured to input the text sequence and the classification path description information of each category node into a large model in an associated manner, and obtain a hierarchical classification result output by the large model after reasoning based on the classification path description information of each category node, where the result indicates the hierarchical classification category to which the file to be classified belongs in the multi-level classification structure.

[0202] In some embodiments, the apparatus further includes:

[0203] For each level of the multi-level classification structure, construct and generate prompt words; wherein, the generated prompt words are used to prompt the large language model to enhance the discrimination between the description information of each category node at the same level;

[0204] Input each category node at the same level in the multi-level classification structure and the generated prompt words into the large language model together, and obtain the description information of each category node at the same level output by the large language model.

[0205] In some embodiments, when there are multiple category nodes at the same level in the multi-level classification structure, the apparatus further includes:

[0206] A batch generation module, configured to obtain multiple batches of description information generated multiple times for the multiple category nodes at the same level; each batch includes the description information of all category nodes at this level;

[0207] A discrimination calculation module, configured to, for the description information in each batch, determine the text description discrimination of this batch based on the text similarity between the description information of any two category nodes at the same level;

[0208] A screening module, configured to use the batch of description information with the highest text description discrimination among the multiple batches of description information as the description information of the multiple category nodes at the same level.

[0209] In some embodiments, the discrimination calculation module is specifically configured to:

[0210] For the multiple pieces of description information corresponding to each category node in the same batch, respectively determine the complement of the text similarity between every two pieces of description information with respect to 1 as the text description discrimination between these two pieces of description information;

[0211] Sum up the text description discriminations between all different pairs of description information, and determine the average discrimination based on the summation result and the logarithm of all different pairs of description information as the text description discrimination of this batch.

[0212] In some embodiments, the classification preprocessing module is specifically configured to:

[0213] For each node passed from the hierarchical root node to this category node, organize the node and the description information of the classification category it represents into an information item;

[0214] According to the path order of the passed nodes, integrate each information item through a set symbol link to form the classification path description information of this category node.

[0215] In some embodiments, based on the classification path description information of each category node, the large model determines the category node to which the text sequence corresponding to the file to be classified belongs, and generates a category sequence formed by the nodes passed from the hierarchical root node to the belonging category node as the hierarchical classification result of the text sequence corresponding to the file to be classified; The apparatus further includes:

[0216] Calculate the character matching similarity of the category nodes at each level between the category sequence output by the large model and the multi-level classification structure to determine the node path in the multi-level classification structure with the highest character similarity to the category sequence;

[0217] Use the category sequence represented by all nodes in the node path as the updated hierarchical classification result of the text sequence corresponding to the file to be classified.

[0218] In some embodiments, the large model classification module is specifically configured to:

[0219] Construct a classification prompt word for document classification according to the text sequence and the classification path description information of each category node; The classification prompt word at least further includes the description information of the classification task executed by the large model and the corresponding classification output requirements;

[0220] Input the classification prompt word into the large model so that the large model generates the hierarchical classification result of the text sequence corresponding to the file to be classified according to the classification prompt word.

[0221] In some embodiments, the file parsing module is specifically configured to:

[0222] When the file to be classified includes a picture, obtain at least one of the text information included in the picture, the file name, and the text description of the content included in the picture as the text sequence of the picture;

[0223] When the file to be classified includes audio or video, obtain the text information obtained by converting the speech included in the audio or video as the text sequence of the audio or video.

[0224] For the implementation processes of the functions and roles of each unit in the above device, please refer to the implementation processes of the corresponding steps in the above method for details, which will not be elaborated here.

[0225] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0226] This application embodiment also provides an electronic device. The structural schematic diagram of the electronic device is as Figure 6 shown. The electronic device 600 includes at least one processor 601, a memory 602, and a bus 603. At least one processor 601 is electrically connected to the memory 602; the memory 602 is configured to store at least one computer-executable instruction, and the processor 601 is configured to execute the at least one computer-executable instruction, so as to execute the steps of any file classification method provided in any embodiment or any optional implementation manner of this application.

[0227] Furthermore, the processor 601 can be an FPGA (Field-Programmable Gate Array) or other devices with logical processing capabilities, such as an MCU (Microcontroller Unit) or a CPU (Central Process Unit).

[0228] This application embodiment also provides another readable storage medium, storing a computer program, which is used to implement the steps of any file classification method provided in any embodiment or any optional implementation manner of this application when executed by a processor.

[0229] The readable storage medium provided by the embodiments of the present application includes, but is not limited to, any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards, or optical cards. That is, the readable storage medium includes any medium that stores or transmits information in a form readable by a device (such as a computer).

[0230] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. In addition, the processes depicted in the figures are not necessarily in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0231] The foregoing is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for classifying documents, characterized in that, The method includes: For a preset multi-level classification structure, determining the classification path description information of each category node in each level; wherein, the classification path description information is determined based on all the nodes passed from the root node of the level to this category node and the description information of each node representing the classification category; Obtaining the text sequence obtained by parsing the file to be classified; Associating and inputting the text sequence with the classification path description information of each category node into a large model, and obtaining the hierarchical classification result output by the large model after reasoning based on the classification path description information of each category node, and this result indicates the hierarchical classification category to which the file to be classified belongs in the multi-level classification structure.

2. The method according to claim 1, characterized in that, The method further includes: For each level of the multi-level classification structure, constructing and generating prompt words; wherein, the generated prompt words are used to prompt the large language model to enhance the distinguishability between the description information of each category node in the same level; Inputting each category node in the same level of the multi-level classification structure and the generated prompt words into the large language model together, and obtaining the description information of each category node in the same level output by the large language model.

3. The method according to claim 1 or 2, characterized in that, In the case where there are multiple category nodes in the same level of the multi-level classification structure, the method further includes: Obtaining multiple batches of description information generated multiple times for the multiple category nodes in the same level; each batch includes the description information of all category nodes in this level; For the description information in each batch, based on the text similarity between the description information of two category nodes in the same level, determining the text description distinguishability of this batch; Taking the batch of description information with the highest text description distinguishability among the multiple batches of description information as the description information of the multiple category nodes in the same level.

4. The method according to claim 3, wherein The determining the text description distinguishability of this batch based on the text similarity between the description information of two category nodes in the same level includes: For the multiple description information corresponding to each category node in the same batch, respectively determining the complement of the text similarity between every two description information with respect to 1 as the text description distinguishability between these two description information; Summing up the text description distinguishability between all different pairs of description information, and determining the average distinguishability according to the summation result and the logarithm of all different pairs of description information as the text description distinguishability of this batch.

5. The method according to claim 1, characterized in that, The determining the classification path description information of each category node in each level includes: For each node passed from the root node of the level to this category node, organizing the node and the description information of the classification category it represents into an information item; According to the path order of the passed nodes, integrating each information item through a set symbol link to form the classification path description information of this category node.

6. The method according to claim 1, characterized in that, The large model determines the category node to which the text sequence corresponding to the file to be classified belongs based on the classification path description information of each category node, and generates a category sequence formed by the nodes passed from the root node of the level to the belonging category node as the hierarchical classification result of the text sequence corresponding to the file to be classified; The method further includes: Calculate the character matching similarity of the category sequence output by the large model with the category nodes at each level of the multi-level classification structure, and determine the node path in the multi-level classification structure with the highest character similarity to the category sequence; Use the category sequences represented by all nodes in the node path as the updated hierarchical classification result of the text sequence corresponding to the file to be classified.

7. The method according to claim 1, characterized in that The step of associatively inputting the text sequence and the classification path description information of each category node into the large model includes: Construct a classification prompt word for document classification based on the text sequence and the classification path description information of each category node; the classification prompt word at least further includes the description information of the classification task performed by the large model and the corresponding classification output requirements; Input the classification prompt word into the large model so that the large model generates the hierarchical classification result of the text sequence corresponding to the file to be classified according to the classification prompt word.

8. The method according to claim 1, characterized in that, The step of obtaining the text sequence obtained by parsing the file to be classified includes: When the file to be classified includes a picture, obtain at least one of the text information included in the picture, the file name, and the text description of the content included in the picture as the text sequence of the picture; When the file to be classified includes audio or video, obtain the text information obtained by converting the speech included in the audio or video as the text sequence of the audio or video.

9. A file classification device, characterized in that, The device includes: A classification preprocessing module, configured to determine the classification path description information of each category node in each level for a preset multi-level classification structure; wherein, the classification path description information is determined based on all nodes passed from the hierarchical root node to the category node and the description information of the classification category represented by each node; A file parsing module, configured to obtain the text sequence obtained by parsing the file to be classified; A large model classification module, configured to associatively input the text sequence and the classification path description information of each category node into the large model, and obtain the hierarchical classification result output by the large model after reasoning according to the classification path description information of each category node, and this result indicates the hierarchical classification category to which the file to be classified belongs in the multi-level classification structure.

10. An electronic device, characterized in that, It includes: A memory and a processor; The memory is used to store computer programs; The processor is used to call the computer program to implement the method according to any one of claims 1-9.

11. A readable storage medium, on which a computer program is stored, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-9.

Citation Information

Cited By

  • Document recommendation problem generation method and device, equipment and storage medium

    CN120764514A

  • Heterogeneous data fusion method and device based on bridging technology, equipment and medium

    CN121834666A

  • Heterogeneous data fusion method and device based on bridging technology, equipment and medium

    CN121834666B