Document Classification Method, Apparatus, Electronic Device, and Storage Medium

By dividing documents into text segments and determining document categories using classification models, the problem of time-consuming and labor-consuming manual document classification is solved, and an automated and high-speed document classification method is realized.

CN117150010BActive Publication Date: 2025-07-08BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311061623.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-22
Publication Date
2025-07-08
Estimated Expiration
2043-08-22

AI Technical Summary

Technical Problem

In the prior art, document classification requires manual operation, which takes a long time and has a large workload, making it difficult to efficiently classify data.

Method used

By dividing the text into multiple text segments based on the catalog in the document, using the classification model to determine the text segment summary and category, and determining the final category of the document based on the text segment summary and category, a multi-stage processing process includes text segments, directory-level candidate category determination and confidence calculation to finally determine the document category.

Benefits of technology

Automatic classification of documents is realized, classification efficiency is improved, labor and cost are reduced, and classification accuracy and efficiency are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117150010B_ABST
    Figure CN117150010B_ABST
Patent Text Reader

Abstract

The present disclosure provides a document classification method, apparatus, electronic device, and storage medium, which relate to the field of artificial intelligence technology, and particularly to the field of natural language processing. The specific implementation solution is as follows: divide the body text in the document into multiple text segments according to the table of contents in the document; for each of the multiple text segments, determine the text segment summary and the text segment category; determine multiple first candidate categories and the confidence level of each first candidate category according to the text segment summaries of the multiple text segments; determine multiple second candidate categories and the confidence level of each second candidate category according to the text segment categories of the multiple text segments; and determine the category of the document according to the multiple first candidate categories, the confidence level of each first candidate category, the multiple second candidate categories, and the confidence level of each second candidate category.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly to the field of natural language processing. More specifically, the present disclosure provides a document classification method, apparatus, electronic device, storage medium, and computer program product. Background Art

[0002] Due to requirements such as regulatory requirements, value requirements, and governance requirements, data classification is needed, such as classifying documents. In practical applications, manual document classification is required, which is time-consuming and labor-intensive. Summary of the Invention

[0003] The present disclosure provides a document classification method, apparatus, electronic device, storage medium, and computer program product.

[0004] According to one aspect of the present disclosure, there is provided a document classification method, including: dividing the body text in a document into a plurality of text segments according to the table of contents in the document; for each text segment among the plurality of text segments, determining a text segment summary and a text segment category; determining a plurality of first candidate categories and the confidence level of each first candidate category according to the text segment summaries of the plurality of text segments; determining a plurality of second candidate categories and the confidence level of each second candidate category according to the text segment categories of the plurality of text segments; and determining the category of the document according to the plurality of first candidate categories, the confidence level of each first candidate category, the plurality of second candidate categories, and the confidence level of each second candidate category.

[0005] According to another aspect of the present disclosure, there is provided a document classification apparatus, including: a dividing module, a first determining module, a second determining module, a third determining module, and a fourth determining module. The dividing module is configured to divide the body text in a document into a plurality of text segments according to the table of contents in the document. The first determining module is configured to determine a text segment summary and a text segment category for each text segment among the plurality of text segments. The second determining module is configured to determine a plurality of first candidate categories and the confidence level of each first candidate category according to the text segment summaries of the plurality of text segments. The third determining module is configured to determine a plurality of second candidate categories and the confidence level of each second candidate category according to the text segment categories of the plurality of text segments. The fourth determining module is configured to determine the category of the document according to the plurality of first candidate categories, the confidence level of each first candidate category, the plurality of second candidate categories, and the confidence level of each second candidate category.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method provided by the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, which implements the method provided by the present disclosure when executed by a processor.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a schematic diagram of an application scenario of a document classification method and apparatus according to an embodiment of the present disclosure;

[0012] Figure 2 is a schematic flowchart of a document classification method according to an embodiment of the present disclosure;

[0013] Figure 3A and Figure 3B is a schematic principle diagram of a document classification method according to an embodiment of the present disclosure;

[0014] Figure 4 is a schematic structural block diagram of a document classification apparatus according to an embodiment of the present disclosure; and

[0015] Figure 5 is a structural block diagram of an electronic device for implementing the document classification method according to an embodiment of the present disclosure. Detailed Embodiments

[0016] The following describes exemplary embodiments of the present disclosure with reference to the drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding and should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0017] In some embodiments, manual classification of documents is required. It can be understood that documents in various industries involve a wide variety of contents, which requires classification staff to have various types of professional knowledge understanding abilities. If the document content is long, the staff needs to spend a long time reading the document and then gradually summarize it to determine the classification result. In addition, the number of classification categories in a certain industry is large, for example, there are 100. The staff needs to match the content of the currently processed document with these 100 classification categories respectively. It can be seen that the manual classification method takes a long time and has a large workload.

[0018] Embodiments of the present disclosure aim to provide a document classification method, which can replace manual operation for automatic classification of documents, thereby saving labor costs and improving classification efficiency.

[0019] The technical solutions provided by the present disclosure will be elaborated in detail below in conjunction with the accompanying drawings and specific embodiments.

[0020] Figure 1 is a schematic diagram of an application scenario of a document classification method and apparatus according to an embodiment of the present disclosure.

[0021] It should be noted that Figure 1 The example shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0022] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0023] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0024] The server 105 may be a server that provides various services, such as a background management server (only an example) that supports the websites browsed by users using the terminal devices 101, 102, 103. The background management server can analyze and process data such as received user requests, and feedback the processing results (such as the document category determined according to the document, the document abstract, etc.) to the terminal device.

[0025] It should be noted that the document classification method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the document classification device provided by the embodiments of the present disclosure can generally be set in the server 105. The document classification method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105. Correspondingly, the document classification device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105.

[0026] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0027] Figure 2 is a schematic flowchart of the document classification method according to the embodiments of the present disclosure.

[0028] As Figure 2 shown, the document classification method 200 may include operations S210 to S250.

[0029] In operation S210, according to the table of contents in the document, the main text in the document is divided into multiple text segments.

[0030] For example, the document may include a table of contents and a main text. The number of tables of contents is one or more, and each table of contents corresponds to a part of the main text.

[0031] For example, for each table of contents, the main text corresponding to the table of contents can be divided into multiple text segments, which can be divided according to a predetermined character length. For example, a text segment is divided every predetermined character length. The predetermined character length can be 512, 1024, etc.

[0032] For another example, according to the predetermined character length and punctuation marks in the table of contents text, the table of contents text can be divided into multiple text segments. The length of each text segment is less than or equal to the predetermined character length, and the last character of the text segment is a punctuation mark. For example, division can be performed at the last punctuation mark before the 512th character each time. In addition, adjacent two text segments can have overlapping text to ensure semantic integrity. For example, the last 100 characters of the first text segment are the same as the first 100 characters of the second text segment.

[0033] For another example, the main text can be divided into a predetermined number of text segments in an equal division manner.

[0034] In operation S220, for each of the multiple text segments, determine a text segment summary and a text segment category.

[0035] For example, the text segment can be input into a classification model, and the classification model outputs the text segment category. The text segment can be input into a text generation model, and the text generation model generates the text segment summary. This embodiment does not limit the classification model and the text generation model.

[0036] In operation S230, based on the text segment summaries of the multiple text segments, determine multiple first candidate categories and the confidence of each first candidate category.

[0037] For example, the text segment summaries can be combined in sequence, and the combined text is input into a classification model, and the classification model outputs the first candidate category and the confidence.

[0038] In operation S240, based on the text segment categories of the multiple text segments, determine multiple second candidate categories and the confidence of each second candidate category.

[0039] For example, the second candidate category can be determined based on the importance parameter of the text segment. The specific processing method will be described in detail below and will not be elaborated here. Another example is that the importance parameter can be ignored, and the text segment category with a higher frequency is determined as the second candidate category, and the confidence is determined based on the frequency of the second candidate category.

[0040] In operation S250, based on the multiple first candidate categories, the confidence of each first candidate category, the multiple second candidate categories, and the confidence of each second candidate category, determine the category of the document.

[0041] For example, the set composed of the multiple first candidate categories and the set composed of the multiple second candidate categories may have an intersection, and the candidate category in the intersection can be determined as the directory category.

[0042] If the document includes only one directory, the directory category of this directory is the category of the document. If the document includes multiple directories, the category of the document can be determined based on the directory categories of the multiple directories. For example, the frequencies of the directory categories of the multiple directories can be counted, and the directory category with a higher frequency is determined as the document category.

[0043] The embodiments of the present disclosure understand the content of the document according to the directory of the document. First, for the text corresponding to each directory, the text is divided into multiple text segments, and then the text segment summary and the text segment category of each text segment are determined. Then, based on these text segment summaries and text segment categories, the first candidate category and the second candidate category are respectively determined, and then the category of the document is determined from the two candidate categories. This method can replace manual document classification, realize automatic document classification, improve classification efficiency, reduce the labor volume of staff, and reduce classification costs.

[0044] Figure 3A and Figure 3B is a schematic diagram of a document classification method according to an embodiment of the present disclosure.

[0045] As Figure 3A shown, in this embodiment, a document may include multiple directories and multiple main texts. The directories and the main texts correspond one by one, and the main text is the specific content under the directory. For example, the document includes directory A 3011, directory B 3012, and directory C 3013, and the three directories correspond to main text A 3021, main text B 3022, and main text C 3023 respectively.

[0046] Taking the processing process of main text A 3021 as an example, main text A 3021 can be divided into multiple text segments, such as being divided into text segment A 3031, text segment B 3232, and text segment C 3033.

[0047] Input text segment A and the first prompt information template into a predetermined model, and the predetermined model outputs a text segment summary 304 and a text segment category 305. Similar processing is performed on text segment B 3232 and text segment C 3033 as on text segment A 3231, so as to obtain multiple text segment summaries 304 and text segment categories 305. The first spliced text 306 can be determined based on multiple text segment summaries 304, and then the target summary 307 and the first candidate category 308 can be determined based on the first spliced text 306. The second candidate category 309 can also be determined according to multiple text segment categories 305. Then, the directory category 310 is determined based on the first candidate category 308 and the second candidate category 309.

[0048] As Figure 3B shown, a similar processing process as that of main text A 3021 is adopted for main text B 3022 and main text C 3023, so as to obtain multiple directory summaries 307 and multiple directory categories 310. Next, the second spliced text 311 can be determined based on multiple directory summaries 307, and then the document summary 312 and the third candidate category 313 can be determined. The fourth candidate category 314 can be determined based on multiple directory categories. Then, the category 315 of the document can be determined based on the third candidate category 313 and the fourth candidate category 314.

[0049] This embodiment briefly introduces the segmented classification method. In the following, each process in the document classification method will be described in detail in combination with other embodiments.

[0050] In this embodiment, a document may include multiple directories and multiple main texts. The directories and the main texts correspond one by one, and the main text is the specific content under the directory. This embodiment may include the following stages.

[0051] In the first stage, configure a prompt information template.

[0052] For example, multiple prompt message templates can be pre-configured. Each prompt message template can include general prompt information, and each prompt message template can also include non-general prompt information according to actual needs. The general prompt information can include multiple sub-information. The following takes the general prompt information as an example for illustration.

[0053] For example, the general prompt information can include the document name of the document and the directory name of the directory. Sometimes, the general classification category can be determined from the document name and the directory name. Therefore, the document name and the directory name are described in the prompt message template to make the information input into the model have a context semantic environment and ensure the processing effect of the model. The document name and the directory name can be automatically extracted and added to the prompt message template.

[0054] For example, the general prompt information can include multiple reference categories related to the scenario. Common classification categories in the industry can be added to the prompt to enable the predetermined model to better determine the output category information.

[0055] For example, the general prompt information can include category constraint information, which characterizes that the predetermined model is used to determine category information based on multiple reference categories. The category information includes at least one of the text segment category, the first candidate category, and the third candidate category. For example, the following content can be added to the prompt message template: "Judge the relevance of the following text to these 20 common classifications in the industry. The relevance is judged using a decimal number greater than or equal to 0 and less than or equal to 1, with two decimal places reserved after the decimal point. If there are two identical relevance decimal numbers, it is recommended to re-judge the relative relevance between the two classifications and re-generate the relevance. Finally, sort the results in descending order of relevance and output the TOP 3 classification categories and the relevance of each classification category."

[0056] For example, the general prompt information can include a quantity threshold representing the maximum number of characters in the abstract. The abstract includes at least one of the text segment abstract, the directory abstract, and the document abstract. For example, the following content can be added to the prompt message template: "The abstract needs to be concise and to the point, limited to within 50 words."

[0057] For example, the general prompt information can include processing order constraint information, which characterizes that the predetermined model is used to generate an abstract based on the category information. For example, the following content can be added to the prompt message template: "When generating the abstract, it is necessary to summarize from the perspective of the TOP 3 classification categories; of course, if there are no relevant classification categories for the text, a direct summary can be made. It can be seen that this general prompt information will enable the predetermined model to first determine the classification information and then determine the abstract, so that the abstract is a targeted summary from the perspective of classification, thereby improving the accuracy of the abstract.

[0058] In addition, the sentence pattern of the general prompt information can be consistent with the sentence pattern of the training data used to train the predetermined model, so as to better exert the effect of the predetermined model.

[0059] In the second stage, based on the table of contents, the corresponding text is divided into multiple text segments.

[0060] For example, divide according to a predetermined character length, and at the same time, there are some overlapping characters between two adjacent text segments to ensure semantic integrity. In addition, punctuation marks are also considered during the division process.

[0061] In some embodiments, when the character length of the text corresponding to the table of contents is greater than or equal to the predetermined length, the second stage can be entered. If the character length is less than the predetermined length, no division is required, but the text corresponding to the table of contents is directly used as a text segment and the subsequent third stage is entered.

[0062] In the third stage, the text segment summary and the text segment category are determined.

[0063] For example, for each text segment, a first prompt information template corresponding to the text segment can be determined from multiple prompt information templates. For example, the first prompt information template can be selected according to the process identifier, and the first prompt information template can include general prompt information. The process identifier indicates that the current process is to determine the text segment summary and the text segment category. Then, the first prompt information is determined according to the text segment and the first prompt information template. For example, the text segment and the first prompt information template are combined into the first prompt information. Then, the first prompt information is input into the predetermined model, and the predetermined model outputs the text segment summary and the text segment category. The predetermined model can include a text generation model or a classification model, and a large language model (LLM) can be used as the predetermined model. The structure and working principle of the predetermined model are not limited in this embodiment. The first prompt information template can improve the generalization ability and processing effect of the pre-trained predetermined model, so as to obtain a relatively accurate text segment summary and text segment category.

[0064] In the fourth stage, the table of contents summary and the table of contents category of the table of contents are determined based on multiple text segments. The fourth stage can include the following multiple sub-stages.

[0065] In the first sub-stage, the table of contents summary and the first candidate category can be determined first. The table of contents summary can represent the total summary corresponding to multiple text segments under the table of contents.

[0066] For example, multiple text segment summaries of multiple text segments can be concatenated to obtain a first concatenated text. Then, a second prompt information template corresponding to the first concatenated text is determined from multiple prompt information templates. For example, the second prompt information template can be selected according to a process identifier, and the process identifier indicates that the current process being processed is to determine a catalog summary and a first candidate category. Next, according to the first concatenated text and the second prompt information template, second prompt information is determined. For example, the first concatenated text and the second prompt information template are combined into the second prompt information. Then, the second prompt information is input into a predetermined model, and the predetermined model outputs a catalog summary. In addition, the predetermined model can also output multiple first candidate categories and the confidence levels of each first candidate category. In this embodiment, the number of characters in the text segment summary is small, and concatenating multiple text segment summaries can obtain information with relatively complete semantics. Based on this information, the catalog summary and the first candidate category can be accurately determined. In addition, the second prompt information template can improve the generalization ability and processing effect of the pre-trained predetermined model.

[0067] The second prompt information template can include general prompt information and can also include non-general prompt information. The non-general prompt information in the second prompt information template can include importance constraint information, and the importance constraint information represents: the relative importance relationship between multiple text segment summaries, and the relative importance relationship is related to the positions of multiple text segments in the catalog content text. The predetermined model generates a catalog summary based on the importance constraint information. For example, the following content is added to the second prompt information template: "Pay special attention to the first paragraph when generating a summary summary, and then pay attention to the last paragraph." It should be noted that since the second prompt information is the text segment summary of all text segments under a catalog and has relatively complete semantics, according to the writing habit of Chinese, the first paragraph in a complete semantic paragraph is the most important, followed by the last paragraph. Therefore, setting the importance constraint information can improve the accuracy of the catalog summary. In addition, there are duplicate characters between two adjacent text segments. Therefore, the non-general prompt information in the second prompt information template can also include semantic overlap information. For example, the following content can be included in the second prompt information target: "There is generally semantic overlap between the endings and beginnings of adjacent paragraphs."

[0068] In the second sub-phase, a second candidate category can be determined.

[0069] For example, according to the position information of multiple text segments in the catalog content text, the importance parameter of each text segment is determined. Then, according to the importance parameter of each text segment and the text segment category of each text segment, the confidence level of each text segment category is determined. After that, according to the ranking of the confidence levels of each text segment category, multiple second candidate categories are determined.

[0070] For example, the importance parameter represents the importance of a text segment in the main text of a directory. A text segment has a specific weight corresponding to it. With the weight as the importance parameter, the weight w1 of the first text segment, the weight w2 of the last text segment, and the weight w3 of each remaining text segment can be different. For example, the weight w1 is greater than the weight w2 and greater than the weight w3, for example, w1:w2:w3=1.2:1.1:1, and the sum of the weights of all text segments can be 1. Taking the main text of a directory as an example, different text segments can belong to the same text segment category. If a text segment category corresponds to the first text segment and the second text segment, the sum of the weights of the text segment category is w1+w3, and the sum of the weights is the confidence of the text segment category. Some text segment categories with greater confidence can be used as the second candidate category. This embodiment determines the importance of the text segment based on the position information of multiple text segments in the directory content text, and then determines the confidence of the text segment category based on the importance, so that the second candidate category can be accurately determined from the categories of multiple texts.

[0071] In the third sub-stage, the catalog category may be determined based on the first candidate category and the second candidate category.

[0072] For example, the evaluation value of the candidate category can be determined based on the weight. The first weight of the first candidate category and the second weight of the second candidate category can be preconfigured, the first weight is, for example, 0.6, and the second weight is, for example, 0.4, and the first product of the confidence of each first candidate category and the first weight can be calculated. The second product of the confidence of each second candidate category and the second weight can be calculated. The set consisting of multiple first candidate categories and the set consisting of multiple second candidate categories can have an intersection. For the candidate categories in the intersection, the sum of the first product and the second product of the candidate category is used as the evaluation value of the candidate category. For the first candidate categories outside the intersection, the first product is used as the evaluation value. For the second candidate categories outside the intersection, the second product is used as the evaluation value. Then the evaluation values ​​are sorted in order from large to small, and the top candidate categories are used as catalog categories.

[0073] In the fifth stage, based on the directory summaries and directory categories of the multiple directories, the document summaries and document categories are determined. The fifth stage may include the following multiple sub-stages.

[0074] In the first sub-phase, based on the directory summaries of multiple directories, a document summary and multiple third candidate categories can be determined. The first sub-phase in this fifth phase can refer to the process of determining the first candidate category in the first sub-phase of the fourth phase above. For example, the directory summaries of at least one directory can be concatenated to obtain a second concatenated text, then a third prompt information template corresponding to the second concatenated text can be determined from multiple prompt information templates, and then based on the second concatenated text and the third prompt information template, third prompt information can be determined. After that, the third prompt information is input into a predetermined model to obtain a document summary and multiple third candidate categories.

[0075] It should be noted that in other embodiments, the above first sub-phase can also adopt other solutions. For example, the second concatenated text can be used as the document summary.

[0076] It should be noted that the third prompt information template can include general prompt information and non-general prompt information. The non-general prompt information in the third prompt information template can include first auxiliary information and / or second auxiliary information. The first auxiliary information indicates that the directory summary of the first directory is summary content, and the second auxiliary information indicates that the directory summary of the last directory is summary content. For example, the content of the first directory and the last directory is sometimes summary content such as abstracts and introductions. Therefore, it is described in the third prompt information template to improve the accuracy of the document summary. In addition, the document name, directory name, directory summary of each directory, and directory category of each directory can be arranged in an orderly manner according to a predetermined format for the predetermined model to understand semantic information.

[0077] In the second sub-phase, based on the directory categories of multiple directories, multiple fourth candidate categories can be determined. The second sub-phase in this fifth phase can refer to the process of determining the second candidate category in the second sub-phase of the fourth phase above. The processing process can be the same, except for the data processed. For example, based on the position information of multiple directories in the document, the importance parameter of each directory can be determined, and then based on the importance parameter of each directory and the directory category of each directory, the confidence level of each directory category can be determined. After that, based on the ranking of the confidence levels of each directory category, multiple fourth candidate categories can be determined.

[0078] In the third sub-phase, based on multiple third candidate categories and multiple fourth candidate categories, the category of the document can be determined. The third sub-phase in this fifth phase can refer to the process of determining the second candidate category in the third sub-phase of the fourth phase above. The processing process can be the same, except for the data processed. For example, the evaluation value of each candidate category can be determined, and then the candidate categories can be ranked according to the evaluation value, and several candidate categories with relatively high rankings can be used as the category of the document.

[0079] It can be seen that in the fifth stage, the document category is comprehensively determined from two dimensions: the table of contents abstract and the table of contents category. Therefore, a relatively accurate document category can be obtained.

[0080] Through the above first stage to the fifth stage, a classification result can be obtained, and then the classification result can be presented to the user.

[0081] It should be noted that the classification result can include the category of the document. By directly informing the user of the category of the document, the user sometimes does not trust the processing result of the pre - defined model. Therefore, the confidence level of each category of the document and the document abstract can also be output, so as to improve the interpretability, reliability, and sense of security of the pre - defined model. In this way, the user not only knows the classification result of the pre - defined model but also the classification logic of the pre - defined model. The user can check some documents to determine whether the classification result of the pre - defined model and the document abstract are accurate. If they are accurate, the user can no longer view the document content subsequently and directly trust the output result of the pre - defined model.

[0082] It should be noted that using the prompt information template in the above first stage to the fifth stage can improve the processing effect of the pre - defined model. In some embodiments, the prompt information template can be omitted, and the model can directly process information such as text segments, text segment summaries, and table of contents summaries.

[0083] It should be noted that in the above - mentioned embodiments, the prompt information template is first configured, and then the data to be processed is combined with the prompt information template to form prompt information, and the prompt information is input into the pre - defined model to obtain the information output by the pre - defined model. In actual applications, the prompt information template is sometimes not completely accurate.

[0084] Taking the first prompt information template and the second prompt information template as an example, for example, the optimized prompt information template can be evaluated. For example, it can be evaluated using N tables of contents, where N is an integer greater than or equal to 1. For example, when N is 50, it is determined through evaluation that the information output by the pre - defined model for 40 of the tables of contents is relatively accurate, but the pre - defined model cannot output accurate information for the other 10 tables of contents. Therefore, the pre - defined model can be optimized.

[0085] During the model optimization process, the 50 pieces of information evaluated above can be used as training samples. Among them, 40 samples can be directly used without modification. For the other 10 samples, since the accuracy of the output result of the pre - defined model does not meet the requirements, they can be modified based on the model output for the incorrect parts.

[0086] However, the pre - trained large - model is learned from a vast amount of general knowledge corpus, and 50 samples are not sufficient to train the pre - defined model. For example, when using SFT (supervised fine - tuning of all parameters) for downstream task training, the training effect is poor under the condition of 50 samples, and the capabilities of the pre - defined model in other general knowledge aspects also degrade.

[0087] Therefore, AdaLORA local parameter fine-tuning can be used for model training of downstream tasks. When performing AdaLORA fine-tuning training, instead of updating all parameters, only a very small number of parameters are updated. At this time, a small number of samples can be used to train the model and obtain good training results.

[0088] Taking the first prompt information template and the second prompt information template as examples above, the model training process has been described. For the third prompt information template, AdaLORA local parameter fine-tuning can also be used for model training of downstream tasks.

[0089] In some embodiments, in addition to the categories and abstracts of the above documents, prompt information can also be presented to the user. If there are errors in the prompt information, the user can modify it according to actual needs. In addition, after modification, when the user clicks on a predetermined option, the modified prompt information can be used to iteratively optimize a predetermined model. During use, the user can spot-check some of the currently processed documents. If it is found that the processing results do not meet expectations, the model can be iteratively re-trained to optimize the performance of the model.

[0090] Figure 4 It is a schematic structural block diagram of a document classification device according to an embodiment of the present disclosure.

[0091] As Figure 4 shown, the document classification device 400 may include a division module 410, a first determination module 420, a second determination module 430, a third determination module 440, and a fourth determination module 450.

[0092] The division module 410 is configured to divide the body text in the document into multiple text segments according to the table of contents in the document.

[0093] The first determination module 420 is configured to determine a text segment abstract and a text segment category for each of the multiple text segments.

[0094] The second determination module 430 is configured to determine multiple first candidate categories and the confidence level of each first candidate category according to the text segment abstracts of the multiple text segments.

[0095] The third determination module 440 is configured to determine multiple second candidate categories and the confidence level of each second candidate category according to the text segment categories of the multiple text segments.

[0096] The fourth determination module 450 is configured to determine the category of the document according to the multiple first candidate categories, the confidence level of each first candidate category, the multiple second candidate categories, and the confidence level of each second candidate category.

[0097] In this embodiment, the first determination module includes: a first template determination sub-module, a first prompt information determination sub-module, and a first input sub-module. The first template determination sub-module is configured to determine, for each text segment, a first prompt information template corresponding to the text segment from multiple prompt information templates. The first prompt information determination sub-module is configured to determine first prompt information according to the text segment and the first prompt information template. The first input sub-module is configured to input the first prompt information into a predetermined model to obtain a text segment summary and a text segment category.

[0098] In this embodiment, the second determination module includes: a splicing sub-module, a second template determination sub-module, a second prompt information determination sub-module, and a second input sub-module. The splicing sub-module is configured to splice the text segment summaries of multiple text segments to obtain a first spliced text. The second template determination sub-module is configured to determine a second prompt information template corresponding to the first spliced text from multiple prompt information templates. The second prompt information determination sub-module is configured to determine second prompt information according to the first spliced text and the second prompt information template. The second input sub-module is configured to input the second prompt information into a predetermined model to obtain output information, where the output information includes multiple first candidate categories and the confidence of each first candidate category.

[0099] In this embodiment, the output information further includes a table of contents summary. The second prompt information further includes importance constraint information, where the importance constraint information characterizes: the relative importance relationship between multiple text segment summaries, and the relative importance relationship is related to the positions of multiple text segments in the table of contents text, and the predetermined model generates a table of contents summary based on the importance constraint information.

[0100] In this embodiment, the third determination module includes: a parameter determination sub-module, a confidence determination sub-module, and a second candidate category determination sub-module. The parameter determination sub-module is configured to determine the importance parameter of each text segment according to the position information of multiple text segments in the table of contents text. The confidence determination sub-module is configured to determine the confidence of each text segment category according to the importance parameter of each text segment and the text segment category of each text segment. The second candidate category determination sub-module is configured to determine multiple second candidate categories according to the sorting of the confidence of each text segment category.

[0101] In this embodiment, the document includes at least one table of contents, and the fourth determination module includes: a table of contents category determination sub-module and a document category determination sub-module. The table of contents category determination sub-module is configured to determine, for each table of contents in at least one table of contents, the table of contents category for the table of contents according to multiple first candidate categories, the confidence of each first candidate category, multiple second candidate categories, and the confidence of each second candidate category. The document category determination sub-module is configured to determine the category of the document according to the table of contents categories of at least one table of contents.

[0102] In this embodiment, the document includes a plurality of directories, and the document category determination sub-module includes: a first determination unit, a second determination unit, and a document category determination unit. The first determination unit is configured to determine a document abstract and a plurality of third candidate categories according to the directory abstracts of the plurality of directories. The second determination unit is configured to determine a plurality of fourth candidate categories according to the directory categories of the plurality of directories. The document category determination unit is configured to determine the category of the document according to the plurality of third candidate categories and the plurality of fourth candidate categories.

[0103] In this embodiment, the first determination unit includes: a splicing subunit, a third template determination subunit, a third prompt information determination subunit, and an information determination subunit. The splicing subunit is configured to splice the directory abstracts of at least one directory to obtain a second spliced text. The third template determination subunit is configured to determine a third prompt information template corresponding to the second spliced text from a plurality of prompt information templates. The third prompt information determination subunit is configured to determine third prompt information according to the second spliced text and the third prompt information template in the plurality of prompt information templates. The information determination subunit is configured to input the third prompt information into a predetermined model to obtain a document abstract and a plurality of third candidate categories.

[0104] In this embodiment, the third prompt information template includes at least one of first auxiliary information and second auxiliary information. The first auxiliary information represents that the directory abstract of the first directory is summary content, and the second auxiliary information represents that the directory abstract of the last directory is summary content.

[0105] In this embodiment, the prompt information template includes at least one of the following sub-information: the document name of the document. The directory name of the directory. A plurality of reference categories and category constraint information related to the scenario. The category constraint information represents that the predetermined model is used to determine category information according to the plurality of reference categories, and the category information includes at least one of a text segment category, a first candidate category, and a third candidate category. A quantity threshold representing the maximum number of characters of the abstract, and the abstract includes at least one of a text segment abstract, a directory abstract, and a document abstract. Processing order constraint information, and the processing order constraint information represents that the predetermined model is used to generate an abstract according to the category information.

[0106] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, including at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned document classification method.

[0107] According to an embodiment of the present disclosure, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above-mentioned document classification method.

[0108] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, including a computer program which, when executed by a processor, implements the above-mentioned document classification method.

[0109] Figure 5 is a structural block diagram of an electronic device for implementing the document classification method of the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0110] As Figure 5 shown, the device 500 includes a computing unit 501 which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0111] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0112] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the document classification method. For example, in some embodiments, the document classification method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the document classification method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the document classification method in any other suitable way (e.g., by means of firmware).

[0113] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0114] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0115] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0116] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0117] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0118] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0119] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0120] In the technical solutions of this disclosure, the processing of the user's personal information, such as collection, storage, use, processing, transmission, provision and disclosure, all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0121] In the technical solutions of this disclosure, before obtaining or collecting the user's personal information, the user's authorization or consent has been obtained.

[0122] The above specific implementation manners do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A document classification method, comprising: Dividing the main text in the document into multiple text segments according to the table of contents in the document; For each of the multiple text segments, determining a text segment summary and a text segment category; Determining multiple first candidate categories and the confidence level of each first candidate category according to the text segment summaries of the multiple text segments; Determining multiple second candidate categories and the confidence level of each second candidate category according to the text segment categories of the multiple text segments; And Determining the category of the document according to the multiple first candidate categories, the confidence level of each first candidate category, the multiple second candidate categories, and the confidence level of each second candidate category.

2. The method according to claim 1, wherein The step of, for each of the multiple text segments, determining a text segment summary and a text segment category includes: For each text segment, determining a first prompt information template corresponding to the text segment from multiple prompt information templates; Determining first prompt information according to the text segment and the first prompt information template; and Inputting the first prompt information into a predetermined model to obtain the text segment summary and the text segment category.

3. The method according to claim 1, wherein The step of determining multiple first candidate categories and the confidence level of each first candidate category according to the text segment summaries of the multiple text segments includes: Concatenating the text segment summaries of the multiple text segments to obtain a first concatenated text; Determining a second prompt information template corresponding to the first concatenated text from multiple prompt information templates; Determining second prompt information according to the first concatenated text and the second prompt information template; and Inputting the second prompt information into a predetermined model to obtain output information, where the output information includes the multiple first candidate categories and the confidence level of each first candidate category.

4. The method according to claim 3, wherein The output information further includes a table of contents summary; Wherein, the second prompt information further includes importance constraint information, and the importance constraint information characterizes: the relative importance relationship between the multiple text segment summaries, and the relative importance relationship is related to the positions of the multiple text segments in the table of contents text, and the predetermined model generates the table of contents summary based on the importance constraint information.

5. The method according to claim 1, wherein, The step of determining multiple second candidate categories and the confidence level of each second candidate category according to the text segment categories of the multiple text segments includes: Determining the importance parameter of each text segment according to the position information of the multiple text segments in the table of contents text; Determining the confidence level of each text segment category according to the importance parameter of each text segment and the text segment category of each text segment; and Determining the multiple second candidate categories according to the sorting of the confidence levels of each text segment category.

6. The method according to any one of claims 1 to 5, wherein The document includes at least one table of contents, and the step of determining the category of the document according to the multiple first candidate categories, the confidence level of each first candidate category, the multiple second candidate categories, and the confidence level of each second candidate category includes: For each of the at least one directory, determine a directory category for the directory according to the multiple first candidate categories, the confidence of each first candidate category, the multiple second candidate categories, and the confidence of each second candidate category; and Determine the category of the document according to the directory categories of the at least one directory respectively.

7. The method according to claim 6, wherein The document includes multiple directories, and the determining the category of the document according to the directory categories of the at least one directory respectively includes: Determine a document abstract and multiple third candidate categories according to the directory abstracts of the multiple directories respectively; Determine multiple fourth candidate categories according to the directory categories of the multiple directories respectively; and Determine the category of the document according to the multiple third candidate categories and the multiple fourth candidate categories.

8. The method according to claim 7, wherein The determining the document abstract and the multiple third candidate categories according to the directory abstracts of the multiple directories respectively includes: Concatenate the directory abstracts of the at least one directory respectively to obtain a second concatenated text; Determine a third prompt information template corresponding to the second concatenated text from multiple prompt information templates; Determine third prompt information according to the second concatenated text and the third prompt information template in the multiple prompt information templates; and Input the third prompt information into a predetermined model to obtain the document abstract and the multiple third candidate categories.

9. The method according to claim 8, wherein, The third prompt information template includes at least one of the following: First auxiliary information, characterizing that the directory abstract of the first directory is summary content; And Second auxiliary information, characterizing that the directory abstract of the last directory is summary content.

10. The method according to any one of claims 2, 3, 8 and 9, wherein, The prompt information template includes at least one of the following sub-information: The document name of the document; The directory name of the directory; Multiple reference categories and category constraint information related to the scenario, where the category constraint information characterizes that the predetermined model is used to determine category information according to the multiple reference categories, and the category information includes at least one of the text segment category, the first candidate category, and the third candidate category; A quantity threshold characterizing the maximum number of characters of the abstract, where the abstract includes at least one of the text segment abstract, the directory abstract, and the document abstract; And Processing order constraint information, where the processing order constraint information characterizes that the predetermined model is used to generate the abstract according to the category information.

11. A document classification device, comprising: A division module, configured to divide the body text in the document into multiple text segments according to the directories in the document; A first determination module, configured to determine a text segment abstract and a text segment category for each of the multiple text segments; A second determination module, configured to determine multiple first candidate categories and the confidence of each first candidate category according to the text segment abstracts of the multiple text segments respectively; A third determination module, configured to determine multiple second candidate categories and the confidence of each second candidate category according to the text segment categories of the multiple text segments respectively; And A fourth determination module, configured to determine the category of the document according to the multiple first candidate categories, the confidence of each first candidate category, the multiple second candidate categories, and the confidence of each second candidate category.

12. The apparatus according to claim 11, wherein, The first determination module includes: A first template determination sub-module, configured to determine, for each text segment, a first prompt information template corresponding to the text segment from multiple prompt information templates; A first prompt information determination sub-module, configured to determine first prompt information according to the text segment and the first prompt information template; and A first input sub-module, configured to input the first prompt information into a predetermined model to obtain the text segment summary and the text segment category.

13. The device according to claim 11, wherein The second determination module includes: A splicing sub-module, configured to splice the text segment summaries of the multiple text segments to obtain a first spliced text; A second template determination sub-module, configured to determine a second prompt information template corresponding to the first spliced text from multiple prompt information templates; A second prompt information determination sub-module, configured to determine second prompt information according to the first spliced text and the second prompt information template; and A second input sub-module, configured to input the second prompt information into a predetermined model to obtain output information, where the output information includes the multiple first candidate categories and the confidence level of each first candidate category.

14. The apparatus according to claim 13, wherein The output information further includes a table of contents summary; Wherein, the second prompt information further includes importance constraint information, and the importance constraint information characterizes: the relative importance relationship between the text segment summaries of the multiple text segments, and the relative importance relationship is related to the positions of the multiple text segments in the table of contents text, and the predetermined model generates the table of contents summary based on the importance constraint information.

15. The apparatus according to claim 11, wherein, The third determination module includes: A parameter determination sub-module, configured to determine the importance parameter of each text segment according to the position information of the multiple text segments in the table of contents text; A confidence level determination sub-module, configured to determine the confidence level of each text segment category according to the importance parameter of each text segment and the text segment category of each text segment; and A second candidate category determination sub-module, configured to determine the multiple second candidate categories according to the sorting of the confidence levels of each text segment category.

16. The device according to any one of claims 11 to 15, wherein The document includes at least one table of contents, and the fourth determination module includes: A table of contents category determination sub-module, configured to determine, for each table of contents in the at least one table of contents, the table of contents category for the table of contents according to the multiple first candidate categories, the confidence level of each first candidate category, the multiple second candidate categories, and the confidence level of each second candidate category; and A document category determination sub-module, configured to determine the category of the document according to the table of contents categories of the at least one table of contents.

17. The device according to claim 16, wherein, The document includes multiple tables of contents, and the document category determination sub-module includes: A first determination unit, configured to determine a document summary and multiple third candidate categories according to the table of contents summaries of the multiple tables of contents; A second determination unit, configured to determine multiple fourth candidate categories according to the table of contents categories of the multiple tables of contents; and A document category determination unit, configured to determine the category of the document according to the multiple third candidate categories and the multiple fourth candidate categories.

18. The apparatus according to claim 17, wherein, The first determination unit includes: A splicing subunit, configured to splice the directory summaries of the at least one directory respectively to obtain a second spliced text; A third template determination subunit, configured to determine a third prompt information template corresponding to the second spliced text from a plurality of prompt information templates; A third prompt information determination subunit, configured to determine third prompt information according to the second spliced text and the third prompt information template in the plurality of prompt information templates; and An information determination subunit, configured to input the third prompt information into a predetermined model to obtain the document summary and the plurality of third candidate categories.

19. The apparatus according to claim 18, wherein, The third prompt information template includes at least one of the following: First auxiliary information, indicating that the directory summary of the first directory is summary content; And Second auxiliary information, indicating that the directory summary of the last directory is summary content.

20. The apparatus according to any one of claims 12, 13, 18 and 19, wherein, The prompt information template includes at least one of the following sub-information: The document name of the document; The directory name of the directory; Multiple reference categories and category constraint information related to the scenario, where the category constraint information indicates that the predetermined model is used to determine category information according to the multiple reference categories, and the category information includes at least one of the text segment category, the first candidate category, and the third candidate category; A quantity threshold representing the maximum number of characters of the summary, where the summary includes at least one of the text segment summary, the directory summary, and the document summary; And Processing order constraint information, where the processing order constraint information indicates that the predetermined model is used to generate the summary according to the category information.

21. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 10.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 10.

23. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Text information recommendation method and device, electronic equipment and storage medium

    CN116150497A

  • Systems and Methods for Data Evaluation and Classification

    US20170017387A1