Document classification method and device and computer program product

By extracting feature words and calculating feature distinction degrees on the document data set, filtering out relevant feature words and training the classifier model, the problem of low efficiency and accuracy of document classification in the existing technology is solved, and more efficient document classification effect is achieved.

CN119939425APending Publication Date: 2025-05-06HUANENG CLEAN ENERGY RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510018030.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has large amount of calculation, high model complexity, unsatisfactory classification effect in document classification, and ignores the relative frequency and non-existence of terms, resulting in low classification accuracy and efficiency.

Method used

By extracting feature words from the document data set, the feature distinction of each feature word is calculated, and feature words whose feature distinction is greater than the preset value are selected to form a second feature word set, and the classifier model is trained for document classification using this feature word set.

Benefits of technology

Improves the efficiency and accuracy of document classification, and overcomes the limitations of only considering the absolute frequency of the term and ignoring relative frequency and non-existence in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939425A_ABST
    Figure CN119939425A_ABST
Patent Text Reader

Abstract

The invention discloses a document classification method and device and a computer program product, and relates to the technical field of text analysis, and the method comprises the steps that feature word extraction is carried out on a document data set to obtain a first feature word set, the document data set comprises documents of which categories are marked, and each document comprises a plurality of feature words; the feature distinction degree of each feature word in the first feature word set is calculated, a second feature word set is determined according to the feature distinction degree, the feature distinction degree represents the relevancy between the feature words and the document category, and the feature distinction degree of the feature words in the second feature word set is larger than a preset value; the to-be-classified documents are classified through a classifier model, a classification result is obtained, and the classifier model is obtained by training the initial model with the feature words in the second feature word set as input and the document categories corresponding to the feature words in the second feature word set as output. By adopting the technical scheme, the problem of how to improve the document classification efficiency is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text analysis, and in particular to a document classification method, device and computer program product. Background Art

[0002] In the field of scientific research, there are many types of document data, including research project applications, research reports, academic papers, etc. These document data contain a lot of information, but due to their unstructured characteristics, there are certain challenges in extracting useful information and classifying them. For data of scientific and technological document types, directly using classification algorithms often faces problems such as large amount of calculation, high model complexity, and unsatisfactory classification results. In addition, existing document classification methods usually only focus on the absolute frequency of terms for feature selection, while ignoring the relative frequency of terms in documents of different categories and the non-existence of terms in specific document categories, which leads to low accuracy and efficiency of document classification.

[0003] Therefore, in the related art, there is a problem of how to improve the efficiency of document classification.

[0004] Regarding the problem of how to improve the efficiency of document classification in related technologies, no effective solution has been proposed yet.

[0005] Therefore, it is necessary to improve the related technology to overcome the above-mentioned defects in the related technology. Summary of the invention

[0006] The embodiments of the present application provide a document classification method, device and computer program product to at least solve the problem of how to improve the efficiency of document classification in the related art.

[0007] According to one aspect of an embodiment of the present application, a document classification method is provided, comprising: extracting feature words from a document data set to obtain a first feature word set, wherein the document data set includes documents of marked categories, and each document includes multiple feature words; calculating the feature discrimination of each feature word in the first feature word set, and determining a second feature word set based on the feature discrimination, wherein the feature discrimination represents the correlation between the feature word and the document category, and the feature discrimination of the feature words in the second feature word set is greater than a preset value; classifying the documents to be classified through a classifier model to obtain a classification result, wherein the classifier model is obtained by training the initial model with the feature words in the second feature word set as input and the document category corresponding to the feature words in the second feature word set as output.

[0008] In an exemplary embodiment, the feature discrimination of each feature word in the first feature word set is calculated, including: counting the number of true positive examples, false positive examples, true negative examples and false negative examples of each feature word according to the document data set to obtain a confusion matrix for each feature word; calculating the true positive rate and false positive rate of each feature word according to the confusion matrix; and calculating the feature discrimination of each feature word according to the true positive rate and false positive rate.

[0009] In an exemplary embodiment, the number of true positive examples, false positive examples, true negative examples and false negative examples for each feature word is counted according to the document data set, including: for a target feature word, determining a target category corresponding to the target feature word; counting the number of documents in the target category that include the target feature word to obtain the number of true positive examples; counting the number of documents in documents not in the target category that include the target feature word to obtain the number of false positive examples; counting the number of documents in documents not in the target category that do not include the target feature word to obtain the number of true negative examples; counting the number of documents in the target category that do not include the target feature word to obtain the number of false negative examples.

[0010] In an exemplary embodiment, calculating the true positive rate and the false positive rate of each feature word according to the confusion matrix includes: determining the true positive rate according to the ratio of the number of true positive examples to a first sum value, wherein the first sum value represents the sum of the number of true positive examples and the number of false negative examples; determining the false positive rate according to the ratio of the number of false positive examples to a second sum value, wherein the second sum value represents the sum of the number of false positive examples and the number of true negative examples.

[0011] In an exemplary embodiment, calculating the feature discrimination of each feature word according to the true positive rate and the false positive rate includes: calculating the feature discrimination of each feature word by the following formula:

[0012]

[0013] Among them, TCM represents feature discrimination, TDR represents true positive rate, and FPR represents false positive rate.

[0014] In an exemplary embodiment, the classifier model is obtained by training in the following manner, including: taking the feature words in the second feature word set as input and the document categories corresponding to the feature words in the second feature word set as output, training the initial model; testing the trained initial model with a test feature word set, and when it is determined that the accuracy of the initial model is greater than a preset accuracy, determining the trained initial model as the classifier model, wherein the test feature word set includes test feature words and the document categories corresponding to the test feature words.

[0015] According to another aspect of an embodiment of the present application, a document classification device is also provided, including: an extraction module, used to extract feature words from a document data set to obtain a first feature word set, wherein the document data set includes documents of marked categories, and each document includes multiple feature words; a calculation module, used to calculate the feature discrimination of each feature word in the first feature word set, and determine a second feature word set based on the feature discrimination, wherein the feature discrimination represents the correlation between the feature word and the document category, and the feature discrimination of the feature words in the second feature word set is greater than a preset value; a classification module, used to classify the documents to be classified through a classifier model to obtain a classification result, wherein the classifier model is obtained by training the initial model with the feature words in the second feature word set as input and the document category corresponding to the feature words in the second feature word set as output.

[0016] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above document classification method when run.

[0017] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the document classification method through the computer program.

[0018] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of the method described in each embodiment of the present application are implemented.

[0019] Through this application, a first feature word set can be extracted from a document data set that has been marked with a category, the feature discrimination of each feature word in the first feature word set can be calculated, and then the feature words with a feature discrimination greater than a preset value are determined as a second feature word set, and the second feature word set is used to train the initial model to obtain a classifier model, and then the classifier model is used to classify the documents to be classified to obtain a document classification result. This solves the problem of how to improve the efficiency of document classification in the related art and achieves the effect of improving the efficiency of document classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0022] Figure 1 It is a hardware structure block diagram of a computer terminal of a document classification method according to an embodiment of the present application;

[0023] Figure 2 is a flow chart of a document classification method according to an embodiment of the present application;

[0024] Figure 3 is a schematic flow chart of a document classification method according to an embodiment of the present application;

[0025] Figure 4 It is a structural block diagram of a document classification device according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 1 is a hardware structure block diagram of a computer terminal for a document classification method according to an embodiment of the present application. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor (Central Processing Unit, MCU) or a programmable logic device (Field Programmable Gate Array, FPGA)) and a memory 104 for storing data, wherein the above-mentioned computer terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.

[0029] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the document classification method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0030] A wireless network provided by a communication provider of a computer terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.

[0031] In this embodiment, a document classification method is provided. Figure 2 is a flow chart of a document classification method according to an embodiment of the present application, such as Figure 2 As shown, the process includes the following steps:

[0032] Step S202, extracting feature words from a document dataset to obtain a first feature word set, wherein the document dataset includes documents that have been marked with categories, and each document includes a plurality of feature words;

[0033] Optionally, in the above step S202, for example, the document dataset includes documents in the artificial intelligence category and documents in the biotechnology category, and the feature words represent professional terms in the documents. For example, documents in the artificial intelligence category may include the professional term "neural network", and documents in the biotechnology category may include the professional term "gene editing".

[0034] Step S204, calculating the characteristic discrimination degree of each characteristic word in the first characteristic word set, and determining a second characteristic word set according to the characteristic discrimination degree, wherein the characteristic discrimination degree represents the relevance between the characteristic word and the document category, and the characteristic discrimination degree of the characteristic word in the second characteristic word set is greater than a preset value;

[0035] Optionally, in the above step S204, the feature discrimination represents a TCM (Trigonometric Comparison Measure, a feature selection method) value. The TCM method is a method specifically for the comparison and classification tasks of scientific and technological documents. It optimizes the feature selection process by considering the frequency characteristics of documents in different categories, especially the important information of the non-existence of terms in specific categories, thereby improving the accuracy and efficiency of text classification. Specifically, the TCM method evaluates the discrimination of terms by calculating the difference between the True Positive Rate (TPR) and the False Positive Rate (FPR) of the terms, and combines the trigonometric function transformation to select the most valuable features for classification.

[0036] Step S206, classifying the documents to be classified through a classifier model to obtain a classification result, wherein the classifier model is obtained by training the initial model with feature words in the second feature word set as input and document categories corresponding to the feature words in the second feature word set as output.

[0037] Optionally, in the above step S206, the initial model may be a support vector machine, a naive Bayes model, or the like.

[0038] Through the above steps, a first feature word set can be extracted from a document data set of a marked category, the feature discrimination of each feature word in the first feature word set is calculated, and then the feature words with a feature discrimination greater than a preset value are determined as a second feature word set, and the second feature word set is used to train the initial model to obtain a classifier model, and then the classifier model is used to classify the documents to be classified to obtain a document classification result. This solves the problem of how to improve the efficiency of document classification in the related art and achieves the effect of improving the efficiency of document classification.

[0039] In an exemplary embodiment, the feature discrimination of each feature word in the first feature word set is calculated, including: counting the number of true positive examples, false positive examples, true negative examples and false negative examples of each feature word according to the document data set to obtain a confusion matrix for each feature word; calculating the true positive rate and false positive rate of each feature word according to the confusion matrix; and calculating the feature discrimination of each feature word according to the true positive rate and false positive rate.

[0040] In an exemplary embodiment, the number of true positive examples, false positive examples, true negative examples and false negative examples for each feature word is counted according to the document data set, including: for a target feature word, determining a target category corresponding to the target feature word; counting the number of documents in the target category that include the target feature word to obtain the number of true positive examples; counting the number of documents in documents not in the target category that include the target feature word to obtain the number of false positive examples; counting the number of documents in documents not in the target category that do not include the target feature word to obtain the number of true negative examples; counting the number of documents in the target category that do not include the target feature word to obtain the number of false negative examples.

[0041] Optionally, in the above embodiment, the confusion matrix is ​​a table used to describe the relationship between the actual category of the document and the labeled category. For a binary classification problem, the confusion matrix contains four main parts:

[0042] True Positive (TP): The number of instances correctly labeled as the positive class.

[0043] False Positive (FP): The number of instances that are incorrectly labeled as positive.

[0044] True Negative (TN): The number of instances correctly labeled as negative class.

[0045] False Negative (FN): The number of instances that are incorrectly labeled as negative classes.

[0046] Following are the steps to construct a confusion matrix:

[0047] 1. Determine the categories: First, determine the positive and negative categories in the classification task. For example, in the scientific document classification task, the positive category can be defined as the "artificial intelligence" category and the negative category can be defined as the "biotechnology" category.

[0048] 2. Collect data: Collect the number of times each term appears in positive and negative documents.

[0049] 3. Count the number of occurrences:

[0050] True Positives (TP): count the number of times a term appears in the positive class documents.

[0051] False Positives FP: counts the number of times a term appears in negative documents.

[0052] True negative examples (TN): count the number of times a term does not appear in negative documents.

[0053] False negative examples FN: count the number of times a term does not appear in the positive class document.

[0054] Optionally, assume that there is a small-scale scientific document dataset containing the following documents:

[0055] Positive (artificial intelligence) documents: contain terms such as "neural network", "machine learning", and "deep learning".

[0056] Negative (biotechnology) documents: contain terms such as "gene editing", "cell culture", and "bioinformatics".

[0057] Then build the confusion matrix for the term "neural network":

[0058] TP statistics: Among all the "artificial intelligence" documents, 10 documents mentioned "neural networks".

[0059] Statistics FP: Of all the “biotechnology” documents, 2 documents mentioned “neural networks”.

[0060] Statistics TN: Among all "biotechnology" documents, 98 documents did not mention "neural networks".

[0061] Statistics FN: Among all the "artificial intelligence" documents, 10 documents did not mention "neural networks".

[0062] In an exemplary embodiment, calculating the true positive rate and the false positive rate of each feature word according to the confusion matrix includes: determining the true positive rate according to the ratio of the number of true positive examples to a first sum value, wherein the first sum value represents the sum of the number of true positive examples and the number of false negative examples; determining the false positive rate according to the ratio of the number of false positive examples to a second sum value, wherein the second sum value represents the sum of the number of false positive examples and the number of true negative examples.

[0063] Optionally, in the above embodiment, the true positive rate represents the frequency of the term appearing in the positive class, and the false positive rate represents the frequency of the term appearing in the negative class. These two indicators reflect the distribution of the terms in the two types of documents. The formulas for calculating the true positive rate (TPR) and the false positive rate (FPR) are as follows: TPR = TP / (TP+FN), FPR = FP / (FP+TN).

[0064] In an exemplary embodiment, calculating the feature discrimination of each feature word according to the true positive rate and the false positive rate includes: calculating the feature discrimination of each feature word by the following formula:

[0065]

[0066] Among them, TCM represents feature discrimination, TDR represents true positive rate, and FPR represents false positive rate.

[0067] Optionally, the above formula utilizes the deformation characteristics of the inverse tangent function in the trigonometric function to map the difference between TPR and FPR into a range of [0,1], and when the difference between TPR and FPR is larger, the TCM value is closer to 1, indicating that the discrimination of the term is higher.

[0068] Through the above embodiment, the terms can be sorted according to the calculated TCM values, and the terms with higher TCM values ​​can be selected as classification features to train the classifier model. The classification effect of the classifier model can be further optimized by adjusting the number of features.

[0069] In an exemplary embodiment, the classifier model is trained in the following manner, including: using the feature words in the second feature word set as input and the document categories corresponding to the feature words in the second feature word set as output to train the initial model; testing the trained initial model with a test feature word set, and when it is determined that the accuracy of the initial model is greater than a preset accuracy, determining the trained initial model as the classifier model, where the test feature word set includes test feature words and the document categories corresponding to the test feature words.

[0070] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. To better understand the above method for determining the power generation area, the following describes the above process in conjunction with embodiments, but does not limit the technical solutions of the embodiments of the present application. Specifically:

[0071] In an alternative embodiment, Figure 3 a further description of the document classification method of the present application is provided. As Figure 3 shown, it specifically includes the following steps:

[0072] Step S301: Data preprocessing;

[0073] Each scientific and technological document contains a large number of words. The preprocessing operations include removing stop words (such as "of", "is", etc.), removing low-frequency words (words that appear very rarely) and high-frequency words (common words such as "technology", "research", etc.), and performing word form reduction or stemming to reduce the noise of the data set and improve the comparability of features.

[0074] Step S302: Construct a confusion matrix;

[0075] For each term, count the number of true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN), and construct a confusion matrix of the term in positive-class documents and negative-class documents.

[0076] Optionally, assume there is a small-scale scientific and technological document data set containing the following documents:

[0077] Positive-class (artificial intelligence) documents: contain terms such as "neural network", "machine learning", "deep learning", etc.

[0078] Negative-class (biotechnology) documents: contain terms such as "gene editing", "cell culture", "bioinformatics", etc.

[0079] Then construct a confusion matrix for the term "neural network":

[0080] TP statistics: Among all the "artificial intelligence" documents, 10 documents mentioned "neural networks".

[0081] Statistics FP: Of all the “biotechnology” documents, 2 documents mentioned “neural networks”.

[0082] Statistics TN: Among all "biotechnology" documents, 98 documents did not mention "neural networks".

[0083] Statistics FN: Among all the "artificial intelligence" documents, 10 documents did not mention "neural networks".

[0084] Step S303: Calculate the true positive rate TPR and the false positive rate FPR;

[0085] Among them, the true positive rate indicates the frequency of terms appearing in the positive class, and the false positive rate indicates the frequency of terms appearing in the negative class. These two indicators reflect the distribution of terms in the two types of documents. The calculation formulas for the true positive rate and false positive rate are as follows: TPR = TP / (TP+FN), FPR = FP / (FP+TN).

[0086] Step S304: Calculate the TCM value;

[0087] The TCM value is calculated using the following formula:

[0088]

[0089] This formula utilizes the deformation characteristics of trigonometric functions, especially the inverse tangent function, to map the difference between TPR and FPR into a range of [0,1]. When the difference between TPR and FPR is larger, the TCM value is closer to 1, indicating that the discrimination of the term is higher.

[0090] Step S305: feature selection;

[0091] According to the calculated TCM value, all terms are sorted and the term with the highest TCM value is selected as the classification feature to train the classifier. The classification training effect can be further optimized by adjusting the number of feature selections.

[0092] Step S306: classifier training and testing;

[0093] Use the selected classification features to train a classifier (such as support vector machine SVM, naive Bayes NB, etc.), and evaluate the performance of the classifier on the test set, such as using indicators such as accuracy and recall to verify the effectiveness of the classifier.

[0094] Through the above embodiments, a trigonometric function comparison mechanism is introduced into text classification, and the frequency of occurrence of terms relative to documents is taken into account, so that more discriminating features can be selected to train the classifier model, overcoming the limitation of traditional document classification methods that only consider the absolute frequency of term occurrence and ignore the relative frequency, thereby improving the processing performance and efficiency of text classification tasks.

[0095] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0096] In this embodiment, a document classification device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0097] Figure 4 : is a structural block diagram of a document classification device according to an embodiment of the present application, the device comprising:

[0098] An extraction module 42 is used to extract feature words from a document data set to obtain a first feature word set, wherein the document data set includes documents that have been marked with categories, and each document includes a plurality of feature words;

[0099] A calculation module 44 is used to calculate the characteristic discrimination degree of each characteristic word in the first characteristic word set, and determine the second characteristic word set according to the characteristic discrimination degree, wherein the characteristic discrimination degree represents the relevance between the characteristic word and the document category, and the characteristic discrimination degree of the characteristic word in the second characteristic word set is greater than a preset value;

[0100] The classification module 46 is used to classify the documents to be classified through a classifier model to obtain a classification result, wherein the classifier model is obtained by training the initial model with the feature words in the second feature word set as input and the document category corresponding to the feature words in the second feature word set as output.

[0101] Through the above device, a first feature word set can be extracted from a document data set of a marked category, the feature discrimination of each feature word in the first feature word set can be calculated, and then the feature words with a feature discrimination greater than a preset value can be determined as a second feature word set, and the second feature word set can be used to train the initial model to obtain a classifier model, and then the classifier model can be used to classify the documents to be classified to obtain a document classification result. This solves the problem of how to improve the efficiency of document classification in the related art and achieves the effect of improving the efficiency of document classification.

[0102] In an exemplary embodiment, the calculation module 44 is also used to count the number of true positive examples, false positive examples, true negative examples and false negative examples of each feature word according to the document data set to obtain a confusion matrix for each feature word; calculate the true positive rate and false positive rate of each feature word according to the confusion matrix; and calculate the feature discrimination of each feature word according to the true positive rate and false positive rate.

[0103] In an exemplary embodiment, the calculation module 44 is also used to determine the target category corresponding to the target feature word; count the number of documents in the target category that include the target feature word to obtain the number of true positive examples; count the number of documents in documents other than the target category that include the target feature word to obtain the number of false positive examples; count the number of documents in documents other than the target category that do not include the target feature word to obtain the number of true negative examples; count the number of documents in the target category that do not include the target feature word to obtain the number of false negative examples.

[0104] In an exemplary embodiment, the calculation module 44 is further used to determine the true positive rate according to the ratio of the number of true positive examples to a first sum value, wherein the first sum value represents the sum of the number of true positive examples and the number of false negative examples; and determine the false positive rate according to the ratio of the number of false positive examples to a second sum value, wherein the second sum value represents the sum of the number of false positive examples and the number of true negative examples.

[0105] In an exemplary embodiment, the calculation module 44 is further configured to calculate the feature discrimination of each feature word by using the following formula:

[0106]

[0107] Among them, TCM represents feature discrimination, TDR represents true positive rate, and FPR represents false positive rate.

[0108] In an exemplary embodiment, the classification module 46 is also used to train the initial model using the feature words in the second feature word set as input and the document categories corresponding to the feature words in the second feature word set as output; the trained initial model is tested by using a test feature word set, and when it is determined that the accuracy of the initial model is greater than a preset accuracy, the trained initial model is determined as the classifier model, wherein the test feature word set includes test feature words and the document categories corresponding to the test feature words.

[0109] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.

[0110] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:

[0111] S1, extracting feature words from a document data set to obtain a first feature word set, wherein the document data set includes documents of marked categories, and each document includes a plurality of feature words;

[0112] S2, calculating the feature discrimination of each feature word in the first feature word set, and determining a second feature word set according to the feature discrimination, wherein the feature discrimination represents the relevance between the feature word and the document category, and the feature discrimination of the feature words in the second feature word set is greater than a preset value;

[0113] S3, classifying the documents to be classified through the classifier model to obtain the classification results, wherein the classifier model is obtained by training the initial model with the feature words in the second feature word set as input and the document categories corresponding to the feature words in the second feature word set as output.

[0114] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0115] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail herein.

[0116] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0117] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0118] S1, extracting feature words from a document data set to obtain a first feature word set, wherein the document data set includes documents of marked categories, and each document includes a plurality of feature words;

[0119] S2, calculating the feature discrimination of each feature word in the first feature word set, and determining a second feature word set according to the feature discrimination, wherein the feature discrimination represents the relevance between the feature word and the document category, and the feature discrimination of the feature words in the second feature word set is greater than a preset value;

[0120] S3, classifying the documents to be classified through the classifier model to obtain the classification results, wherein the classifier model is obtained by training the initial model with the feature words in the second feature word set as input and the document categories corresponding to the feature words in the second feature word set as output.

[0121] An embodiment of the present application further provides a computer program product, comprising a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program product, and when the computer program is executed by a processor, the steps of the method described in each embodiment of the present application are implemented.

[0122] Optionally, in this embodiment, the above computer program may be configured to implement the following steps when executed by a processor:

[0123] S1, extracting feature words from a document data set to obtain a first feature word set, wherein the document data set includes documents of marked categories, and each document includes a plurality of feature words;

[0124] S2, calculating the feature discrimination of each feature word in the first feature word set, and determining a second feature word set according to the feature discrimination, wherein the feature discrimination represents the relevance between the feature word and the document category, and the feature discrimination of the feature words in the second feature word set is greater than a preset value;

[0125] S3, classifying the documents to be classified through the classifier model to obtain the classification results, wherein the classifier model is obtained by training the initial model with the feature words in the second feature word set as input and the document categories corresponding to the feature words in the second feature word set as output.

[0126] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail herein.

[0127] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0128] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A document classification method, characterized in that: include: Extracting feature words from a document data set to obtain a first feature word set, wherein the document data set includes documents of marked categories, and each document includes a plurality of feature words; Calculating the characteristic discrimination degree of each characteristic word in the first characteristic word set, and determining the second characteristic word set according to the characteristic discrimination degree, wherein the characteristic discrimination degree indicates the relevance between the characteristic word and the document category, and the characteristic discrimination degree of the characteristic word in the second characteristic word set is greater than a preset value; The documents to be classified are classified by a classifier model to obtain a classification result, wherein the classifier model is obtained by training the initial model with feature words in the second feature word set as input and document categories corresponding to the feature words in the second feature word set as output.

2. The method according to claim 1, characterized in that Calculating the characteristic discrimination degree of each characteristic word in the first characteristic word set includes: Counting the number of true positive examples, false positive examples, true negative examples and false negative examples of each feature word according to the document data set to obtain a confusion matrix for each feature word; Calculate the true positive rate and false positive rate of each feature word according to the confusion matrix; The feature discrimination degree of each feature word is calculated according to the true positive rate and the false positive rate.

3. The method according to claim 2, characterized in that Counting the number of true positive examples, false positive examples, true negative examples, and false negative examples of each feature word according to the document data set includes: For a target feature word, determining a target category corresponding to the target feature word; Counting the number of documents in the target category that include the target feature word to obtain the number of true positive examples; Counting the number of documents including the target feature word in documents not belonging to the target category to obtain the number of false positive examples; Counting the number of documents that do not include the target feature word in documents that are not in the target category, to obtain the number of true negative examples; The number of documents in the target category that do not include the target feature word is counted to obtain the number of the false negative examples.

4. The method according to claim 3, characterized in that Calculating the true positive rate and the false positive rate of each feature word according to the confusion matrix includes: Determine the true positive rate according to a ratio of the number of true positive examples to a first sum value, wherein the first sum value represents a sum of the number of true positive examples and the number of false negative examples; The false positive rate is determined according to a ratio of the number of false positive examples to a second sum value, wherein the second sum value represents a sum of the number of false positive examples and the number of true negative examples.

5. The method according to claim 2, characterized in that: Calculating the feature discrimination of each feature word according to the true positive rate and the false positive rate includes: The characteristic discrimination degree of each characteristic word is calculated by the following formula: Among them, TCM represents feature discrimination, TDR represents true positive rate, and FPR represents false positive rate.

6. The method according to claim 1, characterized in that The classifier model is obtained by training in the following manner, including: Taking the feature words in the second feature word set as input and the document categories corresponding to the feature words in the second feature word set as output, training the initial model; The trained initial model is tested by using a test feature word set, and when it is determined that the accuracy of the initial model is greater than a preset accuracy, the trained initial model is determined as the classifier model, wherein the test feature word set includes test feature words and document categories corresponding to the test feature words.

7. A document classification device, characterized in that: include: An extraction module is used to extract feature words from a document data set to obtain a first feature word set, wherein the document data set includes documents of a marked category, and each document includes a plurality of feature words; a calculation module is used to calculate a feature discrimination degree of each feature word in the first feature word set, and determine a second feature word set according to the feature discrimination degree, wherein the feature discrimination degree indicates a correlation degree between a feature word and a document category, and the feature discrimination degree of a feature word in the second feature word set is greater than a preset value; A classification module is used to classify documents to be classified through a classifier model to obtain a classification result, wherein the classifier model is obtained by training the initial model with feature words in the second feature word set as input and document categories corresponding to the feature words in the second feature word set as output.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 6 when executed.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 6 through the computer program.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.