A text classification method

By constructing a category lexicon and optimizing classification thresholds, and updating keyword weights with user feedback, the problems of low efficiency and high cost in the existing technology are solved, and efficient and accurate text classification is achieved.

CN114090774BActive Publication Date: 2025-07-11CHENGDU AEROSPACE SCI & TECH BIG DATA RES INST CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111372486.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2025-07-11
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

In the prior art, text classification is inefficient, costly and poorly effective, and it is difficult to effectively classify especially in the absence of sample data.

Method used

Build a category vocabulary, and use crawling category keywords and preprocessing training texts to obtain classification thresholds, and update category keyword weights and tolerances based on user feedback, optimize classification thresholds to improve text classification accuracy.

Benefits of technology

It improves the efficiency and accuracy of text classification, reduces the cost of manual labeling, solves the cold start problem of text classification, and improves the recognition rate of text classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114090774B_ABST
    Figure CN114090774B_ABST
Patent Text Reader

Abstract

The present invention discloses a text classification method, including: constructing a number of categories, and based on each of the categories, crawling corresponding category keywords from the network, and constructing a category thesaurus with the category keywords; collecting a number of training texts, and preprocessing the training texts to obtain the preprocessed training texts; according to the preprocessed training texts, obtaining a classification threshold corresponding to each category; obtaining a text to be classified, and according to the classification threshold, obtaining the category corresponding to the text to be classified. The text classification method provided by the present invention improves the efficiency and recognition rate of text classification, and reduces the cost of text classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of text classification, and particularly relates to a text classification method. Background Art

[0002] Text classification refers to the process of automatically classifying texts by a computer according to a certain classification system or rules. It is not only a natural language processing problem but also a pattern recognition problem. Chinese text classification generally includes processes such as text preprocessing, word segmentation, model construction, and classification.

[0003] Text classification is generally divided into: classification methods based on knowledge engineering (KE), classification methods based on machine learning (ML), and classification methods based on deep learning (DL).

[0004] 1. Classification methods based on knowledge engineering

[0005] Classification methods based on knowledge engineering refer to classifying texts by relying on expert experience, manually extracting rules, and designing a set of computer programs to simulate the reasoning and decision-making of human experts. Classification methods based on knowledge engineering rely on the empirical knowledge accumulated in long-term practice and do not require complex mathematical modeling, so they have been widely used. However, it is difficult to obtain the required empirical knowledge for such methods, and the classification accuracy depends on the richness of expert experience and the perfection of the extracted rules. At the same time, when there are many rules and categories, problems such as classification conflicts and combinatorial explosion in the classification process make the reasoning process slow and inefficient.

[0006] 2. Classification methods based on machine learning

[0007] Classification methods based on machine learning refer to classifying texts by preprocessing the text (word segmentation and stop word removal), representing the text (numerically transforming the text through methods such as One-hot (one-hot encoding), Bag of Words (bag-of-words model), N-gram (language model), TF-IDF (term frequency-inverse document frequency index), etc.) and feature extraction based on the already classified text category samples, and then classifying through the computer's autonomous learning and extraction of rules. It mainly includes classification methods based on naive Bayes, classification methods based on support vector machines, classification methods based on decision trees, classification methods based on fuzzy decision-making, and classification methods based on neural networks.

[0008] The classification method based on Naive Bayes has the characteristics of low complexity, stable classification effect, being suitable for the training of small-scale data, requiring fewer parameters to be estimated, and being insensitive to missing data. However, the premise of this classification method is to assume that attributes are independent of each other, and this premise is often difficult to hold in practical applications. When there are too many attributes or the correlation between attributes is large, the classification effect is not good. At the same time, the classification effect depends on the prior probability and is sensitive to the expression form of the input data.

[0009] The classification method based on Support Vector Machine has high generalization ability, can be used for the calculation of high-dimensional data, and avoids the problems of neural network structure selection and local minimum points. However, this method is sensitive to missing data and cannot solve non-linear problems.

[0010] The classification method based on Decision Tree is easy to understand, the generation of logical expressions is relatively simple, the requirement for data preprocessing is low, and it can handle irrelevant features. It can evaluate the model through static testing, can process large-scale data in a short time, can handle both numerical and regular attributes simultaneously, and can construct multi-attribute decision trees. However, this method tends to be more inclined to features with more values, has difficulties in dealing with missing data, is prone to overfitting, and tends to ignore the correlation of dataset attributes.

[0011] The classification method based on Fuzzy Decision has strong robustness, can solve problems such as non-linearity and strong coupling, and has strong fault tolerance. However, this method is not suitable for the processing of simple information and cannot precisely define the target.

[0012] The classification method based on Neural Network has the characteristics of strong parallel processing ability, strong learning ability, and high classification accuracy. It has strong robustness and fault tolerance to data noise, can solve complex non-linear relationships, and has the function of memory. However, there are a large number of parameters to be determined in the neural network training process, the learning process between networks cannot be observed, the output results are difficult to explain, the learning time is long, and the classification effect cannot be guaranteed.

[0013] The classification method based on Machine Learning depends on a certain number of labeled sample sets and needs to extract corresponding text features through feature engineering. For different texts, the feature engineering of this classification method cannot be unified, and generally, multi-faceted feature engineering will be carried out to extract complete features. Therefore, it consumes a large amount of manpower and has a high usage cost.

[0014] 3. Classification Method Based on Deep Learning

[0015] The classification method based on deep learning refers to using a deep learning model to solve the problem of large-scale text classification based on the pre-divided text category samples. However, this classification method often requires a large amount of labeled sample data, and the annotation of sample data requires a lot of manpower. At the same time, there are a large number of parameters to be determined during the training process of the deep learning model. The learning process of the model cannot be observed, the learning time is long, and the classification effect cannot be guaranteed. Summary of the Invention

[0016] In view of the above deficiencies in the prior art, a text classification method provided by the present invention solves the problems of low efficiency, high cost, and poor effect in text classification in the prior art.

[0017] To achieve the above invention purpose, the technical solution adopted by the present invention is: A text classification method, including:

[0018] Construct several categories, and based on each category, crawl the corresponding category keywords from the network, and construct a category thesaurus with the category keywords;

[0019] Collect several training texts, and preprocess the training texts to obtain preprocessed training texts;

[0020] According to the preprocessed training texts, obtain the classification threshold corresponding to each category;

[0021] Obtain the text to be classified, and according to the classification threshold, obtain the category corresponding to the text to be classified.

[0022] Further, the collecting several training texts, and preprocessing the training texts to obtain preprocessed training texts includes:

[0023] Collect several training texts, the training texts include a title and a body text, and the words in the title and the body text all correspond to entity attributes;

[0024] Screen the words in the training texts that are the same as the category keywords to obtain screened words;

[0025] Add the category corresponding to the category keyword to the screened words to obtain the training texts after primary processing;

[0026] Perform word segmentation, stop word removal, and named entity recognition operations on the training texts after primary processing to obtain preprocessed training texts;

[0027] The preprocessing result of the training text includes the preprocessing result corresponding to the title and the preprocessing result corresponding to the body text. Among them, the preprocessing result corresponding to the title is [d i p i c i , where d iDenote the i-th entity in the title, where i = 1, 2, …, I, and I represents the number of entities in the title, p i Denote the position of entity d i in the title, c i Denote the set of categories corresponding to entity d i ;

[0028] The preprocessing result corresponding to the body text is [d j p j c j , where d j Denote the j-th entity in the body text, where j = 1, 2, …, J, and J represents the number of entities in the body text, p j Denote the position of entity d j in the body text, c j Denote the set of categories corresponding to entity d i .

[0029] Furthermore, obtaining the classification threshold corresponding to each category according to the preprocessed training text includes:

[0030] Randomly divide the preprocessed training text into n groups;

[0031] Obtain the scores of the preprocessed training text in each group under each category;

[0032] Obtain the classification threshold corresponding to each category according to the scores.

[0033] Furthermore, the score is:

[0034]

[0035] where y a Denote the score of the text under category a, where category a is a category in the category thesaurus, k t Denote the influence coefficient of the text title, S ta Denote the set of entities belonging to category a in the preprocessing result corresponding to the title of the training text, d i Denote the set S ta The i-th entity in, count(d i ) Denote the entity d i The number of times it appears in S ta , w i Denote the entity d i The weight in category a, k c Denote the influence coefficient of the text body, S ca Denote the set of entities belonging to category a in the preprocessing result corresponding to the body text, d j Denote the set S ca The j-th entity in, count(dj ) represents entity d j The number of occurrences in set S ca , w j represents entity d j The weight of entity d in category a, span(d j ) represents entity d j The span of entity d in set S ca , S represents the entity sequence obtained after preprocessing the main text of the training text, last dj represents entity d j The last position where entity d appears in entity sequence S, first dj represents entity d j The first position where entity d appears in the entity sequence, len(S) represents the total number of entities in the entity sequence.

[0036] Furthermore, obtaining the classification threshold corresponding to each category according to the score includes:

[0037] For a certain category, arrange the scores of the preprocessed training texts in each group in descending order within the group, take out the score of the last preprocessed training text belonging to the current category in each group, and use the taken-out score as the classification critical value of each group in the current category;

[0038] According to the classification critical values of each group in the current category, obtain the average value of the classification critical values, and use the current average value as the classification threshold of the current category;

[0039] According to the above steps, traverse all categories to obtain the classification threshold corresponding to each category;

[0040] The classification threshold is:

[0041]

[0042] where v represents the classification threshold, y i' represents the classification critical value of the i'-th group, i' = 1, 2,..., n.

[0043] Furthermore, obtaining the text to be classified and obtaining the category corresponding to the text to be classified according to the classification threshold includes:

[0044] Perform word segmentation, stop word removal, and named entity recognition operations on the text to be classified to obtain the processed text to be classified;

[0045] Obtain the scores of the processed text to be classified under each category, and use the categories with scores higher than the classification threshold as the categories corresponding to the text to be classified.

[0046] Furthermore, after obtaining the category corresponding to the text to be classified, it further includes:

[0047] Send the category corresponding to the text to be classified to the user side, send a category confirmation message to the user side, and receive the feedback message from the user side. The feedback message includes that the classification of the text to be classified is correct and incorrect.

[0048] Update the weights and tolerance times of the category keywords according to the feedback message. The initial weights of the category keywords are 1 and the initial tolerance times are 0.

[0049] Select the text to be classified with correct classification according to the feedback message.

[0050] Remove the words identical to the category keywords and repeated words in the text to be classified with correct classification, and obtain the importance degree of the entities corresponding to the remaining words.

[0051] Add the words with importance degree greater than the set threshold to the category thesaurus, and remove the category keywords with updated weights less than the first threshold T and updated tolerance times greater than the second threshold R to complete the update of the category thesaurus.

[0052] Further, the update of the weights and tolerance times of the category keywords includes:

[0053] If the feedback message is that the classification is correct, update the weights and tolerance times of the category keywords to:

[0054]

[0055]

[0056] If the feedback message is that the classification is incorrect, update the weights and tolerance times of the category keywords to:

[0057]

[0058]

[0059] Among them, w' represents the updated weight, r' represents the updated tolerance time, σ represents the moving step size, w represents the weight, and r represents the tolerance time.

[0060] Further, the importance degree is:

[0061]

[0062] Among them, count(d) represents the number of times the entity d appears in the text to be classified, the entity d represents the entity corresponding to the remaining words, C represents the set of texts to be classified, count(C) represents the total number of entities in the set C, count(D d) represents the total number of documents containing entity d in the text to be classified, last id represents the last occurrence position of entity d in the text to be classified C i first id represents the first occurrence position of entity d in the text to be classified C i count(C i ) represents the text to be classified C i the total number of entities in it.

[0063] Furthermore, after receiving the feedback message from the receiving client, it further includes:

[0064] According to the feedback message, update the classification threshold to:

[0065]

[0066] where v represents the classification threshold, v' represents the updated classification threshold, τ represents the classification threshold update step size, m represents the number of texts to be classified with incorrect classification, n represents the number of texts to be classified with correct classification, y i ' represents the score of the text to be classified with incorrect classification under the category, y j' represents the score of the text to be classified with correct classification under the category, i” = 1, 2, …, m, j' = 1, 2, …, n.

[0067] The beneficial effects of the present invention are:

[0068] (1) The present invention provides a text classification method, which improves the efficiency and recognition rate of text classification and reduces the cost of text classification.

[0069] (2) The present invention obtains category keywords based on the category, constructs a category thesaurus with the category keywords, and obtains the category score of the text based on the constructed category thesaurus, and classifies the text in combination with the set category, solves the cold start problem of text classification, and avoids the problem that cannot be carried out due to lack of samples in machine learning.

[0070] (3) The present invention preprocesses the training text to quickly obtain the categories of a certain number of training texts, thereby reducing the manual annotation cost.

[0071] (4) The present invention combines the classification score and category information of the training text to construct a classification threshold, and can quickly classify the text to be classified through the classification threshold, improving the efficiency and accuracy of text classification.

[0072] (5) After the text classification of the present invention is initially started, it determines whether the classification result of the text to be classified is correct based on user feedback, sets a penalty and reward mechanism, updates the weights of the category keywords, and eliminates the category keywords whose weights reach the set threshold, thereby improving the text classification accuracy.

[0073] (6) The present invention extracts new category keywords based on the correctly classified text to be classified, and expands the category thesaurus through the extracted category keywords, thereby improving the text classification accuracy.

[0074] (7) The present invention updates the classification threshold based on the classification result of the text to be classified, thereby improving the text classification accuracy. Description of the Drawings

[0075] Figure 1 It is a flowchart of a text classification method proposed by an embodiment of the present application. Detailed Embodiment

[0076] The following describes the detailed embodiment of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiment. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0077] The following describes the embodiment of the present invention in detail with reference to the drawings.

[0078] Embodiment 1

[0079] As Figure 1 shown, a text classification method includes:

[0080] Construct several categories, and based on each category, crawl the corresponding category keywords from the network, and construct a category thesaurus with the category keywords.

[0081] For example, when the category is the automotive category, the corresponding category keywords crawled can be SAIC-GM and Dongfeng Nissan, etc.

[0082] Collect several training texts, and preprocess the training texts to obtain the preprocessed training texts.

[0083] According to the preprocessed training texts, obtain the classification threshold corresponding to each category.

[0084] Obtain the text to be classified, and according to the classification threshold, obtain the category corresponding to the text to be classified.

[0085] In this embodiment, the initial weight of the category keyword is set to 1, and the initial tolerance times are set to 0.

[0086] In a possible implementation, collecting a number of training texts, and preprocessing the training texts to obtain preprocessed training texts, including:

[0087] Collecting a number of training texts, where the training texts include a title and a body text, and the words in the title and the body text all correspond to entity attributes.

[0088] Screening the words in the training texts that are the same as the category keywords to obtain screened words.

[0089] Adding the category corresponding to the category keywords to the screened words to obtain the training texts after primary processing.

[0090] Performing word segmentation, stop word removal, and named entity recognition operations on the training texts after primary processing to obtain preprocessed training texts.

[0091] The preprocessing result of the training texts includes the preprocessing result corresponding to the title and the preprocessing result corresponding to the body text. Among them, the preprocessing result corresponding to the title is [d i p i c i , where d i represents the i-th entity in the title, i = 1, 2,..., I, and I represents the number of entities in the title, p i represents the position of the entity d i in the title, and c i represents the category set corresponding to the entity d i .

[0092] The preprocessing result corresponding to the body text is [d j p j c j , where d j represents the j-th entity in the body text, j = 1, 2,..., J, and J represents the number of entities in the body text, p j represents the position of the entity d j in the body text, and c j represents the category set corresponding to the entity d i .

[0093] In this embodiment, if the entity d i and the entity d i do not have corresponding categories, then c i and c j are empty.

[0094] In a possible implementation, obtaining the classification threshold corresponding to each category according to the preprocessed training texts, including:

[0095] Randomly divide the preprocessed training text into n groups.

[0096] Obtain the scores of the preprocessed training text in each group under each category.

[0097] According to the scores, obtain the classification threshold corresponding to each category.

[0098] In a possible implementation manner, the score is:

[0099]

[0100] where y a represents the score of the text under category a, and category a is a category in the category library, k t represents the influence coefficient of the text title, S ta represents the set of entities belonging to category a in the preprocessing result corresponding to the title of the training text, d i represents the i-th entity in the set S ta , count(d i ) represents the number of times the entity d i appears in S ta , w i represents the weight of the entity d i in category a, k c represents the influence coefficient of the text body, S ca represents the set of entities belonging to category a in the preprocessing result corresponding to the text body, d j represents the j-th entity in the set S ca , count(d j ) represents the number of times the entity d j appears in the set S ca , w j represents the weight of the entity d j in category a, span(d j ) represents the span of the entity d j in the set S ca , S represents the entity sequence obtained after preprocessing the text body of the training text, last dj represents the last position where the entity d j appears in the entity sequence S, first dj represents the first position where the entity d j appears in the entity sequence, and len(S) represents the total number of entities in the entity sequence.

[0101] In a possible implementation manner, the obtaining the classification threshold corresponding to each category according to the scores includes:

[0102] For a certain category, sort the scores of the preprocessed training texts in each group from largest to smallest within the group, take out the score of the last preprocessed training text belonging to the current category in each group, and use the taken-out score as the classification threshold of each group for the current category. The last training text belonging to the current category refers to the last training text among the consecutive training texts belonging to the current category found starting from the first training text within the current group.

[0103] According to the classification threshold of each group for the current category, obtain the average value of the classification thresholds, and use the current average value as the classification threshold for the current category.

[0104] According to the above steps, traverse all categories to obtain the classification threshold corresponding to each category.

[0105] The classification threshold is:

[0106]

[0107] where v represents the classification threshold, and y i' represents the classification threshold of the i'-th group, i' = 1, 2,..., n.

[0108] In a possible implementation manner, the obtaining of the text to be classified and the obtaining of the category corresponding to the text to be classified according to the classification threshold include:

[0109] Perform word segmentation, stop word removal, and named entity recognition operations on the text to be classified to obtain the processed text to be classified.

[0110] Obtain the scores of the processed text to be classified under each category, and use the categories with scores higher than the classification threshold as the categories corresponding to the text to be classified.

[0111] In a possible implementation manner, after obtaining the category corresponding to the text to be classified, it further includes:

[0112] Send the category corresponding to the text to be classified to the user side, send a category confirmation message to the user side, and receive the feedback message from the user side. The feedback message includes that the classification of the text to be classified is correct and incorrect.

[0113] Optionally, if no feedback message is received within the specified time, send the category corresponding to the text to be classified to another user side, send a category confirmation message to the other user side, and receive the feedback message from the other user side.

[0114] Update the weights and tolerance times of the category keywords according to the feedback message. The initial weights of the category keywords are 1 and the initial tolerance times are 0.

[0115] Select the text to be classified with the correct classification according to the feedback message.

[0116] Remove the words that are the same as the category keywords and duplicate words in the text to be classified with the correct classification, and obtain the importance of the entities corresponding to the remaining words.

[0117] The words that are the same as the category keywords indicate that there are words corresponding to the corresponding category in the training text. After removing the words that are the same as the category keywords and duplicate words, the remaining words have an empty category and are non-repetitive.

[0118] Add the words with an importance greater than the set threshold to the category thesaurus, and remove the category keywords whose updated weight is less than the first threshold T and whose updated tolerance count is greater than the second threshold R to complete the update of the category thesaurus.

[0119] In this embodiment, the first threshold T is set to 0.5, and the second threshold R is set to 100.

[0120] Optionally, a shielding thesaurus can be constructed, add the removed category keywords to the shielding thesaurus, and the words existing in the shielding thesaurus are not allowed to be added to the category thesaurus again.

[0121] In a possible implementation manner, the updating of the weight and tolerance count of the category keywords includes:

[0122] If the feedback message is that the classification is correct, then update the weight and tolerance count of the category keywords to:

[0123]

[0124]

[0125] If the feedback message is that the classification is incorrect, then update the weight and tolerance count of the category keywords to:

[0126]

[0127]

[0128] Among them, w' represents the updated weight, r' represents the updated tolerance count, σ represents the moving step size, w represents the weight, and r represents the tolerance count.

[0129] In this embodiment, the moving step size σ is set to 0.001.

[0130] In a possible implementation manner, the importance is:

[0131]

[0132] Among them, count(d) represents the number of times the entity d appears in the text to be classified, the entity d represents the entity corresponding to the remaining vocabulary, C represents the set of texts to be classified, count(C) represents the total number of entities in the set C, and count(D d ) represents the total number of documents in the text to be classified that contain the entity d, and last id represents the last position where the entity d appears in the text to be classified C i , and first id represents the first position where the entity d appears in the text to be classified C i , and count(C i ) represents the total number of entities in the text to be classified C i .

[0133] In a possible implementation manner, after receiving the feedback message from the user terminal, it further includes:

[0134] According to the feedback message, update the classification threshold to:

[0135]

[0136] Among them, v represents the classification threshold, v' represents the updated classification threshold, τ represents the classification threshold update step size, m represents the number of texts to be classified with incorrect classification, n represents the number of texts to be classified with correct classification, y i' represents the score of the text to be classified with incorrect classification under the category, and y j' represents the score of the text to be classified with correct classification under the category, where i” = 1, 2, …, m and j' = 1, 2, …, n.

[0137] Embodiment 2

[0138] In this embodiment, a text classification device is provided, including a construction module, a processing module, an acquisition module, and a classification module;

[0139] The construction module is used to construct several categories, and based on each category, crawl the corresponding category keywords from the network, and construct a category thesaurus with the category keywords;

[0140] The processing module is used to collect several training texts and preprocess the training texts to obtain the preprocessed training texts;

[0141] The acquisition module is used to obtain the classification threshold corresponding to each category according to the preprocessed training texts;

[0142] The classification module is used to obtain the text to be classified and obtain the category corresponding to the text to be classified according to the classification threshold.

[0143] Embodiment 3

[0144] In this embodiment, a text classification device is provided, including a processor and a memory;

[0145] The memory stores computer-executable instructions;

[0146] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the text classification method described in Embodiment 1.

[0147] Embodiment 4

[0148] In this embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions, which are used to implement the text classification method described in Embodiment 1 when the computer-executable instructions are executed by a processor.

[0149] Embodiment 5

[0150] In this embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it implements the text classification method described in Embodiment 1.

[0151] The present invention provides a text classification method, which improves the efficiency and recognition rate of text classification and reduces the cost of text classification. The present invention obtains category keywords based on categories, constructs a category thesaurus with the category keywords, and obtains the category score of the text based on the constructed category thesaurus, and classifies the text in combination with the set categories, solving the cold start problem of text classification and avoiding the problem that cannot be carried out due to lack of samples in machine learning.

[0152] The present invention preprocesses the training text to quickly obtain the categories of a certain number of training texts, thereby reducing the manual annotation cost. The present invention combines the classification score and category information of the training text to construct a classification threshold, and can quickly classify the text to be classified through the classification threshold, improving the efficiency and accuracy of text classification.

[0153] After the text classification is initially started, the present invention judges whether the classification result of the text to be classified is correct based on user feedback, sets a penalty and reward mechanism, updates the weights of the category keywords, and eliminates the category keywords whose weights reach the set threshold, improving the text classification accuracy.

[0154] The present invention extracts new category keywords based on the correctly classified text to be classified, expands the category thesaurus with the extracted category keywords, and improves the text classification accuracy. The present invention updates the classification threshold based on the classification result of the text to be classified, improving the text classification accuracy.

Claims

1. A text classification method, characterized in that, Including: Construct several categories, and based on each category, crawl the corresponding category keywords from the network, and construct a category thesaurus with the category keywords; Collect several training texts, and preprocess the training texts to obtain the preprocessed training texts; According to the preprocessed training texts, obtain the classification threshold corresponding to each category; Obtain the text to be classified, and according to the classification threshold, obtain the category corresponding to the text to be classified; The obtaining the classification threshold corresponding to each category according to the preprocessed training texts includes: Randomly divide the preprocessed training text into n groups; Obtain the scores of the preprocessed training texts in each group under each category; According to the scores, obtain the classification threshold corresponding to each category; The score is: Among them, represents the score of the text under the category a where the category a is a category in the category thesaurus, represents the influence coefficient of the text title, represents the set of entities belonging to the category a in the preprocessing result corresponding to the title of the training text, represents the in the set i th entity, represents the entity appears times, represents the entity in the category a weight, represents the influence coefficient of the text body, represents the set of entities belonging to the category a in the preprocessing result corresponding to the text body, represents the in the set j th entity, represents the entity appears times, represents the entity in the category a weight, represents the entity in the set span, S represents the entity sequence obtained after preprocessing the text body of the training text, represents the entity in the entity sequence S the last position where it appears, represents the entity the first position where it appears in the entity sequence, represents the total number of entities in the entity sequence.

2. The text classification method according to claim 1, characterized in that The collecting several training texts and preprocessing the training texts to obtain the preprocessed training texts includes: Collect several training texts, the training texts include a title and a body, and the words in the title and the body all correspond to entity attributes; Screen the words in the training texts that are the same as the category keywords to obtain the screened words; Add the category corresponding to the category keyword to the screened words to obtain the initially processed training texts; Perform word segmentation, stop word removal, and named entity recognition operations on the initially processed training texts to obtain the preprocessed training texts; The preprocessing results of the training text include the preprocessing results corresponding to the title and the preprocessing results corresponding to the body text. Among them, the preprocessing results corresponding to the title are , where represents the i th entity in the title, i = 1, 2, …, I , I represents the number of entities in the title, represents the entity 's position in the title, represents the entity 's corresponding category set; The preprocessing result corresponding to the above text is , where represents the j th entity in the above text, j = 1, 2, …, J , J represents the number of entities in the above text, represents the position of entity in the above text, represents the set of categories corresponding to entity .

3. The text classification method according to claim 1, characterized in that, The obtaining the classification threshold corresponding to each category according to the scores includes: For a certain category, arrange the scores of the preprocessed training texts in each group in descending order within the group, take out the score of the last preprocessed training text belonging to the current category in each group, and use the taken-out score as the classification critical value of each group in the current category; According to the classification critical values of each group in the current category, obtain the average value of the classification critical values, and use the current average value as the classification threshold of the current category; According to the above steps, traverse all categories to obtain the classification threshold corresponding to each category; The classification threshold is: Among them, represents the classification threshold, represents the classification critical value of the th group, = 1, 2, …, n .

4. The text classification method according to claim 1, characterized in that The obtaining the text to be classified and according to the classification threshold, obtaining the category corresponding to the text to be classified includes: Perform word segmentation, stop word removal, and named entity recognition operations on the text to be classified to obtain the processed text to be classified; Obtain the scores of the processed text to be classified under each category, and use the categories with scores higher than the classification threshold as the categories corresponding to the text to be classified.

5. The text classification method according to claim 3, characterized in that After obtaining the category corresponding to the text to be classified, it further includes: Send the category corresponding to the text to be classified to the user terminal, send a category confirmation message to the user terminal, and receive the feedback message from the user terminal, the feedback message includes that the classification of the text to be classified is correct and incorrect; According to the feedback message, update the weights and tolerance times of the category keywords, the initial weight of the category keywords is 1 and the initial tolerance time is 0; Select the text to be classified with correct classification according to the feedback message; Remove the words and repeated words in the text to be classified with correct classification that are the same as the category keywords, and obtain the importance degree of the entities corresponding to the remaining words; Add the words with importance degree greater than the set threshold to the category thesaurus, and remove the category keywords with updated weights less than the first threshold T and updated tolerance times greater than the second threshold R to complete the update of the category thesaurus.

6. The text classification method according to claim 5, characterized in that, Updating the weight and tolerance times of the category keywords includes: If the feedback message indicates correct classification, update the weight and tolerance times of the category keywords to: If the feedback message indicates incorrect classification, update the weight and tolerance times of the category keywords to: Among them, represents the updated weight, represents the updated tolerance count, represents the moving step size, represents the weight, represents the tolerance count.

7. The text classification method according to claim 6, wherein The importance level is: Among them, represents the number of occurrences of the entity d in the text to be classified. The entity d represents the corresponding entity of the remaining words, C represents the set of texts to be classified, represents the set C the total number of entities in it, represents the number of documents containing the entity d in the text to be classified, represents the entity d the last occurrence position of the entity in the text to be classified is, represents the entity d the first occurrence position of the entity in the text to be classified is, represents the text to be classified the total number of entities in it.

8. The text classification method according to claim 5, characterized in that, After receiving the feedback message from the user terminal, it further includes: According to the feedback message, update the classification threshold to: Among them, represents the classification threshold, represents the updated classification threshold, represents the classification threshold update step size, m represents the number of texts to be classified with incorrect classification, n represents the number of texts to be classified with correct classification, represents the score of the text to be classified with incorrect classification under the category, represents the score of the text to be classified with correct classification under the category, = 1, 2, …, m , = 1, 2, …, n .

Citation Information

Patent Citations

  • Internet website classification method and device

    CN106156372A

  • LDA and word2vec algorithm-based news text classification method

    CN107609121A

  • Warning condition classification method and system

    CN110990562A