Subject Classification Method, System, Electronic Device and Storage Medium for Multi-Label Text
By combining the text classification model and weighted loss function based on the RoBerta-base algorithm, the problem of data imbalance in multi-label text topic classification is solved, and the recall rate and classification accuracy of the model are improved, especially in the multi-label text topic classification task.
Patent Information
- Application Number
- CN202411755212.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The prior art has problems with multi-label classification and long-tail distribution in multi-label text theme classification, resulting in poor model effects. Especially in multi-label classification tasks, the coupling between labels is high and data is unbalanced, which affects the classification effect.
The text classification model based on the RoBerta-base algorithm is adopted, combining the classification layer and weighted loss function, and the text to be predicted is processed through document segmentation, special symbol removal and unitization, and a fully connected layer and multiple binary classifiers are added to the model. The improved weighted loss function is used for training, which improves the model's attention to a small number of samples.
It effectively alleviates the problem of poor model effectiveness caused by data imbalance, improves the recall rate and classification accuracy of the model, especially in the multi-label text topic classification task.
Smart Images

Figure CN119226961B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of topic classification, and specifically relates to a method, a system, an electronic device and a storage medium for classifying topics of multi-label texts. Background Art
[0002] Topic classification is a task of classifying text data according to topic categories. There are currently many methods for implementing topic classification, including statistical learning methods such as Naive Bayes and decision trees, as well as some deep learning methods. In the process of topic classification, there are two relatively difficult problems to be solved: 1. Multi-label classification. The same document may involve multiple topics, and there is a relatively high coupling between some topics. 2. Long-tail distribution. This problem is mainly reflected in the data, that is, the data of different topics is extremely unbalanced.
[0003] Deep long-tail learning is a very challenging task in visual recognition and often appears in natural language processing tasks. The long-tail distribution means that a small number of categories in the data set account for a large proportion of the total data set, that is, the so-called "head" labels, while the number of samples of most categories is small and concentrated in the "tail". There are currently many methods for solving the long-tail distribution problem, including sampling methods, mutual information methods, etc. There are many sampling methods, including oversampling by copying existing samples or synthesizing new data and undersampling by reducing the number of most samples. Sampling to process data imbalance problems is the most direct, but there are also many problems, including feature loss, overfitting, increased noise impact, and inapplicability to multi-label processing.
[0004] Sampling can solve the problem of sample imbalance to a certain extent, but for multi-label classification tasks, there may be a relatively high coupling between labels. Blind sampling not only cannot achieve the optimization effect, but may also cause greater impacts. For example, for news about some technology companies releasing new products, this news involves the theme of technological innovation and contains information about the released new products. In addition, the news also mentions the content of the stock price increase due to the release of new products, and involves both the technology and finance themes at the same time. When there are multiple documents that simultaneously have the themes of technological innovation and finance stocks, directly sampling the finance stock theme will also affect the data of the technological innovation theme. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a method, a system, an electronic device and a storage medium for classifying topics of multi-label texts, which are used to solve all or at least part of the technical problems existing in the above-mentioned prior art.
[0006] To achieve the above purpose, the embodiments of the present invention provide a method for classifying topics of multi-label texts, including:
[0007] Obtain the multi-label text to be predicted;
[0008] Input the multi-label text to be predicted into a pre-constructed text classification model, and output the topic classification result of the multi-label text to be predicted, where the text classification model is constructed based on the RoBerta-base algorithm and combined with a classification layer.
[0009] Optionally, the topic classification method for the multi-label text further includes:
[0010] After obtaining the multi-label text to be predicted, segment the multi-label text to be predicted according to a preset length, remove special symbols from the segmented segments and unitize them to obtain the target multi-label text data to be predicted.
[0011] Optionally, the construction process of the text classification model includes:
[0012] Select the structure of the Roberta-base model as the basic architecture of the text classification model, and load the Roberta-base model through transformers;
[0013] Add a fully connected layer after encoding by the Roberta-base model to adjust the vector dimension;
[0014] Add multiple independent binary classifiers after the fully connected layer for classification;
[0015] Define a weighted loss function as the loss function of the text classification model.
[0016] Optionally, the loss function of the text classification model is as follows:
[0017] ;
[0018] In the formula, is used to correct the cross-entropy loss generated in different situations, is the true label, is the prediction result.
[0019] Optionally, ;
[0020] In the formula, is the threshold of different classifiers, is the true label, is the prediction result, and K represents a real number.
[0021] Optionally, inputting the multi-label text to be predicted into a pre-constructed text classification model and outputting the topic classification result of the multi-label text to be predicted includes:
[0022] Segment the multi-label text to be predicted into documents, remove special symbols, and after unitization, convert the multi-label text to be predicted into an input format suitable for the text classification model;
[0023] Input the multi-label text to be predicted in the input format suitable for the text classification model into the Roberta-base model, output the embedding vectors, and input the embedding vectors into the fully connected layer to map the dimension of the embedding vectors to a low dimension to make it consistent with the classification category dimension, and then classify through multiple binary classifiers, where each classifier outputs a classification category;
[0024] Map the results output by multiple binary classifiers to determine the theme of the multi-label text to be predicted.
[0025] On the other hand, the present invention also provides a theme classification system for multi-label text, including:
[0026] An acquisition unit for acquiring the multi-label text to be predicted;
[0027] A classification unit for inputting the multi-label text to be predicted into a pre-constructed text classification model and outputting the theme classification result of the multi-label text to be predicted, where the text classification model is constructed based on the RoBerta-base algorithm and combined with a classification layer.
[0028] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the above-mentioned theme classification method for multi-label text are implemented.
[0029] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned theme classification method for multi-label text are implemented.
[0030] Through the above technical solutions, by using the improved text classification model, the model can pay more attention to a small number of samples, reduce the impact of poor model performance caused by the long-tailed distribution of data, and improve the model recall rate. In the theme classification task, it can effectively alleviate the situation of poor model performance caused by data imbalance.
[0031] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. Description of the Drawings
[0032] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification, and are used to explain the embodiments of the present invention together with the following specific implementation manners, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0033] Figure 1 It is a flowchart of the implementation of a method for theme classification of multi-label texts provided by an embodiment of the present invention;
[0034] Figure 2 It is a schematic diagram of the construction process of a text classification model provided by an embodiment of the present invention;
[0035] Figure 3 It is a flowchart of using a text classification model for theme classification provided by an embodiment of the present invention;
[0036] Figure 4 It is a detailed flowchart of the implementation of a method for theme classification of multi-label texts provided by an embodiment of the present invention;
[0037] Figure 5 It is a schematic diagram of the structure of a system for theme classification of multi-label texts provided by an embodiment of the present invention. Specific implementation manners
[0038] The following will describe in detail the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described here are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.
[0039] Refer to Figure 1 As shown, it is a flowchart of the implementation of a method for theme classification of multi-label texts provided by an embodiment of the present invention, including the following execution steps:
[0040] Step 100: Obtain the multi-label text to be predicted.
[0041] In some implementation manners, after obtaining the multi-label text to be predicted, the multi-label text to be predicted is segmented into documents according to a preset length, and special symbols are removed from the segmented segments and unitized to obtain the target multi-label text data to be predicted.
[0042] In a specific embodiment, 1) Document segmentation: For overly long documents, perform length segmentation by intercepting the content with a length of N at the beginning and end of the document. The value of N is generally 512 and can be adjusted according to different situations. 2) Removal of special symbols: Remove long special symbols through regular expressions. That is, if there is a long text segment (with a length exceeding M, generally 6 and can be adjusted according to different situations) consisting only of special symbols or alphanumeric characters (including: ~!@#$%^&*()_+,. / >< / ’;《》【】[]{}a-zA-Z), replace it with an empty string, i.e., directly delete it. 3) Tokenization: Represent the original text as smaller units (tokens), which can be mapped to numbers and vectors and applied to downstream NLP tasks, namely text topic classification, which can be specifically implemented through AutoTokenizer in transformers.
[0043] Step 101: Input the multi-label text to be predicted into a pre-constructed text classification model to output the topic classification result of the multi-label text to be predicted, where the text classification model is constructed based on the RoBerta-base algorithm combined with a classification layer.
[0044] In some embodiments, the RoBerta pre-trained model is selected. As a robustly optimized version of the Bert method, RoBerta can achieve good results in the Chinese language domain. This is because the model has the following improvements based on Bert: 1) Dynamic masking: During training, each sentence is randomly masked for multiple words, while Bert only masks 15%. This enables the model to learn more complex language patterns. 2) Training batch size: Larger batch sizes are used for training, improving the training efficiency and performance of the model. 3) Training scale: Larger model parameters are used, and the training time is also increased, allowing the model to learn richer language knowledge. 4) Excluding the Next Sentence Prediction (NSP) task: Roberta removes the NSP task in BERT, reducing complexity. Research shows that this task does not significantly contribute to the model's performance. The specific usage of Roberta is as follows: 1) First, download the Roberta-base model. 2) Load the pre-trained Roberta-base model through transformers. 3) Construct the classification model structure. The classification model consists of Roberta-base and a classification layer. Given text features, i.e., token embeddings, are obtained through Roberta-base. Since downstream classification tasks are required, the token embedding of the CLS is taken. After encoding by Roberta-base, a fully connected layer is added to adjust the vector dimension. Finally, N independent binary classifiers are added after the fully connected layer for classification. 4) Fine-tune the text topic classification data, and the loss function used during the fine-tuning process is implemented using the loss function defined in this application.
[0045] In some embodiments, referring to Figure 2 as shown, the construction process of the text classification model includes the following steps:
[0046] S1: Select the structure of the Roberta-base model as the basic architecture of the text classification model, and load the Roberta-base model through transformers;
[0047] S2: Add a fully connected layer after encoding by the Roberta-base model to adjust the vector dimension;
[0048] S3: Add multiple independent binary classifiers after the fully connected layer for classification;
[0049] S4: Define a weighted loss function as the loss function of the text classification model.
[0050] By using the improved loss function, during the model training process, the model can pay more attention to a small number of samples, reduce the impact of the poor model performance caused by the long-tailed distribution of data, and improve the model recall rate. In addition, since the function is differentiable, during the model training process, for the samples with correct classification, the model can still maintain the correct classification of the samples. In the topic classification task, it can effectively alleviate the poor model performance caused by data imbalance.
[0051] Specifically, the loss function of the text classification model is as follows:
[0052] ;
[0053] In the formula, is used to correct the cross-entropy loss generated under different circumstances, is the true label, is the prediction result.
[0054] Among them, ;
[0055] In the formula, is the threshold of different classifiers, is the true label, is the prediction result, represents a real number.
[0056] When , , at this time the loss approaches 0. When , , keep the cross-entropy loss. When , , the cross-entropy loss is of the original loss. During the model training process, when the number of negative samples is much larger than that of positive samples, the model will be more inclined to negative class samples. However, with the improved loss function, it can make the negative samples generate a smaller loss and the positive samples generate a larger loss at this time. For example, when the number of negative samples is much larger than that of positive samples, the model will be more inclined to judge the samples as the negative samples with a larger number. Suppose there are 990 negative samples among 1000 samples. If all samples are predicted as negative samples, the accuracy rate is as high as 99%. However, by using the improved loss function, in this case, the of the negative class samples is very small, while the of the positive class samples remains unchanged, which will prompt the model to start concentrating on positive samples.
[0057] In some embodiments, configure the following optimization strategy:
[0058] 1) Set evaluation metrics: Precision, Recall, and Accuracy are selected as the evaluation metrics for the final results, and the training process is evaluated by the validation set loss and accuracy;
[0059] 2) Define hyperparameters: including Epochs (the size is defined according to the scale of the training set), Batch size
[0060] (defined according to the dataset and hardware device, usually select 16, 32, 64), learning rate (select 4e-5), Hidden size (defined according to the pre-trained model, select the Boberta-base model, so the hidden layer is 1024);
[0061] 3) Loss function: Adopt a weighted loss function;
[0062] 4) Optimizer: Generally, the Adam optimizer is selected.
[0063] In some embodiments, for the training and evaluation of the text classification model: Set the number of batches according to the number of training samples as the number of training set samples / batch_size, evaluate the validation set data, and the evaluation method is to calculate the loss of the current stage model on the validation set data, and predict the accuracy and recall rate of the validation set. If the current loss is less than the loss of the previous stage model, save the model. The same goes for the subsequent steps. If the current stage loss is less than the loss of the previous stage model, save the model, otherwise, do not save. If after 1000 batch iterations, the validation set loss does not decrease, end the training.
[0064] In some embodiments, as shown in Figure 3 When performing step 101, the following steps can be specifically executed:
[0065] S1010: Segment the multi-label text to be predicted into documents, remove special symbols, and after unitization, convert the multi-label text to be predicted into the input format that conforms to the text classification model.
[0066] S1011: Input the multi-label text to be predicted in the input format that conforms to the text classification model into the Roberta-base model, output the embedding vector, and input the embedding vector into the fully connected layer to map the embedding vector dimension to a low dimension to make it consistent with the classification category dimension, and then perform classification through multiple binary classifiers, where each classifier outputs a classification category.
[0067] S1012: Map the results output by multiple binary classifiers to determine the theme of the multi-label text to be predicted.
[0068] In a specific embodiment, the text prediction process is as follows: 1) Input the text to be predicted. 2) The text is preprocessed by a preprocessing module to remove special symbols and determine whether the text length is too long. 3) After Tokenization, the text is converted into an input format acceptable to the model. 4) Use the trained model for prediction: Input the output of the previous step into the BoBerta model. After obtaining the model output, that is, the token embedding, take the CLS token embedding, input it into the fully connected layer, map the vector dimension to a low dimension (consistent with the classification category), and then classify it through multiple classifiers. Each classifier corresponds to a classification category, and each classifier will output 0 or 1, that is, whether it belongs to this category. If the classification result is 1, it belongs to this category. 5) Map according to the final result of the classifier and return the theme of this document.
[0069] In some embodiments, refer to Figure 4 As shown, it is a detailed implementation flowchart of a multi-label text theme classification method provided by an embodiment of the present invention, including the following execution steps:
[0070] S400: Obtain the multi-label text to be predicted.
[0071] S401: Segment the multi-label text to be predicted according to a preset length, remove special symbols from the segmented segments and unitize them to obtain the target multi-label text data to be predicted.
[0072] S402: Select the structure of the Roberta-base model as the basic architecture of the text classification model, and load the Roberta-base model through transformers.
[0073] S403: Add a fully connected layer after encoding by the Roberta-base model to adjust the vector dimension.
[0074] S404: Add multiple independent binary classifiers after the fully connected layer for classification.
[0075] S405: Define a weighted loss function as the loss function of the text classification model.
[0076] S406: Input the multi-label text to be predicted that conforms to the input format of the text classification model into the Roberta-base model, output the embedding vector, and input the embedding vector into the fully connected layer to map the embedding vector dimension to a low dimension to make it consistent with the classification category dimension, and then classify it through multiple binary classifiers. Among them, each classifier corresponds to output a classification category.
[0077] S407: Map the results output by multiple binary classifiers to determine the theme of the multi-label text to be predicted.
[0078] By adopting a weighted loss function, it is possible to solve to a certain extent the problem of poor model performance caused by the data long-tail phenomenon in the multi-label text classification task, improve the recall rate of the model, and the improved loss function can effectively improve the classification effect of theme classification.
[0079] Refer to Figure 5 As shown in the figure, it is a schematic structural diagram of a theme classification system for multi-label text provided by an embodiment of the present invention, including:
[0080] An acquisition unit 50, configured to acquire the multi-label text to be predicted;
[0081] A classification unit 51, configured to input the multi-label text to be predicted into a pre-constructed text classification model, and output the theme classification result of the multi-label text to be predicted, where the text classification model is constructed based on the RoBerta-base algorithm and combined with a classification layer.
[0082] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the theme classification method for multi-label text described in any one of the above embodiments are implemented.
[0083] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the theme classification method for multi-label text described in any one of the above embodiments are implemented.
[0084] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0085] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 a means for implementing the functions specified in one block or multiple blocks.
[0086] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 a means for implementing the functions specified in one block or multiple blocks.
[0087] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 a means for implementing the functions specified in one block or multiple blocks.
[0088] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0089] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0090] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0091] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0092] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for topic classification of multi-label text, characterized in that, Including: Obtain the multi-label text to be predicted; Input the multi-label text to be predicted into a pre-constructed text classification model, and output the theme classification result of the multi-label text to be predicted, where the text classification model is constructed based on the RoBerta-base algorithm and combined with a classification layer; Among them, the construction process of the text classification model includes: Select the structure of the Roberta-base model as the basic architecture of the text classification model, and load the Roberta-base model through transformers; Add a fully connected layer after encoding by the Roberta-base model to adjust the vector dimension; Add multiple independent binary classifiers after the fully connected layer for classification; Define a weighted loss function as the loss function of the text classification model; Among them, the loss function of the text classification model is as follows: ; In the formula, used to correct the cross-entropy loss generated in different situations, is the true label, is the predicted result; ; Wherein, is the threshold of different classifiers, is the true label, is the prediction result, and K represents a real number.
2. The method for topic classification of multi-label text according to claim 1, characterized in that, The theme classification method of the multi-label text further includes: After obtaining the multi-label text to be predicted, segment the multi-label text to be predicted according to a preset length, remove special symbols from the segmented segments, and unitize them to obtain the target multi-label text data to be predicted.
3. The method for subject classification of multi-label text according to claim 1, characterized in that, Input the multi-label text to be predicted into a pre-constructed text classification model, and output the theme classification result of the multi-label text to be predicted, including: Segment the multi-label text to be predicted, remove special symbols, and after unitization, convert the multi-label text to be predicted into an input format that conforms to the text classification model; Input the multi-label text to be predicted in the input format that conforms to the text classification model into the Roberta-base model, output the embedding vector, and input the embedding vector into the fully connected layer to map the embedding vector dimension to a low dimension to make it consistent with the classification category dimension, and then classify it through multiple binary classifiers, where each classifier outputs a classification category; Map the results output by multiple binary classifiers to determine the theme of the multi-label text to be predicted.
4. A multi-label text theme classification system for the multi-label text theme classification method according to any one of claims 1-3, characterized in that, Including: An acquisition unit for obtaining the multi-label text to be predicted; A classification unit for inputting the multi-label text to be predicted into a pre-constructed text classification model, and outputting the theme classification result of the multi-label text to be predicted, where the text classification model is constructed based on the RoBerta-base algorithm and combined with a classification layer.
5. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the theme classification method of the multi-label text according to any one of claims 1-3.
6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the theme classification method of the multi-label text according to any one of claims 1-3.
Citation Information
Patent Citations
Text classification method, system and equipment based on multi-label association and medium
CN118227790A