A text classification method, system, terminal and storage medium
By extracting sentence or word vectors from the Bert model and performing weighted summation, combined with the Attention mechanism and multi-layer loss function optimization, the problem of insufficient utilization of the Bert model's coarse classification results is solved, and the accuracy of text classification and model convergence are improved.
Patent Information
- Application Number
- CN202310714387.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-06-15
AI Technical Summary
The coarse classification results of the Bert model in the existing technology are difficult to be fully utilized, resulting in insufficient text classification performance.
By extracting the sentence vector or word vector of the Bert model, combining it with the category vector for weighted summation, building a neural network and optimizing the loss function, the model is optimized using the Attention mechanism and multi-layer loss function.
It improves the accuracy of text classification and the convergence of the model, fully utilizes the text representation capabilities of the Bert model, and is suitable for optimization in vertical fields.
Smart Images

Figure CN116662551B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text classification and relates to a text classification method, system, terminal and storage medium. Background Art
[0002] Text classification is a crucial component in text processing, with a wide range of applications, such as spam filtering, news classification, and part-of-speech tagging. It's essentially the same as other classification methods. The core approach involves first extracting features from the classified data, converting the raw data into a high-dimensional vector composed of eigenvalues, and then selecting an appropriate similarity metric for classification. However, text also has its own unique characteristics. Based on these characteristics, the general process for text classification is: 1. Preprocessing; 2. Text representation and feature selection; 3. Classifier construction; 4. Classification. Generally speaking, text classification involves assigning text to one or more categories within a given classification system.
[0003] With the rapid development of deep learning technology in recent years, a large number of text classification methods based on character embeddings, word embeddings, and sentence embeddings have emerged. Bert (BERT) is a prominent representative of these methods. Developed by Google researchers in 2018, it has proven to be a state-of-the-art technology for various natural language processing tasks, such as text classification, text summarization, and text generation.
[0004] BERT's superior performance stems from its massive network parameters and vast amounts of training data. However, this also creates difficulties for companies or organizations with limited hardware or data to train and optimize BERT. Using the rough classification results obtained by the BERT model and performing secondary optimization on the feature vectors has become a new research direction in text classification. Summary of the Invention
[0005] The purpose of the present invention is to solve the problem in the prior art that the rough classification results obtained by the Bert model cannot be well utilized, and to provide a text classification method, system, terminal and storage medium.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a text classification method, comprising the following steps:
[0008] Input the text information to be classified into the Bert model to obtain several classification results;
[0009] Extract the sentence vector CLS of the Bert model output layer to obtain several sentence vectors; or aggregate all the word vectors of the Bert model output layer to obtain several sentence vectors;
[0010] According to the classification results, the sentence vectors of the same category are averaged to obtain the category vector. All sentence vectors are associated with all category vectors, and the category vectors are weighted summed to obtain a new sentence vector. The new sentence vector is used as the input of the additional classification, followed by a softmax layer to obtain the category distribution, and the loss of the sentence vector is calculated in combination with the true label of the category distribution.
[0011] The calculated sentence vector loss is superimposed to construct a neural network;
[0012] The weighted sum of the losses of several superimposed sentence vectors is used to obtain the total loss, and the total loss is optimized by the optimizer corresponding to the constructed neural network until convergence.
[0013] Furthermore, the present invention inputs the text information to be classified into the Bert model to obtain several classification results, including:
[0014] Split the sentence to be classified into several sub-vectors, input the word vectors into the Bert model, and obtain several classification results, which are recorded as C1, C2, ..., C n , where n represents the number of categories.
[0015] Furthermore, the aggregation of the present invention adopts the summing, averaging or pooling method.
[0016] Furthermore, the present invention calculates the loss of the sentence vector, including:
[0017] According to the classification results, the sentence vector A of the same category m Find the mean and get the category vector, recorded as B1, B2, ..., B n , integration to obtain matrix B;
[0018] Attention is associated with each sentence vector and all category vectors, and the category vectors are weighted summed to obtain a new sentence vector A. ′ m :
[0019]
[0020]
[0021] Among them, α n represents the weight of B, Indicates α1~α n , i is a positive integer, s() represents a certain attention scoring mechanism, and A m Replace with A ′ m ;
[0022] Am As the input of an additional classifier, followed by a softmax layer, we get the category distribution L m , and its true label Calculate the loss of sentence vector.
[0023] Furthermore, the neural network optimizer of the present invention adopts a stochastic gradient descent optimizer SGD or an adaptive moment estimation optimizer Adam.
[0024] In a first aspect, the present invention provides a text classification system, comprising:
[0025] The classification result calculation module is used to input the text information to be classified into the Bert model to obtain several classification results;
[0026] The sentence vector calculation module is used to extract the sentence vector CLS of the Bert model output layer to obtain several sentence vectors; or to aggregate all word vectors of the Bert model output layer to obtain several sentence vectors;
[0027] The loss calculation module is used to calculate the average of the sentence vectors of the same category based on the classification results to obtain the category vector, associate all sentence vectors with all category vectors, and perform weighted summation on the category vectors to obtain a new sentence vector. The new sentence vector is used as the input for additional classification, followed by a softmax layer to obtain the category distribution, and the loss of the sentence vector is calculated in combination with the true label of the category distribution.
[0028] The neural network construction module is used to superimpose the loss of the calculated sentence vectors to construct a neural network;
[0029] The optimization module is used to perform weighted summation of the losses of several superimposed sentence vectors to obtain the total loss. The total loss is optimized by the optimizer corresponding to the constructed neural network until convergence.
[0030] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0031] In a fourth aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] The present invention uses the Bert model as the backbone to achieve the purpose of extracting text sentence vectors, and at the same time uses the classification results predicted by the Bert model as a rough result for subsequent optimization. The new network structure is bridged on the back end, which can fully utilize the text representation ability of pre-trained Bert and better optimize the classification performance in vertical fields. The present invention draws on the idea of Attention and calculates the deep semantic features of the Bert model with the probability distribution of coarse classification labels, which can effectively reduce the error in the coarse classification labels. In addition, with the help of the update of the class vector, the model will consider all known sentence vectors at the same time. Compared with the traditional classification model that can only see a batch of samples at the same time, this strategy can obtain a more accurate optimization direction. In addition, the addition of multiple layers of loss can also better record the changes in the sentence vector in the process, which plays a positive role in the convergence of the model.
[0034] At the same time, the present invention only uses the Bert model as the backbone, so in theory the Bert model can also be replaced by other text feature extraction models, such as Bi-LSTM, Elmo, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0036] Figure 1 This is a flow chart of a text classification method according to an embodiment of the present invention.
[0037] Figure 2 This is a flow chart of a text classification method according to another embodiment of the present invention.
[0038] Figure 3 FIG. 4 is a schematic diagram of a text classification system according to an embodiment of the present invention.
[0039] Figure 4 This is a flow chart of a text classification method according to another embodiment of the present invention.
[0040] Figure 5 This is a principle diagram of a text classification method according to another embodiment of the present invention. DETAILED DESCRIPTION
[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0042] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0043] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0044] In the description of the embodiments of the present invention, it should be noted that if the terms "upper," "lower," "horizontal," "inner," etc. appear, the orientation or positional relationship indicated is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship in which the inventive product is typically placed when in use. These terms are merely for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or component referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. In addition, the terms "first," "second," etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0045] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", and does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0046] In the description of the embodiments of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0047] The present invention is described in further detail below with reference to the accompanying drawings:
[0048] BERT is a pre-trained language model. A sentence isn't simply a collection of words; it must conform to human linguistic conventions. A language model is used to determine whether a sentence conforms to human conventions. Specifically, given the words preceding a word in a sentence (or the context), it predicts the probability of that word appearing. For a sentence, if the conditional probabilities of each position are multiplied, the higher the probability, the more likely it is to be spoken by a human. To train a language model, all we need to do is select a word, use its context to predict its probability, and then maximize this probability.
[0049] In addition to determining whether sentences are more consistent with human language, language models also utilize word embeddings. During model training, the first words in a sentence are converted into corresponding vectors. Then, the hidden layer and softmax are used to predict the probability of subsequent words. If the model predicts well, the vectors obtained in the first step can be considered to accurately reflect the meaning of the corresponding words. This creates a vocabulary, where each word can be assigned a corresponding vector. This process is called word embedding. After this, word embeddings can be applied to specific downstream tasks, enabling pre-training for NLP tasks through language models.
[0050] It can be seen that the training of the language model does not require pre-labeling of the data, so the corpus that can be used is quite large. This is also the reason why the language model is used for NLP pre-training. A large amount of training data makes the word embedding effect better, thereby making the downstream task model after pre-training better.
[0051] However, in actual experiments, the use of word embedding pre-training technology in downstream tasks did not achieve significant improvement. The main reason is that for word embedding, one word uniquely corresponds to one vector, but some words often correspond to different meanings in different sentences. For such polysemous words, it is impossible to use only one vector to represent them. One vector can only average each meaning of the word, which will obviously affect the subsequent model learning. Therefore, it is necessary to improve the pre-trained language model, and BERT is a good solution to this problem.
[0052] BERT (Bidirectional Encoder Representations from Transformer) is based on the Transformer's bidirectional encoder representations. BERT uses the Transformer, and when processing a word, it also considers the words preceding and following it to derive its meaning in context. The Transformer's attention mechanism is very effective in extracting features from words in context, and bidirectional encoding that considers context is more effective than unidirectional encoding that only considers the previous (or next) context. The BERT model uses the Transformer feature extractor and also implements a bidirectional language model, which gives it better performance.
[0053] See also Figure 1 The embodiment of the present invention discloses a text classification method for re-optimization using category information, comprising the following steps:
[0054] S1 inputs the text information to be classified into the trained Bert model to obtain several classification results; the text information to be classified is input into the trained Bert model to obtain several classification results, including:
[0055] The sentence to be classified is input into a trained Bert model in the form of word vectors to obtain several classification results, which are recorded as C1, C2, ..., C n , where n represents the number of categories.
[0056] S2 extracts the CLS vector of the output layer of the Bert model to obtain several sentence vectors; the sentence vectors are A1, A2, ..., A m , where m represents the number of sentences, and the category corresponding to each sentence vector is recorded as L m .
[0057] S3 calculates the sentence vector loss based on the classification results and the sentence vectors, and constructs a neural network. The calculation of the sentence vector loss includes:
[0058] According to the classification results, the sentence vector A of the same category m Find the mean and get the category vector, recorded as B1, B2, ..., B n , integration to obtain matrix B;
[0059] Attention is performed on each sentence vector and all category vectors, and the category vectors are weighted summed to obtain a new sentence vector A. ′ m :
[0060]
[0061]
[0062] Among them, α n represents the weight of B, Indicates α1~α n , i is a positive integer, s() represents a certain attention scoring mechanism, and A m Replace with A ′ m ;
[0063] A m As the input of an additional classifier, followed by a softmax layer, we get the category distribution L m , and its true label Calculate the loss of sentence vector.
[0064] The constructing of the neural network comprises:
[0065] By superimposing the loss process of calculating the sentence vector D times, the neural network is obtained, where D is the hyperparameter of the model.
[0066] S4 performs a weighted summation of the losses of several sentence vectors to obtain a total loss, and optimizes the total loss using a neural network optimizer, such as SGD or Adam, until convergence.
[0067] In this embodiment, attention is performed on each sentence vector and all category vectors. The attention mechanism is that, given a set of vector sets (values) and a vector query, the attention mechanism is a mechanism that calculates the weighted sum of the values according to the query. The focus of attention is the calculation method of the "weight" of each value in this set of values. Sometimes this attention mechanism is also called the output of the query paying attention to (or taking into account) different parts of the original text. The attention mechanism is a method of extracting specific vectors from the vector expression set (values) according to certain rules or some additional information (query) for weighted combination (attention).
[0068] like Figure 2 As shown, in another feasible embodiment of the present invention, step S2 is replaced by aggregating all word vectors of the output layer of the Bert model to obtain several sentence vectors; the sentence vectors are A1, A2, ..., A m , where m represents the number of sentences, and the category corresponding to each sentence vector is recorded as L m The aggregation adopts summation, averaging or pooling method.
[0069] like Figure 3 As shown, an embodiment of the present invention further discloses a text classification system that utilizes category information for re-optimization, including:
[0070] The classification result calculation module is used to input the text information to be classified into the trained Bert model to obtain several classification results;
[0071] The sentence vector calculation module is used to aggregate all word vectors in the output layer of the BERT model to obtain several sentence vectors;
[0072] The neural network construction module is used to calculate the loss of sentence vectors and build a neural network based on several classification results and several sentence vectors;
[0073] The optimization module is used to obtain the total loss by weighted summation of the losses of several sentence vectors, and optimize the total loss through the neural network optimizer until convergence.
[0074] In another possible embodiment, the sentence vector calculation module is used to aggregate all word vectors of the output layer of the Bert model to obtain a number of sentence vectors.
[0075] like Figure 4 As shown, an embodiment of the present invention discloses a text classification method for re-optimization using category information, comprising the following steps:
[0076] S01 inputs the sentence to be classified into a trained Bert model in the form of words and obtains the classification results, which are recorded as C1, C2, ..., C n , where n represents the number of categories;
[0077] S02 extracts the CLS vector of the output layer of the Bert model separately, or aggregates all the word vectors of the output layer (not limited to summation, averaging or other pooling methods) to obtain the sentence vector, which is recorded as A1, A2, ..., A m , where m represents the number of sentences, and the category corresponding to each sentence vector is recorded as L m ;
[0078] S03 will be the same category of sentence vector A m Find the mean and get the category vector, recorded as B1, B2, ..., B n , integration to obtain matrix B;
[0079] S04 pays attention to each sentence vector and all category vectors, and performs weighted summation of the category vectors to obtain a new sentence vector, that is, s(x,y) represents a certain attention scoring mechanism, and A mReplace with A ′ m ;
[0080] S05 simultaneously transforms the sentence vector A m As the input of an additional classifier, followed by a softmax layer, we get the category distribution L m , and its true label Calculate loss;
[0081] S06 Figure 5 As shown, steps S03-S05 are assembled into a module and stacked D times to obtain a neural network;
[0082] S07 performs weighted summation of the losses of D modules (the weight is a hyperparameter and can be determined according to the specific situation. For example, the closer the module is to the front, the lower the loss weight) to obtain the total loss. It is then optimized using the optimizer commonly used in neural networks (such as SGD, Adam, etc.) until convergence.
[0083] One embodiment of the present invention provides a computer device. The computer device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of each of the aforementioned method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of each module / unit in each of the aforementioned apparatus embodiments are implemented.
[0084] The computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to accomplish the present invention.
[0085] The computer device may be a desktop computer, a notebook computer, a PDA, a cloud server, etc. The computer device may include, but is not limited to, a processor and a memory.
[0086] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0087] The memory may be used to store the computer programs and / or modules, and the processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory.
[0088] If the module / unit integrated in the computer device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0089] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A text classification method, characterized in that: The following steps are involved: Input the text information to be classified into the Bert model to obtain several classification results; Extract the sentence vector CLS of the Bert model output layer to obtain several sentence vectors; or aggregate all the word vectors of the Bert model output layer to obtain several sentence vectors; According to the classification results, the sentence vectors of the same category are averaged to obtain the category vector. All sentence vectors are associated with all category vectors, and the category vectors are weighted summed to obtain a new sentence vector. The new sentence vector is used as the input of the additional classification, followed by a softmax layer to obtain the category distribution, and the loss of the sentence vector is calculated in combination with the true label of the category distribution. The calculated sentence vector loss is superimposed to construct a neural network; The weighted sum of the losses of several superimposed sentence vectors is used to obtain the total loss, and the total loss is optimized by the optimizer corresponding to the constructed neural network until convergence.
2. The text classification method according to claim 1, characterized in that The text information to be classified is input into the Bert model to obtain several classification results, including: Split the sentence to be classified into several sub-vectors, input the word vectors into the Bert model, and obtain several classification results, which are recorded as C1, C2, ..., C n , where n represents the number of categories.
3. The text classification method according to claim 1, characterized in that The aggregation adopts summation, averaging or pooling method.
4. The text classification method according to claim 1, characterized in that The loss of the sentence vector is calculated, including: According to the classification results, the sentence vector A of the same category m Find the mean and get the category vector, recorded as B1, B2, ..., B n , integration to obtain matrix B; Attention is associated with each sentence vector and all category vectors, and the category vectors are weighted summed to obtain the new sentence vector A′ m : Among them, α n represents the weight of B, Indicates α1~α n , i is a positive integer, s() represents a certain attention scoring mechanism, and A m Replace with A ′ m ; A m As the input of an additional classifier, followed by a softmax layer, we get the category distribution L m , and its true label Calculate the loss of sentence vector.
5. The text classification method according to claim 1, characterized in that The neural network optimizer adopts the stochastic gradient descent optimizer SGD or the adaptive moment estimation optimizer Adam.
6. A text classification system, characterized in that include: The classification result calculation module is used to input the text information to be classified into the Bert model to obtain several classification results; The sentence vector calculation module is used to extract the sentence vector CLS of the Bert model output layer to obtain several sentence vectors; or to aggregate all word vectors of the Bert model output layer to obtain several sentence vectors; The loss calculation module is used to calculate the average of the sentence vectors of the same category based on the classification results to obtain the category vector, associate all sentence vectors with all category vectors, and perform weighted summation on the category vectors to obtain a new sentence vector. The new sentence vector is used as the input for additional classification, followed by a softmax layer to obtain the category distribution, and the loss of the sentence vector is calculated in combination with the true label of the category distribution. The neural network construction module is used to superimpose the loss of the calculated sentence vectors to construct a neural network; The optimization module is used to perform weighted summation of the losses of several superimposed sentence vectors to obtain the total loss. The total loss is optimized by the optimizer corresponding to the constructed neural network until convergence.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Long text classification method and device, computer equipment and storage medium
CN110929033A
Method and system for constructing text classification system, medium and electronic equipment
CN111966826A