Text classification method and device, computer equipment and storage medium
By combining lightweight convolutional neural network, hidden Markov model and decision tree model, the problem of large amount of text classification calculation in the existing technology is solved, efficient and accurate text classification is achieved, and overfitting is avoided.
Patent Information
- Application Number
- CN202411764955.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-03
AI Technical Summary
In the prior art, when text classification is performed using Transformer model and other neural networks with complex structures such as Bert, the calculation amount is large and difficult to reduce.
A lightweight target convolutional neural network (CNN) is used to combine the hidden Markov model (HMM) and the decision tree model to extract features through the word embedding model, and the convolutional neural network extracts multi-scale features. The hidden Markov model captures context information, and the decision tree model performs final classification.
It effectively reduces the calculation amount of text classification, and improves the prediction accuracy, avoids overfitting, and improves the interpretability of the model.
Smart Images

Figure CN119938910A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a text classification method, apparatus, computer equipment and storage medium. Background Art
[0002] Text classification is an important task in natural language processing. It can be applied to a variety of scenarios and industries, such as news classification, sentiment analysis, spam filtering, advertising positioning, product positioning, content recommendation, and research literature classification, etc. It has very important research value in practical applications.
[0003] In the related art, complex neural networks such as the Transformer model and Bert are used for text classification. Although a high text classification accuracy can be achieved, the number of parameters in complex neural networks such as the Transformer model and Bert is large, and the amount of computation required for text classification using complex neural networks such as the Transformer model and Bert is large. How to reduce the amount of computation required for text classification has become a problem that needs to be solved. Summary of the invention
[0004] In view of this, embodiments of the present disclosure provide a text classification method, apparatus, computer device, and storage medium.
[0005] In a first aspect, an embodiment of the present disclosure provides a text classification method, the method comprising:
[0006] Get the target text;
[0007] Using the target text classification model, the classification results of the target text are obtained according to the target text, including:
[0008] Using the word embedding model, according to the target text, the word embedding representation corresponding to the target text is obtained;
[0009] Using the target convolutional neural network, according to the word embedding representation corresponding to the target text, the multi-scale features of the target text are obtained, wherein using the target convolutional neural network, according to the word embedding representation corresponding to the target text, the multi-scale features of the target text are obtained including: according to the input of the target convolutional layer in the target convolutional neural network, the target convolutional layer is any one of the convolutional layers in the target convolutional neural network, and according to the input of the target convolutional layer in the target convolutional neural network, the target output channel of the target convolutional layer is obtained including: for each of the multiple convolutional kernels corresponding to the position of the target output channel of the target convolutional layer, using the convolutional kernel to perform a convolution operation on the input of the target convolutional layer, and obtain a sub-output channel of the target output channel corresponding to the convolutional kernel, wherein each of the convolutional kernels has a different convolutional kernel size, and wherein the target output channel is any one of the output channels of the target convolutional layer;
[0010] For each hidden Markov model in a group of hidden Markov models, using the hidden Markov model, according to the multi-scale features of the target text, obtain a prediction result corresponding to the target text output by the hidden Markov model;
[0011] A decision tree model is used to obtain a classification result of the target text according to the prediction result corresponding to the target text output by each hidden Markov model.
[0012] In a possible implementation, using a word embedding model, according to the target text, obtaining a word embedding representation corresponding to the target text includes:
[0013] For each word embedding model among the multiple word embedding models, using the word embedding model, according to the target text, obtain a word embedding matrix corresponding to the target text output by the word embedding model;
[0014] According to the word embedding matrix corresponding to the target text output by each word embedding model, the word embedding representation corresponding to the target text is obtained.
[0015] In a possible implementation, obtaining the word embedding representation corresponding to the target text according to the word embedding matrix corresponding to the target text output by each word embedding model includes:
[0016] According to the word embedding matrix corresponding to the target text output by each word embedding model and the adaptive parameters corresponding to each word embedding model, the word embedding representation corresponding to the target text is obtained.
[0017] In a possible implementation, using the target convolutional neural network, according to the word embedding representation corresponding to the target text, the multi-scale features of the target text are obtained, which also includes:
[0018] Using a pooling layer connected to the target convolutional layer, K-Max pooling is performed on each sub-output channel of the target output channel to obtain a pooling result of each sub-output channel of the target output channel;
[0019] According to the pooling result of each sub-output channel of the target output channel and the adaptive parameter corresponding to each sub-output channel of the target output channel, an input channel in the input of the next unit of the convolution unit to which the target convolution layer belongs is obtained.
[0020] In a second aspect, an embodiment of the present disclosure provides a text classification device, the text classification device comprising:
[0021] An acquisition unit, used for acquiring a target text;
[0022] A classification unit is used to obtain a classification result of the target text according to the target text using a target text classification model. The classification result of the target text obtained according to the target text using the target text classification model includes: obtaining a word embedding representation corresponding to the target text according to the target text using a word embedding model; obtaining a multi-scale feature of the target text according to the word embedding representation corresponding to the target text using a target convolutional neural network, wherein obtaining a multi-scale feature of the target text according to the word embedding representation corresponding to the target text using a target convolutional neural network includes: obtaining a target output channel of the target convolutional layer according to an input of the target convolutional layer in the target convolutional neural network, wherein the target convolutional layer is any convolutional layer in the target convolutional neural network, and obtaining the target according to the input of the target convolutional layer in the target convolutional neural network The target output channel of the convolution layer includes: for each convolution kernel among multiple convolution kernels corresponding to the position of the target output channel of the target convolution layer, using the convolution kernel to perform a convolution operation on the input of the target convolution layer to obtain a sub-output channel of the target output channel corresponding to the convolution kernel, wherein each convolution kernel has a different convolution kernel size, wherein the target output channel is any output channel of the target convolution layer; for each hidden Markov model in a group of hidden Markov models, using the hidden Markov model, according to the multi-scale features of the target text, obtaining a prediction result corresponding to the target text output by the hidden Markov model; using a decision tree model, according to the prediction result corresponding to the target text output by each hidden Markov model, obtaining a classification result of the target text.
[0023] In one possible implementation, the classification unit is further used to, for each word embedding model in the multiple word embedding models, use the word embedding model to obtain, according to the target text, a word embedding matrix corresponding to the target text output by the word embedding model; and obtain the word embedding representation corresponding to the target text according to the word embedding matrix corresponding to the target text output by each word embedding model.
[0024] In a possible implementation, the classification unit is further used to obtain the word embedding representation corresponding to the target text based on the word embedding matrix corresponding to the target text output by each word embedding model and the adaptive parameters corresponding to each word embedding model.
[0025] In one possible implementation, the classification unit is further used to use a pooling layer connected to the target convolutional layer to perform K-Max pooling on each sub-output channel of the target output channel, respectively, to obtain a pooling result of each sub-output channel of the target output channel; and according to the pooling result of each sub-output channel of the target output channel and the adaptive parameters corresponding to each sub-output channel of the target output channel, obtain an input channel in the input of the next unit of the convolutional unit to which the target convolutional layer belongs.
[0026] In a third aspect, an embodiment of the present disclosure provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, computer instructions being stored in the memory, and the processor executing the method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0027] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method of the first aspect or any corresponding implementation manner thereof.
[0028] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the method of the first aspect or any corresponding embodiment thereof.
[0029] The text classification method provided by the embodiment of the present disclosure, on the one hand, uses a much smaller number of parameters of a lightweight target convolutional neural network and HMM compared to the Transformer and Bert models, thereby reducing the amount of computation for text classification. At the same time, the target convolutional neural network can extract rich multi-scale structural features from text data, and the HMM can capture contextual information and perform preliminary classification by determining the evolution law of the hidden state behind the features. The HMM utilizes its efficient learning and reasoning capabilities. At the same time, the efficient learning and reasoning capabilities of the HMM fully utilize the powerful feature extraction capabilities of the neural network model and the efficient learning and reasoning capabilities of the HMM, ensuring a high prediction accuracy.
[0030] On the other hand, the Hidden Markov Model (HMM) mines global contextual information and obtains preliminary prediction results by determining the evolution law of the hidden state behind the features. For different categories of text data, HMMs of corresponding categories are trained, and all trained HMMs constitute an HMM group. Finally, the HMM group is deeply embedded with the Stacming ensemble learning mechanism. Specifically, a group of HMMs is used as a meta-learner, and the decision tree model is used as a top-level learner. The decision tree model makes the final category judgment based on the preliminary prediction results obtained by the HMM group. This ensemble learning mechanism naturally utilizes the characteristics of each HMM in the HMM group, integrates all models in the HMM group for group decision-making, and can effectively avoid the occurrence of overfitting while improving the prediction accuracy. In particular, the parameters and state transition probabilities in the HMM have certain semantics and can be used to explain the model's decisions, which improves the interpretability of the model to a certain extent. HMM is used to capture high-level contextual information, and preliminary classification is performed by determining the evolution law of the hidden state behind the features. The parameters and state transition probabilities in HMM have certain semantics and can be used to explain the model's decisions, which improves the interpretability of the model to a certain extent.
[0031] On the other hand, the differences between each HMM in a group of HMMs are naturally utilized, and multiple models are integrated together through the Stacming ensemble learning mechanism for group decision-making, which can effectively solve the problem of overfitting while improving the prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0033] Figure 1 is a flowchart of a text classification method provided by an embodiment of the present disclosure;
[0034] Figure 2 is a schematic diagram of an example of the structure of a target convolutional neural network;
[0035] Figure 3 is a flowchart of an example of obtaining a prediction result of an HMM output in a set of HMMs;
[0036] Figure 4 It is a flowchart of an example of obtaining the classification result of the target text by using the target text classification model;
[0037] Figure 5It is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0039] refer to Figure 1 , which shows an example flow chart of the text classification method provided by an embodiment of the present disclosure.
[0040] In step S101, the target text is obtained.
[0041] In step S102, the target text classification model is used to obtain the classification result of the target text according to the target text.
[0042] The target text classification model includes: a word embedding model, a target convolutional neural network, a set of hidden Markov models (HMM), and a decision tree model. Each HMM in a set of HMMs in the target text classification model can correspond to different text categories. Each HMM in a set of HMMs in the target text classification model has corresponding parameters. As an example, HMM_1 in a set of HMMs has a parameter λ1, HMM_2 in a set of HMMs has a parameter λ2, and HMM_M in a set of HMMs has a parameter M.
[0043] Step S102 includes: step S1021-step S1024.
[0044] In step S1021, a word embedding model is used to obtain a word embedding representation corresponding to the target text according to the target text.
[0045] In a possible implementation, the word embedding model is one of a Word2Vec model, a GloVe model, a FastText model, etc. The target text is input into the word embedding model to obtain a word embedding representation corresponding to the target text.
[0046] In another possible implementation, using a word embedding model to obtain a word embedding representation corresponding to the target text according to the target text includes: for each word embedding model in multiple word embedding models, using the word embedding model to obtain a word embedding matrix corresponding to the target text output by the word embedding model; obtaining a word embedding representation corresponding to the target text according to the word embedding matrix corresponding to the target text output by each word embedding model.
[0047] The word embedding matrix corresponding to the target text output by each word embedding model can be calculated accordingly, for example, the corresponding matrix elements in the word embedding matrix corresponding to the target text output by each word embedding model can be added in opposite positions to obtain the word embedding representation corresponding to the target text.
[0048] As an example, multiple word embedding models may include: Word2Vec model, GloVe model, FastText model. The word embedding model may perform word vectorization operation on the target text to obtain a word embedding matrix, and each row in the word embedding matrix represents a word in the text data. Padding and truncation operations are performed on the word embedding matrix corresponding to the target text output by the word embedding model to obtain a matrix of the dimension of the input of the convolutional neural network that meets the requirements.
[0049] In one possible implementation, obtaining the word embedding representation corresponding to the target text according to the word embedding matrix corresponding to the target text output by each word embedding model includes: obtaining the word embedding representation corresponding to the target text according to the word embedding matrix corresponding to the target text output by each word embedding model and the adaptive parameters corresponding to each word embedding model.
[0050] The adaptive parameters corresponding to each word embedding model can be used as the weight of each word embedding model, and the weighted sum of the word embedding matrices corresponding to the target text output by each word embedding model can be determined as the word embedding representation corresponding to the target text.
[0051] In the disclosed embodiment, by setting the adaptive parameters corresponding to each word embedding model, it is possible to ensure the generalization ability of the algorithm and allow the model to approach a more appropriate word embedding matrix in a learning manner during the training process.
[0052] In step S1022, the target convolutional neural network is used to obtain multi-scale features of the target text according to the word embedding representation corresponding to the target text.
[0053] It should be noted that the target convolutional neural network provided in the embodiments of the present disclosure can be called mTextCNN.
[0054] The word embedding representation corresponding to the target text is used as the input of the target convolutional neural network. In step S1022, the word embedding representation corresponding to the target text is input to the target convolutional neural network, and the target convolutional neural network outputs the multi-scale features of the target text.
[0055] Step S1022 includes: step S10221.
[0056] In step S10221, using the target convolutional neural network, according to the word embedding representation corresponding to the target text, the multi-scale features of the target text are obtained, including: according to the input of the target convolutional layer in the target convolutional neural network, the target output channel of the target convolutional layer in the target convolutional neural network is obtained.
[0057] Among them, according to the input of the target convolutional layer in the target convolutional neural network, obtaining the target output channel of the target convolutional layer includes: for each convolution kernel among the multiple convolution kernels corresponding to the position of the target output channel of the target convolutional layer, using the convolution kernel to perform a convolution operation on the input of the target convolutional layer, to obtain a sub-output channel of the target output channel of the target convolutional layer corresponding to the convolution kernel, wherein each of the multiple convolution kernels corresponding to the position of the target output channel of the target convolutional layer has a different convolution kernel size; according to each sub-output channel of the target output channel of the target convolutional layer, obtaining the target output channel of the target convolutional layer.
[0058] The output of the target convolutional layer consists of the output channels of the target convolutional layer, and the output channels of the target convolutional layer include the features of the output of the target convolutional layer.
[0059] The target output channel is any output channel of the target convolution layer. The position indication of the target output channel of the target convolution layer: the output channel number of the target output channel of the target convolution layer.
[0060] In an embodiment of the present disclosure, the target convolutional neural network includes one or more convolutional units. The target convolutional neural network includes a fully connected layer. In order to prevent the occurrence of overfitting problems, a random failure module is added to the fully connected network, and the activation function adopts ReLU. A convolutional unit in the target convolutional neural network includes: a convolutional layer, and a pooling layer connected to the convolutional layer. When the target convolutional neural network includes multiple convolutional units, the number of convolutional units in the target convolutional neural network is recorded as n, and the output of the i-th convolutional unit is used as the input of the i+1-th convolutional unit, where i is one of 1...n-1. The input of the convolutional layer in the first convolutional unit in the target convolutional neural network is: the input of the target convolutional neural network, and the output of the last convolutional unit is the input of the fully connected layer in the target convolutional neural network.
[0061] In the embodiment of the present disclosure, for each of the multiple convolution kernels corresponding to the position of the output channel h of the convolution layer j, a convolution operation is performed on the input of the convolution layer j using the convolution kernel to obtain a sub-output channel of the output channel h corresponding to the convolution kernel, wherein each of the multiple convolution kernels corresponding to the position of the output channel h of the convolution layer j has a different convolution kernel size.
[0062] Thus, each sub-output channel of channel h constitutes output channel h.
[0063] As an example, the multiple convolution kernels corresponding to the position of the output channel h of the convolution layer j are composed of a 3×3 convolution kernel corresponding to the position of the output channel m of the convolution layer j, a 5×5 convolution kernel corresponding to the position of the output channel m of the convolution layer j, and a 7×7 convolution kernel corresponding to the position of the output channel h of the convolution layer j. The 3×3 convolution kernel corresponding to the position of the output channel h of the convolution layer j is used to perform a convolution operation on the input of the convolution layer j, and a sub-output channel of the output channel h corresponding to the 3×3 convolution kernel is obtained. The 5×5 convolution kernel corresponding to the position of the output channel h of the convolution layer j is used to perform a convolution operation on the input of the convolution layer j, and a sub-output channel of the output channel h corresponding to the 5×5 convolution kernel is obtained. The 7×7 convolution kernel corresponding to the position of the output channel h of the convolution layer j is used to perform a convolution operation on the input of the convolution layer j, and a sub-output channel of the output channel h corresponding to the 7×7 convolution kernel is obtained.
[0064] In a possible implementation, using the target text classification model, obtaining the classification result of the target text according to the target text also includes: using a pooling layer connected to the target convolutional layer to perform K-Max pooling on each sub-output channel of the target output channel of the target convolutional layer, respectively, to obtain the pooling result of each sub-output channel of the target output channel of the target convolutional layer; according to the pooling result of each sub-output channel of the target output channel of the target convolutional layer and the adaptive parameters corresponding to each sub-output channel of the target output channel of the target convolutional layer, obtaining the input channel in the input of the next unit of the convolutional unit to which the target convolutional layer belongs.
[0065] The adaptive parameter corresponding to the sub-output channel of the target output channel of the target convolution layer can be used as the weight of the sub-output channel of the target output channel of the target convolution layer, and the weighted sum of the pooling results of each sub-output channel of the target output channel of the target convolution layer is determined as the input channel in the input of the next unit of the convolution unit to which the target convolution layer belongs.
[0066] The pooling layer connected to the target convolution layer may be: a convolution kernel in a convolution unit to which the target convolution layer belongs.
[0067] If the convolution unit to which the target convolution layer belongs is the last convolution unit, the next unit of the convolution unit to which the target convolution layer belongs is the fully connected layer of the target convolution neural network. If the convolution unit to which the target convolution layer belongs is not the last convolution unit, the next unit of the convolution unit to which the target convolution layer belongs is the next convolution unit of the convolution unit to which the target convolution layer belongs that is connected to the convolution unit to which the target convolution layer belongs.
[0068] In the disclosed embodiment, the feature expression capability and robustness of the model are further improved by the adaptive parameters corresponding to each sub-output channel of the target output channel. The target convolutional neural network can extract rich multi-scale features with fewer computing resources. The K-Max pooling operation is different from the traditional maximum pooling operation. The top M eigenvalues in the feature vector are retained, which can further enhance the robustness of the model. After the pooling operation, an adaptive parameter is assigned to the feature vectors obtained by the convolutional layers of different scales, and then a splicing operation is performed to obtain a combined feature vector. In particular, the adaptive parameters can improve the feature expression capability of the algorithm.
[0069] refer to Figure 2 , which is a schematic diagram showing an example of the structure of a target convolutional neural network.
[0070] Figure 2 The multi-channel word embedding matrix, i.e., the word embedding representation corresponding to the target text, is shown.
[0071] Figure 2 A convolution layer of a target convolutional neural network, a 3×k convolution kernel corresponding to the position of an output channel of the convolution layer, a 5×k convolution kernel corresponding to the position of the output channel of the convolution layer, and a 7×k convolution kernel corresponding to the position of the output channel of the convolution layer are shown. The 3×k convolution kernel corresponding to the position of the output channel of the convolution layer is used to perform a convolution operation on the input of the convolution layer to obtain a sub-output channel of the output channel corresponding to the 3×k convolution kernel. The 5×k convolution kernel corresponding to the position of the output channel of the convolution layer is used to perform a convolution operation on the input of the convolution layer to obtain a sub-output channel of the output channel corresponding to the 5×k convolution kernel. The 7×k convolution kernel corresponding to the position of the output channel of the convolution layer is used to perform a convolution operation on the input of the convolution layer to obtain a sub-output channel of the output channel corresponding to the 7×k convolution kernel.
[0072] The convolution unit to which this layer belongs is the last convolution unit, and the next unit of the convolution unit to which this convolution layer belongs is the fully connected layer of the target convolutional neural network.
[0073] The adaptive parameter corresponding to the sub-output channel of the target output channel of the target convolution layer is used as the weight of the sub-output channel of the target output channel of the target convolution layer, and the weighted sum of the pooling results of each sub-output channel of the target output channel of the target convolution layer is calculated.
[0074] Figure 2 The weighted sum of the pooling results of each sub-output channel of the target output channel of the target convolutional layer, that is, the combined feature vector combined with the adaptive parameters, is shown.
[0075] The feature vector combined with the adaptive parameters is used as the input channel in the input of the fully connected layer of the target convolutional neural network.
[0076] In step S1023, for each HMM in a set of HMMs in the target text classification model, the HMM is used to obtain a prediction result corresponding to the target text output by the HMM according to the multi-scale features of the target text.
[0077] For each HMM in a set of HMMs in the target text classification model, the multi-scale features of the target text are input into the HMM to obtain a prediction result corresponding to the target text output by the HMM.
[0078] In step S1024, the decision tree model in the target text classification model is used to obtain the classification result of the target text according to the prediction result corresponding to the target text output by each hidden Markov model.
[0079] The decision tree model in the target text classification model receives the prediction result corresponding to the target text output by each hidden Markov model, and the decision tree model in the target text classification model outputs the classification result of the target text.
[0080] In an embodiment of the present disclosure, before the text classification method provided by the embodiment of the present disclosure is executed for the first time, the mTextCNN model is trained using a text training set. During the training of the target convolutional neural network, the training text is input into the mTextCNN model, and the output of the mTextCNN model is used as the input of the Softmax layer, which outputs the probability of each text category. Using the cross entropy function, the loss corresponding to the training text is calculated based on the probability of each text category output by the Softmax layer and the label of the training text. Using the back propagation algorithm, the parameters of the mTextCNN network are updated based on the losses corresponding to multiple training texts.
[0081] In the disclosed embodiment, after the target convolutional neural network training is completed, a set of HMMs is trained using the training corpus.
[0082] For the sake of convenience, the following will explain how HMM works by introducing the training and testing process of HMM. In the previous step, the feature vector output by the mTextCNN model is input into the HMM as an observation sequence.
[0083] In the equation, o=[o 1 ,o 2 ,···,o M ](1)
[0084] It is assumed that there are M observation states and N hidden states in the Marmov chain. During the training phase, the parameters of the HMM corresponding to each type of text are initialized to λ = (π, A, B)
[0085] π=[π n ] represents the initial probability value of each hidden state;
[0086] π n =P(q t =s n ),1≤n≤N
[0087] Among them, q t represents the hidden state variable. In addition, A=[a nh ] N×N represents the state transfer matrix;
[0088] a nh =P(q t+1 =s h |q t =s n ),1≤n,h≤N
[0089] B=[b nm ] N×M represents the observation probability matrix, b nm =P(p t =o m |q t =s n ),1≤n≤N,1≤m≤M, where p t represents observed variables.
[0090] The main task of training HMM is to use the feature vector output by mTextCNN as the observation sequence to update the parameter λ. The training set includes training texts of N text types. During the training of HMM, the training texts in the training text set are input into mTextCNN. The training text set of HMM includes: training texts of N text categories. The multi-scale features of the training texts output by mTextCNN, i.e., the multi-scale feature vectors, are input into HMM as observation vectors. In HMM, the parameter λ corresponding to the highest probability p(o|λ) is obtained according to the Baum-Welch algorithm. It can be considered that the probability p(o|λ) reaches the maximum value when the difference in the probability values obtained in two consecutive iterations is less than a predetermined threshold. For each type of text in the training text set, the optimal parameters of HMM are calculated using the same method as above to obtain the matching HMM corresponding to the text of that type. It can also be said that for each type of text data, an optimal HMM will be trained. It can also be said that the training text set involves each text type, and the training text of that text type is used for training to obtain the HMM corresponding to that text type. Among them, for a text type, if the training text set includes training text of this text type, then this text type is the text type involved in the training text set. If there are N different types of text data in the training set data set, then N HMMs can be obtained, and the N HMMs constitute a group of HMMs. When a new training text is input into a group of HMMs, a group of HMMs will output N prediction probabilities, and the N prediction probabilities correspond to the N text types respectively. Finally, the HMMs corresponding to the training texts of all categories in the training text set constitute a group of HMMs, which are used to distinguish text data of all text categories.
[0091] After a group of HMMs are trained, the decision tree model is trained using the Stacming ensemble learning method. Label inference module based on the Stacming ensemble learning mechanism. Ensemble learning is a technique that uses multiple learners or classifiers to perform the same task and uses a certain strategy to combine the results of all learners or classifiers to make the final decision. At present, common ensemble learning types include Bagging (Bootstrap Aggregating), Boosting, Bayesian Model Combination, and Stacming. There are two reasons for choosing the Stacming ensemble learning mechanism. First, Stacming ensemble learning can effectively suppress the occurrence of overfitting. Second, various types of HMMs in a group of HMMs are natural meta-learners, which fits the framework of the Stacming ensemble learning model. Assuming that a group of HMMs consists of M different groups of HMMs, the number of meta-learners in Stacming is M. For each input sample, a group of HMMs will obtain M different probability values, and these probability values and the labels of the samples will be provided to the top-level learner as training data.
[0092] For any sample in the training set, a set of HMMs will output N predicted probabilities. The training data of the decision tree model is a set of N predicted values of HMMs and the true labels of the samples. These data can be used to complete the training of the decision tree model. The decision tree model uses the 5-fold cross-validation method for training and testing.
[0093] After completing the training of mTextCNN, a set of HMMs, and a decision tree model, testing is performed. During the test, the data in the test set is input into the mTextCNN model, which extracts multi-scale features from the text data. The multi-scale features, i.e., multi-scale feature vectors, are input into each HMM in a set of HMMs as an observation sequence, and a set of HMMs will use a forward-backward algorithm to obtain multiple prediction values. Each type of HMM in a set of HMMs will be combined with the observation sequence, and the forward-backward algorithm will be used to obtain the maximum posterior probability corresponding to each text category, and then the preliminary classification results will be obtained. Finally, the preliminary prediction results of a set of HMMs are input into the decision tree model for the final category judgment.
[0094] refer to Figure 3 , which shows a flowchart of an example of obtaining a prediction result of an HMM output in a set of HMMs.
[0095] Input text data into mTextCNN. mTextCNN outputs a multi-scale feature vector of text data. Input the multi-scale feature vector output by mTextCNN into a set of HMMs. The prediction result output by HMM_1 in a set of HMMs is a probability value of 1. The prediction result output by HMM_2 in a set of HMMs is a probability value of 2. The prediction result output by HMM_M in a set of HMMs is a probability value of M. HMM_1 in a set of HMMs has a parameter λ1, HMM_2 in a set of HMMs has a parameter λ2, and HMM_M in a set of HMMs has a parameter M.
[0096] refer to Figure 4 , which shows a flowchart of an example of obtaining the classification result of the target text using the target text classification model.
[0097] Input the target text into mTextCNN in the target text classification model. mTextCNN outputs the multi-scale feature vector of the target text. Input the multi-scale feature vector output by mTextCNN into each HMM in a set of HMMs in the target text classification model. A set of HMMs in the target text classification model consists of HMM_0 corresponding to category 0, HMM_1 corresponding to category 1…HMM_1_M corresponding to category M. Category 0, category 1…category M are all text categories.
[0098] The decision tree model in the target text classification model receives the prediction results output by each HMM in a set of HMMs in the target text classification model, and predicts the classification results of the target text according to the prediction results output by each HMM in a set of HMMs in the target text classification model.
[0099] The embodiments of the present disclosure provide a text classification device. The device is used to implement the above embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "unit" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0100] The text classification device comprises:
[0101] An acquisition unit, used for acquiring a target text;
[0102] A classification unit is used to obtain a classification result of the target text according to the target text using a target text classification model. The classification result of the target text obtained according to the target text using the target text classification model includes: obtaining a word embedding representation corresponding to the target text according to the target text using a word embedding model; obtaining a multi-scale feature of the target text according to the word embedding representation corresponding to the target text using a target convolutional neural network, wherein obtaining a multi-scale feature of the target text according to the word embedding representation corresponding to the target text using a target convolutional neural network includes: obtaining a target output channel of the target convolutional layer according to an input of the target convolutional layer in the target convolutional neural network, wherein the target convolutional layer is any convolutional layer in the target convolutional neural network, and obtaining the target according to the input of the target convolutional layer in the target convolutional neural network The target output channel of the convolution layer includes: for each convolution kernel among multiple convolution kernels corresponding to the position of the target output channel of the target convolution layer, using the convolution kernel to perform a convolution operation on the input of the target convolution layer to obtain a sub-output channel of the target output channel corresponding to the convolution kernel, wherein each convolution kernel has a different convolution kernel size, wherein the target output channel is any output channel of the target convolution layer; for each hidden Markov model in a group of hidden Markov models, using the hidden Markov model, according to the multi-scale features of the target text, obtaining a prediction result corresponding to the target text output by the hidden Markov model; using a decision tree model, according to the prediction result corresponding to the target text output by each hidden Markov model, obtaining a classification result of the target text.
[0103] In one possible implementation, the classification unit is further used to, for each word embedding model in the multiple word embedding models, use the word embedding model to obtain, according to the target text, a word embedding matrix corresponding to the target text output by the word embedding model; and obtain the word embedding representation corresponding to the target text according to the word embedding matrix corresponding to the target text output by each word embedding model.
[0104] In a possible implementation, the classification unit is further used to obtain the word embedding representation corresponding to the target text based on the word embedding matrix corresponding to the target text output by each word embedding model and the adaptive parameters corresponding to each word embedding model.
[0105] In one possible implementation, the classification unit is further used to use a pooling layer connected to the target convolutional layer to perform K-Max pooling on each sub-output channel of the target output channel, respectively, to obtain a pooling result of each sub-output channel of the target output channel; and according to the pooling result of each sub-output channel of the target output channel and the adaptive parameters corresponding to each sub-output channel of the target output channel, obtain an input channel in the input of the next unit of the convolutional unit to which the target convolutional layer belongs.
[0106] In this embodiment, the device is presented in the form of a functional unit, where the unit refers to an ASIC circuit, a processor and a memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0107] The further functional description of each of the above units is the same as that of the above corresponding embodiments and will not be repeated here.
[0108] refer to Figure 5 , which shows a schematic diagram of the structure of a computer device provided by an embodiment of the present disclosure, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system).
[0109] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0110] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0111] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0112] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0113] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means.
[0114] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0115] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0116] A portion of the embodiments of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes, but is not limited to, source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0117] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A text classification method, characterized in that: The method comprises: Get the target text; Using the target text classification model, the classification results of the target text are obtained according to the target text, including: Using the word embedding model, we can obtain the word embedding representation corresponding to the target text according to the target text. Using the target convolutional neural network, according to the word embedding representation corresponding to the target text, the multi-scale features of the target text are obtained, wherein using the target convolutional neural network, according to the word embedding representation corresponding to the target text, the multi-scale features of the target text are obtained including: according to the input of the target convolutional layer in the target convolutional neural network, the target convolutional layer is any one of the convolutional layers in the target convolutional neural network, and according to the input of the target convolutional layer in the target convolutional neural network, the target output channel of the target convolutional layer is obtained including: for each of the multiple convolutional kernels corresponding to the position of the target output channel of the target convolutional layer, using the convolutional kernel to perform a convolution operation on the input of the target convolutional layer, and obtain a sub-output channel of the target output channel corresponding to the convolutional kernel, wherein each of the convolutional kernels has a different convolutional kernel size, and wherein the target output channel is any one of the output channels of the target convolutional layer; For each hidden Markov model in a group of hidden Markov models, using the hidden Markov model, according to the multi-scale features of the target text, obtain a prediction result corresponding to the target text output by the hidden Markov model; A decision tree model is used to obtain a classification result of the target text according to the prediction result corresponding to the target text output by each hidden Markov model.
2. The method according to claim 1, characterized in that Using the word embedding model, according to the target text, the word embedding representation corresponding to the target text is obtained, including: For each word embedding model among the multiple word embedding models, using the word embedding model, according to the target text, obtain a word embedding matrix corresponding to the target text output by the word embedding model; According to the word embedding matrix corresponding to the target text output by each word embedding model, the word embedding representation corresponding to the target text is obtained.
3. The method according to claim 2, characterized in that According to the word embedding matrix corresponding to the target text output by each word embedding model, obtaining the word embedding representation corresponding to the target text includes: According to the word embedding matrix corresponding to the target text output by each word embedding model and the adaptive parameters corresponding to each word embedding model, the word embedding representation corresponding to the target text is obtained.
4. The method according to claim 1, characterized in that: Using the target text classification model, according to the target text, the classification results of the target text also include: Using a pooling layer connected to the target convolutional layer, K-Max pooling is performed on each sub-output channel of the target output channel to obtain a pooling result of each sub-output channel of the target output channel; According to the pooling result of each sub-output channel of the target output channel and the adaptive parameter corresponding to each sub-output channel of the target output channel, an input channel in the input of the next unit of the convolution unit to which the target convolution layer belongs is obtained.
5. A text classification device, characterized in that: The device comprises: An acquisition unit, used for acquiring a target text; A classification unit is used to obtain a classification result of the target text according to the target text using a target text classification model. The classification result of the target text obtained according to the target text using the target text classification model includes: obtaining a word embedding representation corresponding to the target text according to the target text using a word embedding model; obtaining a multi-scale feature of the target text according to the word embedding representation corresponding to the target text using a target convolutional neural network, wherein obtaining a multi-scale feature of the target text according to the word embedding representation corresponding to the target text using a target convolutional neural network includes: obtaining a target output channel of the target convolutional layer according to an input of the target convolutional layer in the target convolutional neural network, wherein the target convolutional layer is any convolutional layer in the target convolutional neural network, and obtaining the target according to the input of the target convolutional layer in the target convolutional neural network The target output channel of the convolution layer includes: for each convolution kernel among multiple convolution kernels corresponding to the position of the target output channel of the target convolution layer, using the convolution kernel to perform a convolution operation on the input of the target convolution layer to obtain a sub-output channel of the target output channel corresponding to the convolution kernel, wherein each convolution kernel has a different convolution kernel size, wherein the target output channel is any output channel of the target convolution layer; for each hidden Markov model in a group of hidden Markov models, using the hidden Markov model, according to the multi-scale features of the target text, obtaining a prediction result corresponding to the target text output by the hidden Markov model; using a decision tree model, according to the prediction result corresponding to the target text output by each hidden Markov model, obtaining a classification result of the target text.
6. The device according to claim 5, characterized in that The classification unit is also used to obtain, for each word embedding model in the multiple word embedding models, a word embedding matrix corresponding to the target text output by the word embedding model according to the target text; and obtain the word embedding representation corresponding to the target text according to the word embedding matrix corresponding to the target text output by each word embedding model.
7. The device according to claim 6, characterized in that The classification unit is also used to obtain the word embedding representation corresponding to the target text according to the word embedding matrix corresponding to the target text output by each word embedding model and the adaptive parameters corresponding to each word embedding model.
8. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 4 by executing the computer instructions.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 4.
10. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
HMM and decision tree-based Arabic optical alphabet recognition method
CN105023028A
Improved out-of-domain (OOD) detection techniques
CN115398437A
Data classification system based on AI
CN117909507A
Fault prediction method based on AI question-answering system
CN118484763A
Systems and methods for securing information
US20210014205A1