A Chinese text classification method based on deep learning
By combining Word2Vec and LDA word vectors, and using self-attention mechanism and RNN networks in the encoding stage, deep learning methods solve the problem that traditional machine learning is difficult to extract text semantic features, improving the accuracy of text classification and information processing efficiency.
Patent Information
- Application Number
- CN202210614817.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-05-31
AI Technical Summary
Traditional machine learning methods are difficult to extract the semantic features of text, resulting in low accuracy in text classification.
A Chinese text classification method based on deep learning is adopted, combined with Word2Vec word vectors and LDA word vectors as word embeddings, and a self-attention mechanism and RNN network are introduced in the encoding stage to deeply extract text features.
It improves the accuracy of text classification, can process and understand information more effectively, thereby improving the efficiency of information processing.
Smart Images

Figure CN114912461B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence, and in particular to a Chinese text classification method based on deep learning. Background Art
[0002] With the explosive growth of online information in recent years, complex text information has flooded people's lives. In order to improve people's efficiency in processing and understanding information, some people use text classification technology to classify text information, thereby improving reading efficiency. However, text classification models based on traditional machine learning, such as hidden Markov models and decision tree models, are difficult to deeply extract semantic features from texts and cannot meet the requirements of high accuracy in text classification. In recent years, some scholars have proposed text classification based on deep neural networks, which has made certain improvements in classification effects compared to traditional machine learning methods. Summary of the invention
[0003] In order to solve the problem that traditional machine learning methods cannot extract semantic features, resulting in low classification accuracy, the present invention proposes a Chinese text classification method based on deep learning to improve classification accuracy. The method combines Word2Vec word vectors and LDA (Latent Dirichlet Allocation) word vectors as word embeddings to enhance text topic features; in the encoding stage, the self-attention mechanism and the RNN (Recurrent Neural Network) network are combined to achieve deep semantic feature extraction of text while retaining sequence features; the method of the present invention improves the deep learning text classification method, can improve text classification accuracy, and can be well applied in the control field.
[0004] The technical solution adopted by the present invention to solve its technical problem is:
[0005] A Chinese text classification method based on deep learning includes the following steps:
[0006] 1) First, preprocess the input text. The process is as follows: remove stop words and special symbols; use Jieba Chinese word segmentation tool to perform word segmentation and obtain w1, w2, w3, ···, w n Then use the Word2Vec Chinese pre-training model to output the word vector corresponding to each word, using c1, c2, c3, ···, c n Indicates that the word vector group c1,c2,c3,···,c n Denoted as matrix W C ; Use the trained LDA model to output the topic-word matrix of the text, denoted as W T , and finally the matrix W C and the matrix WT Concatenate the corresponding vectors to get n×d model Dimensional matrix W E , d model is the word vector dimension, satisfying
[0007] W E =[W C ,W T ] (1)
[0008] 2) The matrix W E Input to the encoder, first perform position encoding to obtain the matrix P. The calculation process is as follows:
[0009]
[0010] 3) Substitute matrix P with matrix W E Add together to get the final n×d model dimensional word embedding matrix W I , the formula is as follows:
[0011] W I =W E +P (3)
[0012] 4) Input the matrix into W I To the self-attention mechanism, generate n×d model The dimension matrix M is given by the following formula:
[0013]
[0014] Q=W I ×W Q (5)
[0015] K=W I ×W K (6)
[0016] V=W I ×W V (7)
[0017] Among them, W Q , W K , W V For trainable model dimensional parameter matrix, d k is an adjustable hyperparameter;
[0018] 5) Substitute matrices M and W I Perform residual and normalization operations to obtain n×d model The dimension matrix N1 is as follows:
[0019] N1=LayerNorm(Μ+W I ) (8)
[0020] 6) Input the matrix N1 into the feedforward neural network to obtain n×d model dimensional matrix F, the formula is as follows:
[0021] F=max(0,N2W1+B1)W2+B2 (9)
[0022] Where W1 and W2 are neural network weight matrices, and B1 and B2 are neural network bias items;
[0023] 7) Next, transform the matrix W I Input into a single hidden layer recurrent neural network, save the hidden layer output vector at each moment, recorded as matrix R1, d r is the RNN network dimension;
[0024] 8) Linearly transform the matrix R1 into n×d model Dimensional Matrix The formula is as follows
[0025]
[0026] Among them, W L is d r ×d model dimension trainable parameter matrix;
[0027] 9) The matrix F, and N1 are subjected to residual and normalization operations to obtain the matrix N2, the formula is as follows
[0028]
[0029] 10) Take the first vector of matrix N2 and input it into the classifier. It first passes through the feedforward neural network and outputs d f dimensional vector f, the formula is as follows:
[0030] f=(v CLS ·w1+b1)w2+b2 (12)
[0031] Among them, v CLS is the first vector of N2, w1 and w2 are the neural network weights, b1 and b2 are the neural network bias terms, and d f is the dimension of the neural network;
[0032] 11) Perform Softmax operation on the elements of vector f, and the dimension with the largest value corresponds to the text category y p , the formula is as follows:
[0033] y p =softmax(f) (13)
[0034] 12) The model parameters are trained through the cross entropy loss function. The model parameters include matrix elements, neural network weights and bias terms. The loss function is as follows:
[0035]
[0036] Among them, M is the total number of training samples, y t is the true category, y p is the predicted category.
[0037] The beneficial effects of the present invention are: Word2Vec word vectors are fused with LDA word vectors as word embeddings, and the self-attention mechanism is combined with the recurrent neural network in the encoding stage to deeply extract text features, thereby improving the traditional machine learning model to achieve better text classification effects, thereby effectively improving people's efficiency in processing information. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is a schematic diagram of the Chinese text classification system model based on deep learning, which is mainly composed of three parts: embedding layer, encoding layer and classifier. The embedding layer includes the following parts: Word2Vec model and LDA model; the encoding layer mainly includes the following parts: Self-Attention mechanism, RNN network, and feed forward neural network (Feed forward network). DETAILED DESCRIPTION
[0039] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0040] Reference Figure 1 , a Chinese text classification method based on deep learning, the implementation of this method can maximize the accuracy of text classification, the present invention can be applied to the control field, such as Figure 1 As shown, the text classification method for this scenario includes the following steps:
[0041] 1) First, preprocess the input text. The process is as follows: remove stop words and special symbols; use Jieba Chinese word segmentation tool to perform word segmentation and obtain w1, w2, w3, ···, w n Then use the Word2Vec Chinese pre-training model to output the word vector corresponding to each word, using c1, c2, c3, ···, c n Indicates that the word vector group c1,c2,c3,···,c n Denoted as matrix W C ; Use the trained LDA model to output the topic-word matrix of the text, denoted as W T , and finally the matrix W C and the matrix W TConcatenate the corresponding vectors to get n×d model Dimensional matrix W E , d model is the word vector dimension, satisfying
[0042] W E =[W C ,W T ] (1)
[0043] 2) The matrix W E Input to the encoder, first perform position encoding to obtain the matrix P. The calculation process is as follows:
[0044]
[0045] 3) Substitute matrix P with matrix W E Add together to get the final n×d model dimensional word embedding matrix W I , the formula is as follows:
[0046] W I =W E +P (3)
[0047] 4) Input the matrix into W I To the self-attention mechanism, generate n×d model The dimension matrix M is given by the following formula:
[0048]
[0049] Q=W I ×W Q (5)
[0050] K=W I ×W K (6)
[0051] V=W I ×W V (7)
[0052] Among them, W Q , W K , W V For trainable model dimensional parameter matrix, d k is an adjustable hyperparameter;
[0053] 5) Substitute matrices M and W I Perform residual and normalization operations to obtain n×d model The dimension matrix N1 is as follows:
[0054] N1=LayerNorm(Μ+W I ) (8)
[0055] 6) Input the matrix N1 into the feedforward neural network to obtain n×d model dimensional matrix F, the formula is as follows:
[0056] F=max(0,N2W1+B1)W2+B2 (9)
[0057] Where W1 and W2 are neural network weight matrices, and B1 and B2 are neural network bias items;
[0058] 7) Next, transform the matrix W I Input into a single hidden layer recurrent neural network, save the hidden layer output vector at each moment, recorded as matrix R1, d r is the RNN network dimension;
[0059] 8) Linearly transform the matrix R1 into n×d model Dimensional Matrix The formula is as follows
[0060]
[0061] Among them, W L is d r ×d model dimension trainable parameter matrix;
[0062] 9) The matrix F, and N1 are subjected to residual and normalization operations to obtain the matrix N2, the formula is as follows
[0063]
[0064] 10) Take the first vector of matrix N2 and input it into the classifier. It first passes through the feedforward neural network and outputs d f dimensional vector f, the formula is as follows:
[0065] f=(v CLS ·w1+b1)w2+b2 (12)
[0066] Among them, v CLS is the first vector of N2, w1 and w2 are the neural network weights, b1 and b2 are the neural network bias terms, and d f is the dimension of the neural network;
[0067] 11) Perform Softmax operation on the elements of vector f, and the dimension with the largest value corresponds to the text category y p , the formula is as follows:
[0068] y p =softmax(f) (13)
[0069] 12) The model parameters are trained through the cross entropy loss function. The model parameters include matrix elements, neural network weights and bias terms. The loss function is as follows:
[0070]
[0071] Among them, M is the total number of training samples, y t is the true category, y p is the predicted category.
[0072] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept and are for illustrative purposes only. The protection scope of the present invention should not be considered to be limited to the specific forms described in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be thought of by ordinary technicians in this field based on the inventive concept.
Claims
1. A Chinese text classification method based on deep learning, characterized in that: The method comprises the following steps: 1) First, preprocess the input text. The process is as follows: remove stop words and special symbols; use Jieba Chinese word segmentation tool to perform word segmentation and obtain w1, w2, w3, ···, w n Then use the Word2Vec Chinese pre-training model to output the word vector corresponding to each word, using c1, c2, c3, ···, c n Indicates that the word vector group c1,c2,c3,···,c n Denoted as matrix W C ; Use the trained LDA model to output the topic-word matrix of the text, denoted as W T , and finally the matrix W C and the matrix W T Concatenate the corresponding vectors to get n×d model Dimensional matrix W E , d model is the word vector dimension, satisfying IN E =[W C ,IN T ] (1) 2) The matrix W E Input to the encoder, first perform position encoding to obtain the matrix P. The calculation process is as follows: 3) Substitute matrix P with matrix W E Add together to get the final n×d model dimensional word embedding matrix W I , the formula is as follows: W I =W E +P (3) 4) Input the matrix into W I To the self-attention mechanism, generate n×d model The dimension matrix M is given by the following formula: Q=W I ×W Q (5) K=W I ×W K (6) V=W I ×W V (7) Among them, W Q , W K , W V For trainable model dimensional parameter matrix, d k is an adjustable hyperparameter; 5) Substitute matrices M and W I Perform residual and normalization operations to obtain n×d model The dimension matrix N1 is as follows: N1=LayerNorm(Μ+W I ) (8) 6) Input the matrix N1 into the feedforward neural network to obtain n×d model dimensional matrix F, the formula is as follows: F=max(0, N1W1+B1)W2+B2 (9) Where W1 and W2 are neural network weight matrices, and B1 and B2 are neural network bias items; 7) Next, transform the matrix W I Input into a single hidden layer recurrent neural network, save the hidden layer output vector at each moment, recorded as matrix R1, d r is the RNN network dimension; 8) Linearly transform the matrix R1 into n×d model Dimensional Matrix The formula is as follows Among them, W L is d r ×d model dimension trainable parameter matrix; 9) The matrix F, and N1 are subjected to residual and normalization operations to obtain the matrix N2, the formula is as follows 10) Take the first vector of matrix N2 and input it into the classifier. It first passes through the feedforward neural network and outputs d f dimensional vector f, the formula is as follows: f=(v CLS ·w1+b1)w2+b2 (12) Among them, v CLS is the first vector of N2, w1 and w2 are the neural network weights, b1 and b2 are the neural network bias terms, and d f is the dimension of the neural network; 11) Perform Softmax operation on the elements of vector f, and the dimension with the largest value corresponds to the text category y p , the formula is as follows: y p =softmax(f) (13) 12) The model parameters are trained through the cross entropy loss function. The model parameters include matrix elements, neural network weights and bias terms. The loss function is as follows: Among them, S is the total number of training samples, y t is the true category, y p is the predicted category.