Text classification method based on fusion features and improved LSTM

By integrating the features extracted by Word2Vec and BERT in the text classification model and introducing attention mechanisms and residual connections in the LSTM network, the information capture and gradient vanishing problems of existing models when processing long texts are solved, achieving higher text classification accuracy and stability.

CN120030160AActive Publication Date: 2025-05-23NANJING UNIV OF POSTS & TELECOMM
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510144644.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-23
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

Existing text classification models are difficult to effectively capture long-range dependencies and key information when processing long texts, and feature extraction does not take into account both static features and context information, making it easy to encounter the problem of gradient disappearance or explosion.

Method used

Using text classification method based on fusion features and improved LSTM, static features and context features are extracted through Word2Vec and BERT models, and fusion is carried out, combining attention mechanisms and residual connections, the LSTM network is optimized to improve information mobility.

Benefits of technology

It improves the accuracy of text classification, can more effectively capture the basic semantic information of vocabulary and dynamic semantic changes in the context, alleviates the problem of gradient vanishing, and improves the stability and accuracy of the model in long sequence processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030160A_ABST
    Figure CN120030160A_ABST
Patent Text Reader

Abstract

The invention discloses a text classification method based on fusion features and an improved LSTM, and the method comprises the steps: obtaining text data, and dividing the text data into a training set and a test set; preprocessing the text to obtain cleaned text data; extracting features of the text by using a Word2Vec method to obtain a static feature vector; extracting features of the text by using a pre-trained BERT Chinese model to obtain feature vectors containing contexts; fusing the static feature vector and the feature vector containing the context to obtain a fused feature; inputting the fusion features of the training set into the improved LSTM network for model training; performing classification verification on the test set by using the trained classification model so as to evaluate the efficiency of the model; according to the method, static and dynamic feature vectors are combined, the advantages of the static and dynamic feature vectors are utilized, weight distribution of input features is optimized through an attention mechanism, attention of a model to key information is enhanced, and the method is suitable for various fields needing high-precision text classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a text classification method based on fusion features and improved LSTM. Background Art

[0002] With the rapid development of the Internet and information technology, massive amounts of text data such as news reports, social media posts, product reviews, and legal documents are constantly being generated. Text classification, as a basic task in natural language processing, is widely used in various information processing systems, such as news classification, sentiment analysis, spam detection, and public opinion monitoring.

[0003] In recent years, the technology in the field of text classification has developed rapidly, especially the introduction of deep learning models, which has greatly improved the accuracy and efficiency of text classification. Models based on various neural networks can effectively process the sequence information of text and capture the contextual relationships in the text. At the same time, the emergence of some pre-training models using massive unsupervised data has further improved the performance of text classification models.

[0004] However, existing models still face the challenge of not being able to effectively capture long-range dependencies and key information when processing long texts. There are cases where the extraction of text features does not pay attention to semantics or loses context. Most models used for feature extraction are single and cannot take into account both static and dynamic features at the same time. In traditional neural networks, the model treats all input information equally, while in text classification, different words or phrases have different effects on the final classification. At the same time, neural networks with increasing depth will face the problem of gradient vanishing or exploding, making the training process very difficult. The error information of deep networks is difficult to effectively propagate back to the previous layer, making it difficult for the model to capture complex patterns. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention discloses a text classification method based on fused features and improved LSTM, which takes into account both static features and contextual content in feature extraction, and introduces an attention mechanism into the existing LSTM network to increase the weight of key information in the calculation process, so that the network focuses on important information. At the same time, residual connections are used to ensure the fluidity of information and avoid model degradation problems.

[0006] To achieve the above object, the present invention provides the following technical solution: a text classification method based on fusion features and improved LSTM, comprising the following steps:

[0007] S1. Obtain text data and divide it into training set and test set;

[0008] S2, preprocessing the text to obtain cleaned text data;

[0009] S3, use the Word2Vec method to extract the features of the text and obtain a static feature vector;

[0010] S4. Use the pre-trained BERT Chinese model to extract the features of the text and obtain a feature vector containing the context;

[0011] S5, fusing the static feature vector and the feature vector containing the context to obtain a fused feature;

[0012] S6. Input the fusion features of the training set into the improved LSTM network for model training;

[0013] S7. Use the trained classification model to perform classification verification on the test set to evaluate the effectiveness of the model.

[0014] Preferably, in step S2, the text data preprocessing step includes:

[0015] Remove HTML tags, special characters and extra spaces, convert all letters in the text to lowercase, and remove stop words.

[0016] Preferably, in step S3, the Word2Vec feature vector is obtained by training a Word2Vec model, and the Word2Vec model is trained using text data, and the specific steps include:

[0017] The cleaned text data is used to train the Word2Vec model to generate word embedding vectors. The Word2Vec model uses the CBOW method to train each word according to the context window to obtain a 768-dimensional vector representation of each word. The steps include:

[0018] Using the vocabulary in the email text, we construct a context window, select context words within a certain range as input, and target words as output;

[0019] The context word vectors are averaged to obtain the context feature vector h;

[0020] Use the CBOW model to predict the target word through the context vector h, and use the Softmax function to probabilistically process the score of each word. The specific formula of the Softmax function is as follows:

[0021]

[0022] Among them, v ωt Represents the target word ω t The word vector of h is the average vector of the context words, V is the vocabulary, and P(ω t |ω t-n ,…,ωt-n ) is the target word ω under the given context t The predicted probability of .

[0023] Preferably, in step S4, when using the pre-trained BERT Chinese model to extract text features, the specific steps include:

[0024] Use the pre-trained BERT tokenizer to tokenize the email text to obtain a word sequence, which is then passed as input to the BERT model.

[0025] Encode the word sequence using a pre-trained BERT Chinese model to obtain a context-dependent representation of each word, wherein the context-dependent representation is a dynamic word vector, wherein the BERT model adjusts the semantic representation of each word according to the occurrence of the word in different contexts;

[0026] The feature vector of each word is extracted from the encoding result output by the BERT model. The feature vector is marked by the CLS tag of the last layer to obtain the global feature representation of the text, which is a 768-dimensional vector. The global feature representation contains context information.

[0027] Preferably, in step S5, the static feature vector extracted by Word2Vec and the context feature vector extracted by BERT are fused, and the fusion step includes:

[0028] Feature concatenation: The feature vectors extracted by the Word2Vec and BERT models are directly concatenated to form a 1536-dimensional feature representation;

[0029] Adjustment of fusion feature dimension: To avoid the computational overhead caused by too high a dimension, the PCA principal component analysis method is further used to reduce the dimension of the fused feature vector to obtain a reduced dimension feature vector with a dimension of 768. PCA includes the following steps:

[0030] 1) Standardize the fused feature vector;

[0031] 2) Calculate the covariance matrix of the standardized feature matrix;

[0032] 3) Perform eigenvalue decomposition on the covariance matrix and select the first 768 principal components;

[0033] 4) Project the standardized feature vector onto the 768 principal components to obtain the fused feature vector after dimension reduction.

[0034] Preferably, in step S6, the improved LSTM network includes both an attention mechanism and a residual connection, wherein the attention mechanism dynamically calculates the attention weight of each input according to the current input information and the hidden state information at the previous moment, and optimizes the input of the LSTM by weighted fusion of the input information; the residual connection is introduced in each layer of the LSTM, and the input information at the current moment is directly added to the output after the LSTM transformation, thereby ensuring that the information flow is better transmitted in the network and effectively alleviating the gradient vanishing problem.

[0035] Preferably, the attention mechanism calculates the current input vector x t , the hidden state h at the previous moment t-1 and cell state c t-1 To generate the attention weight a t , the current input x t And the calculated attention weight a t Multiply them together to get the optimized input vector x′ t Input into the LSTM unit, the attention weight formula is:

[0036] a t =σ a (W a x t +U a h t-1 +M a c t-1 +b a )

[0037] Weighted input formula:

[0038] x′ t =a t ·x t

[0039] where σ a is the Sigmod activation function, W a , U a 、M a is the weight matrix of the attention mechanism, b a is the bias, a t is the attention weight at the current moment.

[0040] Preferably, the residual connection of the LSTM unit is implemented in the following manner:

[0041] For the input vector x' at time t t , the output after LSTM transformation is F(x' t ), the residual connection is connected by taking the input x' t With the output F(x' t ) are added to get the final output yt :

[0042] y t =F(x' t )+x' t

[0043] F(x' t )=h t

[0044] Output Gate:

[0045] o t =σ(W O x' t +U O h t-1 +b 0 )

[0046] h t =o t ·tanh(c t )

[0047] Cell Status:

[0048]

[0049] Candidate cell states:

[0050]

[0051] Input Gate:

[0052] i t =σ(W i x' t +U i h t-1 +b i )

[0053] Forget Gate:

[0054] f t =σ(W f x' t +U f h t-1 +b f )

[0055] Among them, h t is the current hidden state, c t is the current cell state, W O , W c , W i , W f is the weight matrix of different gates, U O , U c , U i , Uf is the hidden state h at the previous moment t-1 Impact on different gates, b 0 , b c , b i , b f For bias.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] 1. The present invention improves the accuracy of text classification. By integrating the static features extracted by Word2Vec and the contextual features extracted by BERT, it can simultaneously consider the basic semantic information of vocabulary and the dynamic semantic changes in the context, thereby providing a more comprehensive and accurate feature representation for the text classification model.

[0058] 2. The present invention optimizes the performance of LSTM networks. By introducing the attention mechanism and residual connection in the LSTM network, the present invention effectively alleviates the gradient vanishing problem that the LSTM model may encounter when processing long texts. The attention mechanism enables the model to adaptively focus on the important parts of the text, while the residual connection ensures the effective transmission of information between each layer, thereby improving the stability and accuracy of the model in long sequence processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0060] In the attached picture:

[0061] Figure 1 It is a flowchart of a text classification method based on fusion features and improved LSTM provided by an embodiment of the present invention;

[0062] Figure 2 is a schematic diagram of an LSTM unit of an embodiment of the present invention;

[0063] Figure 3 This is a verification accuracy comparison test chart of the improved LSTM network of the present invention and LSTM, CNN, and RNN. DETAILED DESCRIPTION

[0064] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0065] Example: Figure 1-Figure 3 As shown, a text classification method based on fusion features and improved LSTM includes the following steps:

[0066] Step 1: Get the text data and divide it into training set and test set: divide 75% of the text data into the training set and the remaining 25% into the test set.

[0067] Step 2: Preprocess the text to obtain cleaned text data:

[0068] (1) Remove HTML tags and special characters: When processing input text, first remove all HTML tags in the text (such as , , etc.), as well as possible special characters (such as ,, &, <, >, etc.). These HTML tags and special characters are irrelevant to the actual semantics of the text and may interfere with subsequent text analysis, so they need to be removed.

[0069] (2) Remove extra spaces: There may be extra spaces, tab characters, line break characters and other unnecessary characters in the text. These characters not only waste computing resources but may also interfere with subsequent text tokenization and feature extraction. Therefore, remove extra spaces to ensure the standardization and consistency of the text.

[0070] (3) Convert all letters to lowercase: To avoid inconsistencies in vocabulary processing caused by case differences, convert all letters in the text to lowercase. This step helps to simplify text feature extraction and improve model processing efficiency, avoiding redundant features due to different letter cases.

[0071] (4) Remove stop words: Stop words (such as "de", "shi", "zai", "he", etc.) usually do not contribute significantly to the understanding of the text in natural language processing but appear frequently in the text. By removing these stop words, the dimension of the text can be reduced, the amount of calculation can be decreased, and at the same time, the subsequent text representation can be more focused on meaningful content.

[0072] Step 3: Use the Word2Vec method to extract the features of the text and obtain static feature vectors;

[0073] Use the Word2Vec method based on the gensim library to process all the text data in the training set. Set the dimension of the word vector to 768, the size of the context window to 5, the minimum limit of word frequency to 1, and use the CBOW method for training.

[0074] According to the words in the text, construct a context window, and select the context words within a certain range in the window as input and the target word as output for model training.

[0075] Average the word vectors corresponding to the context words to obtain the feature vector h of the context as the context representation of the target word;

[0076] Predict the target word through the context vector h and probabilize the scores of each word through the Softmax function. Among them, the specific formula of the Softmax function is as follows:

[0077]

[0078] Among them, v ωt represents the word vector of the target word ω t h is the average vector of the context words, V is the vocabulary, P(ω t |ω t-n ,…,ω t-n ) is the target word ω under the given context t The predicted probability of

[0079] Step 4: Use the pre-trained BERT Chinese model to extract the features of the text and obtain a feature vector containing the context;

[0080] In this implementation, the BERT model used is derived from the paper "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding". In order to adapt to the processing of Chinese text, the Chinese version of the paper, bert-base-chinese, is used. This model has been specially trained and optimized for Chinese text processing and can effectively capture the contextual information in Chinese text.

[0081] (1) Use the pre-trained BERT word segmenter to segment the email text to obtain a word sequence, which is passed as input to the BERT model;

[0082] (2) using a pre-trained BERT Chinese model to encode the word sequence to obtain a context-dependent representation of each word, wherein the context-dependent representation is a dynamic word vector, wherein the BERT model can adjust the semantic representation of each word according to the occurrence of the word in different contexts;

[0083] (3) Extracting the feature vector of each word from the encoding result output by the BERT model, the feature vector is passed through the CLS tag of the last layer (which is a 768-dimensional vector) to obtain the global feature representation of the text, and the global feature representation contains context information;

[0084] Step 5: Fuse the static feature vector and the feature vector containing the context to obtain a fused feature;

[0085] (1) Feature concatenation: The feature vectors extracted by the Word2Vec and BERT models are directly concatenated to form a 1536-dimensional feature representation, as follows:

[0086] f combined =[f bert ,f w2v ]

[0087] where f bert The 768-dimensional vector generated by BERT, f w2v The 768-dimensional vector generated by Word2Vec, f combined It is the concatenated 1536-dimensional vector.

[0088] (2) Adjustment of fusion feature dimension: In order to avoid the computational overhead caused by high-dimensional features, PCA (principal component analysis) is further used to reduce the dimension of the concatenated fusion feature vector. PCA can project high-dimensional feature vectors into a low-dimensional space and retain the main information of the data as much as possible. After PCA processing, the dimension of the fusion feature vector will be reduced to 768 to meet the needs of subsequent model training and calculation.

[0089] The fused feature vector is standardized so that the feature of each dimension has a standard normal distribution with a mean of 0 and a variance of 1. The standardization formula is:

[0090]

[0091] Where f is the original feature vector, μ is the mean of the feature, and σ is the standard deviation of the feature.

[0092] After the normalization, the feature matrix X norm Calculate the covariance matrix. The covariance matrix reflects the linear relationship between features of different dimensions. The formula is as follows:

[0093]

[0094] Where C is the covariance matrix, n is the number of samples, X norm is the standardized feature matrix.

[0095] Perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvalues ​​and corresponding eigenvectors. Select the first 768 principal components as the new projection space. The formula for eigenvalue decomposition is:

[0096]

[0097] v i is the i-th eigenvector, λ i is the corresponding eigenvalue.

[0098] The standardized fused feature vector is projected onto the first 768 principal components selected to obtain the fused feature vector after dimension reduction. The projection formula is:

[0099] f pca =X norm V 768

[0100] Where V 768 is the eigenvector matrix of the first 768 principal components, f pca is the fused feature vector after dimensionality reduction.

[0101] Finally, the fused feature vector Fpca obtained by PCA dimensionality reduction has 768 dimensions, which effectively reduces the computational complexity while retaining the main semantic information.

[0102] Calculate the covariance matrix of the standardized feature matrix;

[0103] Perform eigenvalue decomposition on the covariance matrix and select the first 768 principal components;

[0104] The standardized feature vector is projected onto the 768 principal components to obtain the fused feature vector after dimension reduction.

[0105] Step 6: Input the fusion features of the training set into the improved LSTM network for model training;

[0106] The improved LSTM network introduces attention mechanism and residual connection to ensure that important information is not lost;

[0107] Depend on Figure 1 It can be seen that the attention mechanism calculates the current input vector x t , the hidden state h at the previous moment t-1 and cell state c t-1 To generate the attention weight a t , the current input x t And the calculated attention weight a t Multiply to get the optimized input vector x' t Input into the LSTM unit, the attention weight formula is:

[0108] a t =σ a (W a x t +U a h t-1 +M a c t-1 +b a )

[0109] Weighted input formula:

[0110] x' t =a t ·x t

[0111] where σ a is the Sigmod activation function, W a , U a 、M a is the weight matrix of the attention mechanism, b a is the bias, a t is the attention weight at the current moment.

[0112] The residual connection of the LSTM unit is realized in the following way: for the input vector x′ at the tth moment t , the output after LSTM transformation is F(x′ t ), the residual connection is connected by taking the input x' t And the output F(x' t ) are added to get the final output y t :

[0113] y t =F(x′ t )+x' t

[0114] F(x′ t )=h t

[0115] Output Gate:

[0116] o t =σ(W O x' t +U O h t-1 +b 0 )

[0117] h t =o t ·tanh(c t )

[0118] Cell Status:

[0119]

[0120] Candidate cell states:

[0121]

[0122] Input Gate:

[0123] i t =σ(W i x' t +U i h t-1 +b i )

[0124] Forget Gate:

[0125] f t =σ(W f x' t +U f h t-1 +b f )

[0126] Among them, h t is the current hidden state, c t is the current cell state, W O , W c , W i , W f is the weight matrix of different gates, U O , U c , U i , U f is the hidden state h at the previous moment t-1 Impact on different gates, b 0 , b c , b i , b f For bias.

[0127] By introducing the attention mechanism and residual connection in the LSTM network, the gradient vanishing problem that the LSTM model may encounter when processing long texts is effectively alleviated. The attention mechanism enables the model to adaptively focus on the important parts of the text, while the residual connection ensures the effective transfer of information between each layer, thereby improving the stability and accuracy of the model in long sequence processing.

[0128] Step 7: Use the trained classification model to verify the classification of the test set.

[0129] To verify the effectiveness of the improved LSTM, we compared it with other neural network models CNN,

[0130] LSTM and RNN are compared and trained for 10, 30, and 50 epochs respectively. The model accuracy is compared. The experimental results are shown in Table 1.

[0131]

[0132]

[0133] As shown in Table 1, when different neural network models are used for experiments, the accuracy of the improved LSTM is the best. When trained 10 times, the accuracy is improved by 2.01% relative to LSTM, when trained 30 times, the accuracy is improved by 1.16% relative to LSTM, and when trained 50 times, the accuracy is improved by 1.08% relative to LSTM.

[0134] Under different training times, the accuracy of RNN and LSTM varies greatly, while the model training effect of improved LSTM is similar. The accuracy difference between training 10 times and training 50 times is only 0.13%, which shows that the improved LSTM has good generalization ability while having higher accuracy. The model is not sensitive to changes in training rounds and can stably adapt to different data and tasks.

[0135] Finally, it should be noted that the above description is only a preferred example of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A text classification method based on fusion features and improved LSTM, characterized in that: The following steps are involved: S1. Obtain text data and divide it into training set and test set; S2, preprocessing the text to obtain cleaned text data; S3, use the Word2Vec method to extract the features of the text and obtain a static feature vector; S4. Use the pre-trained BERT Chinese model to extract the features of the text and obtain a feature vector containing the context; S5, fusing the static feature vector and the feature vector containing the context to obtain a fused feature; S6, input the fusion features of the training set into the improved LSTM network for model training; S7. Use the trained classification model to perform classification verification on the test set to evaluate the effectiveness of the model.

2. A text classification method based on fusion features and improved LSTM according to claim 1, characterized in that: In step S2, the preprocessing steps of the text data include: Remove HTML tags, special characters and extra spaces, convert all letters in the text to lowercase, and remove stop words.

3. A text classification method based on fusion features and improved LSTM according to claim 1, characterized in that: In step S3, the Word2Vec feature vector is obtained by training the Word2Vec model. The Word2Vec model is trained using text data. The specific steps include: The cleaned text data is used to train the Word2Vec model to generate word embedding vectors. The Word2Vec model uses the CBOW method to train each word according to the context window to obtain a 768-dimensional vector representation of each word. The steps include: Using the vocabulary in the email text, we construct a context window, select context words within a certain range as input, and target words as output; The context word vectors are averaged to obtain the context feature vector h; Use the CBOW model to predict the target word through the context vector h, and use the Softmax function to probabilistically process the score of each word. The specific formula of the Softmax function is as follows: in, Represents the target word ω t The word vector of h is the average vector of the context words, V is the vocabulary, and P(ω t |ω t-n ,…,ω t-n ) is the target word ω under the given context t The predicted probability of .

4. A text classification method based on fusion features and improved LSTM according to claim 1, characterized in that: In step S4, when using the pre-trained BERT Chinese model to extract text features, the specific steps include: Use the pre-trained BERT tokenizer to segment the email text to obtain a word sequence, which is passed as input to the BERT model. Encode the word sequence using a pre-trained BERT Chinese model to obtain a context-dependent representation of each word, wherein the context-dependent representation is a dynamic word vector, wherein the BERT model adjusts the semantic representation of each word according to the occurrence of the word in different contexts; The feature vector of each word is extracted from the encoding result output by the BERT model. The feature vector is marked by the CLS tag of the last layer to obtain the global feature representation of the text, which is a 768-dimensional vector. The global feature representation contains context information.

5. A text classification method based on fusion features and improved LSTM according to claim 1, characterized in that: In step S5, the static feature vector extracted by Word2Vec and the context feature vector extracted by BERT are fused. The fusion step includes: Feature concatenation: The feature vectors extracted by the Word2Vec and BERT models are directly concatenated to form a 1536-dimensional feature representation; Adjustment of fusion feature dimension: To avoid the computational overhead caused by too high a dimension, the PCA principal component analysis method is further used to reduce the dimension of the fused feature vector to obtain a reduced dimension feature vector with a dimension of 768. PCA includes the following steps: 1) Standardize the fused feature vector; 2) Calculate the covariance matrix of the standardized feature matrix; 3) Perform eigenvalue decomposition on the covariance matrix and select the first 768 principal components; 4) Project the standardized feature vector onto the 768 principal components to obtain the fused feature vector after dimension reduction.

6. A text classification method based on fusion features and improved LSTM according to claim 1, characterized in that: In step S6, the improved LSTM network includes both an attention mechanism and a residual connection. The attention mechanism dynamically calculates the attention weight of each input according to the current input information and the hidden state information at the previous moment, and optimizes the input of the LSTM by weighted fusion of the input information. The residual connection is introduced in each layer of the LSTM to directly add the input information at the current moment to the output after the LSTM transformation.

7. A text classification method based on fusion features and improved LSTM according to claim 6, characterized in that: The attention mechanism calculates the current input vector x t , the hidden state h at the previous moment t-1 and cell state c t-1 To generate the attention weight a t , the current input x t And the calculated attention weight a t Multiply to get the optimized input vector x' t Input into the LSTM unit, the attention weight formula is: a t =σ a (W a x t +U a h t-1 +M a c t-1 +b a ) Weighted input formula: x’ t =a t ·x t where σ a is the Sigmod activation function, W a , U a 、M a is the weight matrix of the attention mechanism, b a is the bias, a t is the attention weight at the current moment.

8. A text classification method based on fusion features and improved LSTM according to claim 7, characterized in that: The residual connection of the LSTM unit is implemented in the following way: For the input vector x' at time t t , the output after LSTM transformation is F(x' t ), the residual connection is connected by taking the input x' t With the output F(x' t ) are added to get the final output y t : y t =F(x’ t )+x’ t F(x’ t )=h t Output Gate: o t =σ(W O x’ t +U O h t-1 +b0) h t =o t ·tanh(c t ) Cell Status: Candidate cell states: Input Gate: i t =σ(W i x’ t +U i h t-1 +b i ) Forget Gate: f t =σ(W f x’ t +U f h t-1 +b f ) Among them, h t is the current hidden state, c t is the current cell state, W O , W c , W i , W f is the weight matrix of different gates, U O , U c , U i , U f is the hidden state h at the previous moment t-1 Impact on different gates, b0, b c 、b i 、b f For bias.

Citation Information

Patent Citations

  • A multi-feature fusion Chinese news text abstract generation method based on a neural network

    CN109344391A

  • Chinese sentiment analysis method based on BERT, LSTM and CNN fusion

    CN110334210A

  • BERT-A-BiLSTM-based multi-feature patent automatic classification algorithm

    CN113011527A

  • Text classification method and device and storage medium

    CN114860930A

  • Text classification method based on bidirectional long-short term memory network fused attention mechanism

    CN116383384A