Text emotion recognition method and recognition system based on BiLSTM and Transform
By using BiLSTM and Transformer models in parallel in the text emotion recognition system, the timing information and global dependency information of the text are extracted and feature fusion is performed, the problem of low accuracy of long text emotion recognition is solved, and more efficient and accurate emotion recognition effect is achieved.
Patent Information
- Application Number
- CN202411764911.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-05-06
AI Technical Summary
In long text recognition application scenarios, the prior art is difficult to effectively improve the accuracy of text emotion recognition, especially when the same word has different emotional tendencies in different contexts, imbalance in emotional data, and limited ability to capture long text emotional information.
Using the text emotion recognition method based on BiLSTM and Transformer, the text data is converted into word embedding representation through the preprocessing module, and sent to the first recognition module (BiLSTM model) and the second recognition module (Transformer model) running in parallel, and timing information and global dependency information are extracted respectively, and finally feature fusion and emotional recognition are performed in the comprehensive recognition module.
By combining the advantages of BiLSTM and Transformer models, fully integrating timing information and global dependency information, the accuracy of emotion recognition results is improved, the time overhead of traditional serial feature extraction methods is reduced, and the overall processing efficiency of the system is improved.
Smart Images

Figure CN119938909A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of natural language processing, and in particular to a text emotion recognition method and recognition system based on BiLSTM and Transformer. Background Art
[0002] With the rapid development of artificial intelligence interaction and other fields, text emotion recognition has received widespread attention. As a key task in natural language processing, text emotion recognition aims to enable machines to understand and recognize emotional information in text, providing important support for achieving natural human-computer interaction.
[0003] However, this field still faces multiple challenges. For example, the same word may have different emotional tendencies in different contexts, which increases the difficulty of recognition and leads to a decrease in the accuracy of emotion recognition results. Therefore, in the application scenario of long text recognition, how to improve the accuracy of emotion recognition has become a problem that needs to be solved. Summary of the invention
[0004] In view of this, the present disclosure provides a text emotion recognition method and recognition system based on BiLSTM and Transformer to solve the problem of how to improve the accuracy of emotion recognition in the application scenario of long text recognition.
[0005] On the one hand, the present disclosure provides a text emotion recognition method based on BiLSTM and Transformer, which is applied to a recognition system. The recognition system includes: a preprocessing module, a first recognition module, a second recognition module and a comprehensive recognition module. The method includes: the preprocessing module obtains preset text data, preprocesses the preset text data, converts the preset text data into a word embedding representation, and sends the word embedding representation to the first recognition module and the second recognition module at the same time; the first recognition module adopts a BiLSTM model to determine the time series information corresponding to the word embedding representation; the second recognition module adopts a Transformer model to determine the global dependency information corresponding to the word embedding representation; the first recognition module and the second recognition module run in parallel; the comprehensive recognition module performs feature fusion on the time series information and the global dependency information, adopts a convolutional neural network to perform emotion recognition on the result of the feature fusion, and obtains the emotion recognition result of the preset text data.
[0006] On the other hand, the present disclosure also provides a recognition system, which includes a preprocessing module, a first recognition module, a second recognition module and a comprehensive recognition module, wherein: the preprocessing module is used to obtain preset text data, preprocess the preset text data, convert the preset text data into a word embedding representation, and send the word embedding representation to the first recognition module and the second recognition module at the same time; the first recognition module is used to use a BiLSTM model to determine the time series information corresponding to the word embedding representation; the second recognition module is used to use a Transformer model to determine the global dependency information corresponding to the word embedding representation; the first recognition module and the second recognition module run in parallel; the comprehensive recognition module is used to perform feature fusion on the time series information and the global dependency information, use a convolutional neural network to perform emotion recognition on the result of the feature fusion, and obtain the emotion recognition result of the preset text data.
[0007] On the other hand, the present disclosure further provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to enable a computer to implement the above-mentioned text emotion recognition method based on BiLSTM and Transformer.
[0008] On the other hand, the present disclosure further provides a computer program product, including computer instructions, which are used to enable a computer to execute the above-mentioned text emotion recognition method based on BiLSTM and Transformer.
[0009] Through the BiLSTM and Transformer-based text emotion recognition method and recognition system of the above-mentioned embodiments of the present invention, the word embedding representation is sent to the first recognition module and the second recognition module running in parallel, the timing information is obtained based on the first recognition module, the global dependency information is obtained based on the second recognition module, and the text emotion recognition is performed after the timing information and the global dependency information are integrated, which fully integrates the advantages of the two models, avoids the limitations of a single model in understanding emotion information, and thus improves the accuracy of the emotion recognition results.
[0010] In addition, the first recognition module and the second recognition module run in parallel, which reduces the time overhead in the traditional serial feature extraction method, improves the overall processing efficiency of the system, and meets the real-time requirements in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the drawings required for use in the specific embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0012] Figure 1a An exemplary schematic diagram showing the architecture of a recognition system applied to the text emotion recognition method based on BiLSTM and Transformer according to an embodiment of the present disclosure;
[0013] Figure 1b It is a flowchart of a text emotion recognition method based on BiLSTM and Transformer provided in an embodiment of the present disclosure;
[0014] Figure 2 An exemplary schematic diagram showing the architecture of another recognition system applied to the text emotion recognition method based on BiLSTM and Transformer in an embodiment of the present disclosure;
[0015] Figure 3 It is a flowchart of another text emotion recognition method based on BiLSTM and Transformer provided in an embodiment of the present disclosure;
[0016] Figure 4 This is a structural diagram of another identification system provided by the embodiment of the present disclosure.
[0017] Figure 5 It is a structural diagram of another identification system provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0018] With the rapid development of artificial intelligence interaction and other fields, text emotion recognition has received widespread attention. As a key task in natural language processing, text emotion recognition aims to enable machines to understand and recognize emotional information in text, providing important support for achieving natural human-computer interaction.
[0019] At present, although text emotion recognition has made significant progress, there are still some problems:
[0020] 1. When the text sentiment data is too concentrated, that is, when the number of samples of different sentiment categories is unbalanced, the model may be overly biased towards the majority sentiment category and ignore other categories, resulting in reduced accuracy of sentiment recognition results.
[0021] 2. The same word in text sentiment data may have different sentiment tendencies in different contexts. The current model has poor recognition effect in this aspect, which further reduces the accuracy.
[0022] 3. Some emotional expressions may be contained in longer text fragments, and the current model has limited ability to effectively capture key emotional information in long texts.
[0023] To solve the above problems, various embodiments of the present invention provide a text emotion recognition method based on BiLSTM and Transformer, which is applied to a recognition system. The recognition system includes: a preprocessing module, a first recognition module, a second recognition module and a comprehensive recognition module. The method includes: a preprocessing module, which obtains preset text data, preprocesses the preset text data, converts the preset text data into a word embedding representation, and sends the word embedding representation to the first recognition module and the second recognition module at the same time; the first recognition module, which adopts a BiLSTM model to determine the time series information corresponding to the word embedding representation; the second recognition module, which adopts a Transformer model to determine the global dependency information corresponding to the word embedding representation; the first recognition module and the second recognition module run in parallel; the comprehensive recognition module, which performs feature fusion on the time series information and the global dependency information, uses a convolutional neural network to perform emotion recognition on the result of the feature fusion, and obtains the emotion recognition result of the preset text data.
[0024] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0025] Please refer to Figure 1a , Figure 1a An exemplary schematic diagram of the architecture of a recognition system used in the text emotion recognition method based on BiLSTM and Transformer in the embodiment of the present disclosure is shown. Figure 1a As shown, the recognition system may include: a pre-processing module, a first recognition module, a second recognition module and a comprehensive recognition module.
[0026] The recognition system may refer to an intelligent system that automatically recognizes and analyzes the input text data, and can extract emotion-related information from the text and output the final emotion classification result.
[0027] Among them, the emotion classification results may include: happiness, anger, sadness, fear, etc.
[0028] The preprocessing module may refer to the input processing unit of the system, which is responsible for preliminary processing and standardization of the original text data to ensure the efficiency and accuracy of the subsequent recognition process.
[0029] The comprehensive recognition module can be the final processing unit of the system, which classifies the sentiment tendency of the text by fusing the output features of the first recognition module and the second recognition module.
[0030] like Figure 1a The preprocessing module, the first recognition module, the second recognition module and the comprehensive recognition module are deployed in the recognition system. The comprehensive recognition module interacts with the first recognition module and the second recognition module to obtain timing information and global dependency information and perform feature fusion and emotion recognition.
[0031] Further references Figure 1b , Figure 1b is a flowchart of a text emotion recognition method based on BiLSTM and Transformer provided by the embodiment of the present disclosure, which is applied to the above Figure 1a The identification system shown in the figure, the process of the method may include the following steps:
[0032] Step S101, a preprocessing module obtains preset text data, preprocesses the preset text data, converts the preset text data into word embedding representation, and sends the word embedding representation to the first recognition module and the second recognition module at the same time.
[0033] In this embodiment, the preset text data may refer to text data obtained from at least one data source, and the preset text data includes a preset emotion tag.
[0034] Here, in order to ensure that the data set includes various emotion types and language styles and avoids emotion type bias, the data source may include but is not limited to: social media platforms, news websites, blog forums, e-commerce platforms, and movie review websites. The preset emotion labels contained in the preset text data may include positive, negative, and neutral.
[0035] Preset text data containing positive preset emotion labels can be used to represent positive emotions, preset text data containing negative preset emotion labels can be used to represent negative emotions, and preset text data containing neutral preset emotion labels can be used to represent neutral emotions. For example, positive emotions can mean "this movie is really great", negative emotions can mean "this movie is really disappointing", and neutral emotions can mean "the weather is OK today".
[0036] Furthermore, the number of preset text data containing each preset emotion tag can be made equal and cover at least one language style.
[0037] Word embedding representation can refer to a dense vector representation method used to represent words in natural language processing, which can convert high-dimensional sparse representation into low-dimensional dense vectors to capture the semantic and contextual relationships between words.
[0038] Step S102: The first recognition module uses the BiLSTM model to determine the time sequence information corresponding to the word embedding representation.
[0039] In this embodiment, the bidirectional long short-term memory network (BiLSTM) can capture past and future contextual information in an input sequence by simultaneously processing information in the forward direction (from front to back) and the reverse direction (from back to front) of a sequence.
[0040] Here, the BiLSTM model can capture bidirectional context information and improve the semantic understanding ability of the model.
[0041] Temporal information may refer to the adjacency or dependency relationship implied between elements in sequence data when they are arranged in time or order, that is, the contextual dependency of words or phrases represented by word embedding in preset text data.
[0042] Step S103, the second recognition module uses the Transformer model to determine the global dependency information corresponding to the word embedding representation.
[0043] In this embodiment, the first recognition module and the second recognition module run in parallel.
[0044] The Transformer model can be a deep learning model based on the self-attention mechanism. The Transformer model can include at least one encoder and one decoder. Each encoder can include a multi-head self-attention mechanism, and the decoder can include a masked multi-head self-attention mechanism.
[0045] Here, the self-attention mechanism of the Transformer model can directly model the relationship between any two sequence positions, which is particularly suitable for processing long text sequences.
[0046] Global dependency information can refer to the association relationship between distant elements in sequence data. Through global dependency information, semantic or contextual dependencies between any positions in the sequence can be identified.
[0047] Step S104, the comprehensive recognition module performs feature fusion on the temporal information and the global dependency information, uses a convolutional neural network to perform emotion recognition on the result of the feature fusion, and obtains the emotion recognition result of the preset text data.
[0048] In this embodiment, a convolutional neural network (CNN) may refer to a deep learning model designed specifically for processing data with a grid structure (such as sequence data), and a convolution operation may be used to extract local features in the data.
[0049] The feature fusion of the time series information and the global dependency information may include: setting weights for the time series information and the global dependency information for comprehensive representation. The weights corresponding to the time series information and the global dependency information may be set according to the actual application and are not limited here.
[0050] As an example, the output format of the emotion recognition result of the preset text data may include but is not limited to: result label and result probability distribution.
[0051] Here, the result label may refer to whether the emotion recognition result is positive, negative, or neutral, and the result probability distribution may refer to (positive: 0.80, negative: 0.15, neutral: 0.05).
[0052] In the text emotion recognition method and recognition system based on BiLSTM and Transformer of the above-mentioned embodiment of the present disclosure, the complementary design of the two modules enables the convolutional neural network to simultaneously obtain the local temporal features and global semantic dependency features of the text, thereby more comprehensively characterizing the emotional expression. The comprehensive recognition module integrates the features extracted by BiLSTM and Transformer, combines the advantages of both, comprehensively analyzes the text emotion from multiple dimensions, and improves the recognition accuracy. Through the multi-dimensional feature learning ability of the model, even in the situation where the category samples are unevenly distributed, the model can still better identify the emotional features of a few categories and reduce the bias problem.
[0053] In a possible implementation of the above step S101, the preprocessing module obtains preset text data, preprocesses the preset text data, and converts the preset text data into a word embedding representation, including:
[0054] A preprocessing module uses a word segmentation tool to segment the preset text data into word sequences, and removes stop words in the word sequences based on a preset stop word list;
[0055] The preprocessing module converts the letters in the word sequence into uppercase and lowercase letters and cleans special symbols, and uses a pre-trained word embedding model to convert the word sequence into a word embedding representation.
[0056] In this embodiment, the format of the preset text data may include but is not limited to: a single character string, a text file, or a text data set.
[0057] A word segmentation tool may refer to a natural language processing technology for dividing an input text string into smallest language units (ie, words or subwords), which may be called "Tokens".
[0058] As an example, the word segmentation tool may include, but is not limited to: Natural Language Toolkit (NLTK), spaCy.
[0059] Furthermore, stop words may refer to words that appear frequently in a language but are of little significance to text analysis tasks. For example, stop words may include "the", "and", "is", "of" in English, or "的", "和" in Chinese. The preset stop word list can be provided by NLTK or customized according to specific tasks.
[0060] The purpose of removing stop words from the word sequence may be to remove words that are irrelevant to sentiment analysis while retaining potentially useful sentiment words. The new word sequence obtained after deleting stop words only contains meaningful vocabulary, thereby reducing computational overhead and data noise.
[0061] Furthermore, converting the case of letters in the word sequence may refer to converting all uppercase words in the new word sequence to lowercase, or converting all lowercase words to uppercase, so as to unify the vector representation and reduce the sparsity of the data.
[0062] Special symbols in the word sequence may include at least one of the following: “!@#$%^&*()[]{}”. Special symbols do not contribute to sentiment recognition in the sentiment recognition task and instead increase noise. Cleaning special symbols can reduce meaningless features, enabling the CNN to focus more on the actual words and their sentiment tendencies in the text, and optimizing the recognition efficiency.
[0063] The pre-trained word embedding model may refer to a model that learns word vector representations by training on large-scale unsupervised text data. By mapping words to low-dimensional continuous vectors, words with similar semantics are closer in the vector space.
[0064] As an example, the word embedding model may include, but is not limited to: Word2Vec, FastText.
[0065] In the text sentiment recognition method and recognition system based on BiLSTM and Transformer in the above embodiments of the present disclosure, through the design of the preprocessing module, the data quality and feature expression ability can be significantly improved, laying a solid foundation for subsequent sentiment recognition tasks. The optimization of the preprocessing steps reduces the impact of noise on the model, improves the recognition accuracy, training efficiency, and adaptability to diverse text data, and finally achieves a more accurate and efficient sentiment recognition effect.
[0066] In a possible implementation of the above step embodiment, the preprocessing module uses a pre-trained word embedding model to convert the word sequence into a word embedding representation, including:
[0067] The preprocessing module loads the pre-trained word embedding model, queries the word embedding model based on each word in the word sequence, obtains the word embedding vector of each word, averages the word embedding vector of each word, and obtains the word embedding representation of the entire preset text data.
[0068] In this embodiment, loading a pre-trained word embedding model may include: selecting a suitable pre-trained word embedding model according to task requirements, and using a toolkit provided by the model to load a word embedding weight file into memory.
[0069] Querying the word embedding model based on each word in the word sequence to obtain the word embedding vector of each word may include: querying the word sequence obtained after preprocessing word by word, checking whether each word exists in the loaded word embedding model, if the word exists, obtaining its corresponding word embedding vector, if not, a default vector may be used instead.
[0070] Averaging the word embedding vector of each word to obtain the word embedding representation of the entire preset text data may include: summing up the word embedding vectors of each word in the word sequence, dividing the sum by the number of valid words in the sequence, and obtaining the average word embedding representation of the entire text.
[0071] Furthermore, if the word sequence is empty or no word embedding is found for all words, a default vector can be set as the word embedding representation of the text. The default vector can refer to a zero vector or a random vector.
[0072] In the text emotion recognition method and recognition system based on BiLSTM and Transformer in the above embodiment of the present disclosure, by adopting a pre-trained word embedding model, when facing polysemous words, the word embedding model can effectively capture the tendency of different meanings through the statistical characteristics of the context, making the emotion classification more accurate. By averaging the word embedding vectors, the original text is represented as a semantic vector of fixed dimension, which can ensure the consistency of input dimension regardless of the sentence length, thereby simplifying the subsequent model design.
[0073] In a possible implementation of step S102, the first recognition module uses a BiLSTM model to determine the time sequence information corresponding to the word embedding representation, including:
[0074] The first recognition module adopts the BiLSTM model to obtain the bidirectional hidden state sequence of the word embedding representation as the temporal information corresponding to the word embedding representation.
[0075] In this embodiment, the bidirectional hidden state sequence may refer to a set of hidden layer states generated after forward and backward information transfer is performed on the input word embedding representation through the BiLSTM model.
[0076] Here, the bidirectional hidden state at each time step consists of two parts: the forward hidden state and the backward hidden state. A time step can refer to the "step" or "stage" when the model processes each individual element of the input sequence in a sequence model. In natural language processing, a time step usually corresponds to each word or subword in the input sequence. The forward hidden state can be calculated from the start position of the sequence to the end position, and the backward hidden state can be calculated from the end position of the sequence to the start position in reverse.
[0077] In the text emotion recognition method and recognition system based on BiLSTM and Transformer in the above-mentioned embodiments of the present disclosure, a BiLSTM model is used to obtain a bidirectional hidden state sequence represented by word embedding, and it is used as the timing information corresponding to the word embedding representation, which can effectively capture the temporal dependency and contextual information in the text, thereby providing a more accurate feature representation for emotion recognition.
[0078] In a possible implementation of S103, the second recognition module uses a Transformer model to determine the global dependency information corresponding to the word embedding representation, including:
[0079] The second recognition module adopts the Transformer model to obtain the time-step feature representation sequence of the word embedding representation as the global dependency information corresponding to the word embedding representation.
[0080] In this embodiment, the time-step feature representation sequence may refer to the feature representation sequence extracted by the model for each time step during the sequence data processing. For example, when considering the text sequence "I am happy", the time-step feature representation sequence may be represented as a step-by-step feature representation for each word "I", "am", and "happy".
[0081] In the text emotion recognition method and recognition system based on BiLSTM and Transformer in the above embodiment of the present disclosure, the Transformer model is used to capture the global dependency between each word and other words in the text through the self-attention mechanism, so the generated time-step feature representation sequence can fully reflect the semantic structure of the text. Based on the time-step feature representation sequence, the model can better understand the complex emotional expressions and subtle context changes in the text, thereby improving the performance of emotion recognition.
[0082] In a possible implementation of the above step embodiment, the first recognition module adopts a BiLSTM model to obtain a bidirectional hidden state sequence of the word embedding representation as the time sequence information corresponding to the word embedding representation, including:
[0083] The first recognition module uses the forward LSTM in the BiLSTM model to obtain the forward hidden state of each time step in the word embedding representation. The forward hidden state is used to represent the output sequence of the forward LSTM.
[0084] The first recognition module uses the backward LSTM in the BiLSTM model to obtain the backward hidden state of each time step in the word embedding representation. The backward hidden state is used to represent the output sequence of the backward LSTM;
[0085] The first recognition module concatenates the forward hidden state with the backward hidden state to obtain a bidirectional hidden state sequence for each time step in the word embedding representation.
[0086] In this embodiment, the first recognition module uses the BiLSTM model to obtain time series information, which can be implemented based on the following steps:
[0087] Build a BiLSTM model, use the word embedding representation as the input of the BiLSTM model, make the BiLSTM model output the forward hidden state and the backward hidden state, cascade the forward hidden state and the backward hidden state to obtain the bidirectional hidden state sequence as the feature representation output of the BiLSTM module.
[0088] Specifically, building a BiLSTM model may include: building a model including two LSTM layers, one for processing the input sequence from forward (from left to right), and the other for processing the input sequence from backward (from right to left). Each LSTM layer may include a forget gate, an input gate, and an output gate to maintain and update the cell state.
[0089] The forward hidden state at each time step can be calculated using the following formula:
[0090] i t =σ(W xi x t +W hi h t-1 +b i )
[0091] f t =σ(W xf x t +W hf h t-1 +b f )
[0092] o t=σ(W xo x t +W ho h t-1 +b o )
[0093] c t =ft t *c t-1 +i t *tanh(W xc x t +W hc h t-1 +b c )
[0094] h t =o t *tanh(c t )
[0095] Among them, i t can represent the input gate, x t It can represent the input of the current time step, h t-1 can represent the hidden state of the previous time step, σ can represent the activation function, which is used to indicate the importance of the current input, W xi Represents the input weight matrix, W hi The weight matrix representing the hidden state, b i is the bias term. Similarly, f t Can represent the forget gate, W xf and W hf is the corresponding weight matrix, b f is the bias term. t represents the output gate, W xo and W ho is the corresponding weight matrix, b o is the bias term. c t Indicates the cell state, ft t *c t-1 represents the forgotten part of the cell state at the previous time step, i t *tanh(W xc x t +W hc h t-1 +b c ) indicates the increment of the current input information. t represents the hidden state, tanh(c t ) is a nonlinear activation of the cell state, ensuring that the values in the cell state are not too large or too small.
[0096] The calculation formula of the forward hidden state at each time step is similar to the calculation formula of the forward hidden state, but in the opposite direction (i.e., t-1 is changed to t+1).
[0097] Concatenation can refer to connecting multiple tensors together in a certain dimension. Concatenation of bidirectional hidden states can refer to concatenating the forward hidden state and the backward hidden state together to obtain a bidirectional hidden state representation at each time step.
[0098] The cascade of bidirectional hidden states can be expressed as follows:
[0099]
[0100] Among them, h bi It can represent the bidirectional hidden state at the tth time step, including the forward hidden state and the backward hidden state
[0101] In the text emotion recognition method and recognition system based on BiLSTM and Transformer in the above-mentioned embodiment of the present disclosure, the BiLSTM model can obtain the contextual information of each word in the text by using the forward LSTM and the backward LSTM at the same time, which helps the CNN model to better understand the meaning of each word in the global context. After cascading the forward hidden state and the backward hidden state, the model can obtain a more comprehensive text representation, thereby providing more contextual clues and improving the accuracy of emotion recognition. BiLSTM has a certain memory capacity and can adapt to long-distance dependencies in long texts, so that it can process long texts containing complex emotional expressions.
[0102] In a possible implementation of the above step embodiment, the second recognition module uses a Transformer model to obtain a time-step feature representation sequence of the word embedding representation as global dependency information corresponding to the word embedding representation, including:
[0103] The second recognition module performs position encoding on the word embedding representation and introduces the position information of at least one word in the word embedding representation into the Transformer model;
[0104] The second recognition module adopts a multi-head self-attention mechanism in at least one encoder layer of the Transformer model to determine the context dependency between each word in at least one word in the word embedding representation, and determines the time-step-by-time features of the word embedding representation based on the context dependency.
[0105] In this embodiment, the second recognition module uses a Transformer model to obtain global dependency information, which can be implemented based on the following steps: constructing a Transformer model including multiple self-attention layers and feedforward neural network layers, position encoding the word embedding representation, and introducing the position information of at least one word in the word embedding representation into the encoder part of the Transformer model;
[0106] In the Transformer model, the self-attention mechanism of the model is used to calculate a weighted combination representation for each position in each input sequence to capture the dependencies between different positions;
[0107] In each self-attention layer, the self-attention weight of each word is calculated to determine the relationship between this word and other words;
[0108] By stacking multiple self-attention layers, the time-step feature representation sequence of the word embedding representation is obtained as the global dependency information corresponding to the word embedding representation. Among them, the Transformer model is composed of multiple encoder layers stacked.
[0109] Specifically, position encoding can add a position to each word in the word embedding representation through a fixed function, and the position can indicate the position of the word in the sequence.
[0110] The functions here may include but are not limited to: sine function and cosine function.
[0111] For example, the following formula may be used for position encoding:
[0112]
[0113] Among them, PE(pos,2i) can represent the encoding value of position pos in dimension 2i, PE(pos,2i+1) can represent the encoding value of pos in dimension 2i+1, and d model Can represent the dimension of word embedding in the model.
[0114] In the self-attention mechanism, the Transformer model attempts to calculate the association between different words in the input sequence in order to capture the relationship between each word and other words in the sequence. The representation of each word in the sequence depends not only on its own features, but also on its relationship with other words. In this way, the model can understand the context and generate more semantically informative representations. The core of the self-attention mechanism is to calculate the attention distribution and then determine the association between different positions (i.e. words).
[0115] As an example, the attention distribution can be calculated using the following formula:
[0116]
[0117] Q can represent the query matrix, K can represent the key matrix, i.e. the “identity” of each input word, V can represent the value matrix, i.e. the “content” of each input word, d k That is, the dimension of the key, which is used to scale the calculated dot product to prevent the dot product value from being too large.
[0118] Furthermore, the following formula can be used to calculate multi-head self-attention:
[0119] Mul(Q,K,V)=CONCAT(head1,head2,...,head h )W O
[0120] Among them head i represents the output of the i-th attention head, CONCAT represents concatenating the outputs of multiple heads, and W O is the output weight matrix.
[0121] In a possible implementation of the above embodiment, the comprehensive recognition module performs feature fusion on the time series information and the global dependency information, uses a convolutional neural network to perform emotion recognition on the result of the feature fusion, and obtains the emotion recognition result of the preset text data, which may include:
[0122] The comprehensive recognition module uses weighted summation to perform feature fusion.
[0123] Specifically, the calculation formula for feature fusion can be as follows:
[0124] f = α * BiLSTMF t +(1-α)TransformerF t
[0125] Among them, α is the weight of the BiLSTM feature, 1-α is the weight of the Transformer feature, and f is the fused feature vector.
[0126] In the text emotion recognition method and recognition system based on BiLSTM and Transformer of the above-mentioned embodiments of the present disclosure, the second recognition module uses the Transformer model to generate a time-step feature representation with context dependency through position encoding and multi-head self-attention mechanism, which provides more accurate feature information for subsequent sentiment analysis.
[0127] In one embodiment, Figure 2 An exemplary schematic diagram of the architecture of another recognition system applied to the text emotion recognition method based on BiLSTM and Transformer in an embodiment of the present disclosure is shown, as shown in FIG. Figure 2As shown:
[0128] The preprocessing module is used to obtain the original text data (i.e., the preset text data), preprocess the original text data, use the Word2Vec model to convert the text data into a word embedding representation, and send the word embedding representation to the first recognition module and the second recognition module at the same time. The first recognition module is used to determine the temporal information of the input word embedding representation in multiple BiLSTM models, taking into account the forward and backward context information. The second recognition module is used to positionally encode the word embedding representation based on the position encoder, add the position information of each word in the sentence to the word vector, so that the Transformer model can perceive the position of the word in the text, and use the encoder in the Transformer to obtain global dependency information. The comprehensive recognition module is used to perform feature fusion of the temporal information and the global dependency information, use a convolutional neural network to perform sentiment recognition on the result of the feature fusion, and obtain the sentiment recognition result of the original text data.
[0129] In one embodiment, Figure 3 is a flowchart of another text emotion recognition method based on BiLSTM and Transformer provided by the embodiment of the present disclosure, which is applied to the above Figure 1b or Figure 2 In any of the identification systems shown in the figure, the method process may include the following steps:
[0130] Step S301, obtaining preset text data;
[0131] Step S302, preprocessing the preset text data;
[0132] Step S303, the BiLSTM model obtains time series information;
[0133] Step S304, the Transformer model obtains global dependency information;
[0134] Here, step S303 and step S304 are executed in parallel;
[0135] Step S305, feature fusion;
[0136] Step S306: text emotion recognition.
[0137] In one embodiment, a recognition system 40 is provided, which corresponds one-to-one to the text emotion recognition method based on BiLSTM and Transformer in the above embodiment. Figure 4 As shown, the recognition system 40 includes a pre-processing module 401, a first recognition module 402, a second recognition module 403 and a comprehensive recognition module 404, wherein each functional module is described in detail as follows:
[0138] A preprocessing module 401 is used to obtain preset text data, preprocess the preset text data, convert the preset text data into a word embedding representation, and send the word embedding representation to the first recognition module and the second recognition module at the same time;
[0139] A first recognition module 402 is used to determine the time series information corresponding to the word embedding representation using a BiLSTM model;
[0140] A second recognition module 403 is used to determine the global dependency information corresponding to the word embedding representation by using a Transformer model; the first recognition module and the second recognition module are run in parallel;
[0141] The comprehensive recognition module 404 is used to perform feature fusion on the temporal information and the global dependency information, and use a convolutional neural network to perform emotion recognition on the result of the feature fusion to obtain the emotion recognition result of the preset text data.
[0142] In a specific embodiment, the preprocessing module 401 is used to segment the preset text data into word sequences using a word segmentation tool, and remove stop words in the word sequences based on a preset stop word list;
[0143] The preprocessing module 401 is used to convert the letters in the word sequence into uppercase and lowercase letters and clean special symbols, and use a pre-trained word embedding model to convert the word sequence into a word embedding representation.
[0144] In a specific embodiment, the preprocessing module 401 is used to load a pre-trained word embedding model, query the word embedding model based on each word in a word sequence, obtain the word embedding vector of each word, average the word embedding vector of each word, and obtain the word embedding representation of the entire preset text data.
[0145] In a specific embodiment, the first recognition module 402 is used to adopt a BiLSTM model to obtain a bidirectional hidden state sequence of a word embedding representation as the timing information corresponding to the word embedding representation.
[0146] In a specific embodiment, the second recognition module 403 is used to adopt a Transformer model to obtain a time-step feature representation sequence of a word embedding representation as global dependency information corresponding to the word embedding representation.
[0147] In a specific embodiment, the first recognition module 402 is used to use the forward LSTM in the BiLSTM model to obtain the forward hidden state of each time step in the word embedding representation, and the forward hidden state is used to characterize the output sequence of the forward LSTM;
[0148] Using the backward LSTM in the BiLSTM model, obtaining a backward hidden state of each time step in the word embedding representation, wherein the backward hidden state is used to represent an output sequence of the backward LSTM;
[0149] The forward hidden state is cascaded with the backward hidden state to obtain a bidirectional hidden state sequence for each time step in the word embedding representation.
[0150] In a specific embodiment, the second recognition module 403 is used to perform position encoding on the word embedding representation, and introduce the position information of at least one word in the word embedding representation into the Transformer model;
[0151] A multi-head self-attention mechanism is used in at least one encoder layer of the Transformer model to determine a context dependency relationship between each word in at least one word in the word embedding representation, and a time-step feature of the word embedding representation is determined based on the context dependency relationship.
[0152] It should be noted that: the recognition system 40 provided in the above embodiment only uses the division of the above program modules as an example to implement the corresponding text emotion recognition method based on BiLSTM and Transformer. In actual application, the above processing can be assigned to different program modules as needed, that is, the internal structure of the above system can be divided into different program modules to complete all or part of the above-described processing. Figure 1b The embodiments of the method shown belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0153] The present disclosure also provides a computer device having the above Figure 4 The use case execution system shown.
[0154] See also Figure 5 , Figure 5 is a schematic diagram of the structure of another identification system provided by an embodiment of the present disclosure, such as Figure 5As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.
[0155] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0156] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0157] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0158] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0159] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 5 The example of connecting through bus is taken in the following.
[0160] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0161] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0162] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0163] A part of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes, but is not limited to, source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0164] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A text emotion recognition method based on BiLSTM and Transformer, characterized in that: Applied to an identification system, the identification system comprises: a preprocessing module, a first identification module, a second identification module and a comprehensive identification module, the method comprises: The preprocessing module acquires preset text data, preprocesses the preset text data, converts the preset text data into word embedding representation, and sends the word embedding representation to the first recognition module and the second recognition module at the same time; The first recognition module uses a BiLSTM model to determine the time sequence information corresponding to the word embedding representation; The second recognition module uses a Transformer model to determine the global dependency information corresponding to the word embedding representation; the first recognition module and the second recognition module run in parallel; The comprehensive recognition module performs feature fusion on the time series information and the global dependency information, uses a convolutional neural network to perform emotion recognition on the result of the feature fusion, and obtains the emotion recognition result of the preset text data.
2. The method according to claim 1, characterized in that The preprocessing module obtains preset text data, preprocesses the preset text data, and converts the preset text data into a word embedding representation, including: A preprocessing module, using a word segmentation tool to segment the preset text data into word sequences, and based on a preset stop word list, removing stop words in the word sequence; The preprocessing module performs case conversion and special symbol cleaning on the letters in the word sequence, and uses a pre-trained word embedding model to convert the word sequence into a word embedding representation.
3. The method according to claim 2, characterized in that The preprocessing module uses a pre-trained word embedding model to convert the word sequence into a word embedding representation, including: The preprocessing module loads a pre-trained word embedding model, queries the word embedding model based on each word in the word sequence, obtains the word embedding vector of each word, averages the word embedding vector of each word, and obtains the word embedding representation of the entire preset text data.
4. The method according to claim 2, characterized in that: The first recognition module uses a BiLSTM model to determine the time sequence information corresponding to the word embedding representation, including: The first recognition module adopts the BiLSTM model to obtain the bidirectional hidden state sequence of the word embedding representation as the temporal information corresponding to the word embedding representation.
5. The method according to claim 4, characterized in that The second recognition module uses a Transformer model to determine the global dependency information corresponding to the word embedding representation, including: The second recognition module adopts the Transformer model to obtain the time-step feature representation sequence of the word embedding representation as the global dependency information corresponding to the word embedding representation.
6. The method according to claim 4, characterized in that The first recognition module adopts the BiLSTM model to obtain the bidirectional hidden state sequence of the word embedding representation as the time sequence information corresponding to the word embedding representation, including: The first recognition module uses the forward LSTM in the BiLSTM model to obtain the forward hidden state of each time step in the word embedding representation, where the forward hidden state is used to represent the output sequence of the forward LSTM; The first recognition module adopts the backward LSTM in the BiLSTM model to obtain the backward hidden state of each time step in the word embedding representation, wherein the backward hidden state is used to represent the output sequence of the backward LSTM; The first recognition module cascades the forward hidden state and the backward hidden state to obtain a bidirectional hidden state sequence for each time step in the word embedding representation.
7. The method according to any one of claim 5, characterized in that: The second recognition module adopts the Transformer model to obtain a time-step feature representation sequence of the word embedding representation as the global dependency information corresponding to the word embedding representation, including: The second recognition module performs position encoding on the word embedding representation and introduces the position information of at least one word in the word embedding representation into the Transformer model; The second recognition module adopts a multi-head self-attention mechanism in at least one encoder layer of the Transformer model to determine the context dependency between each word in at least one word in the word embedding representation, and determines the time-step features of the word embedding representation based on the context dependency.
8. A recognition system, characterized in that: The recognition system comprises a preprocessing module, a first recognition module, a second recognition module and a comprehensive recognition module, wherein: The preprocessing module is used to obtain preset text data, preprocess the preset text data, convert the preset text data into word embedding representation, and send the word embedding representation to the first recognition module and the second recognition module at the same time; The first recognition module is used to determine the time sequence information corresponding to the word embedding representation using a BiLSTM model; The second recognition module is used to determine the global dependency information corresponding to the word embedding representation using a Transformer model; the first recognition module and the second recognition module run in parallel; The comprehensive recognition module is used to perform feature fusion on the temporal information and the global dependency information, use a convolutional neural network to perform emotion recognition on the result of the feature fusion, and obtain the emotion recognition result of the preset text data.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the text emotion recognition method based on BiLSTM and Transformer described in claims 1-7.
10. A computer program product, characterized in that It includes computer instructions, which are used to enable a computer to execute the text emotion recognition method based on BiLSTM and Transformer as described in any one of claims 1-7.