Sentiment Recognition Method Based on Convolutional Recurrent Neural Network and Multi-Head Self-Attention

By combining convolutional recursive neural networks and multi-head self-attention mechanisms, the problem of feature extraction and classification in EEG emotional recognition is solved, the recognition accuracy is improved and the training time is reduced, and the effective processing of EEG signals is achieved.

CN115238731BActive Publication Date: 2025-07-08BEIJING SHUFENG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210665185.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2025-07-08
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

The existing EEG emotion recognition methods have difficulties in feature extraction and classification, especially the lack of time series processing capabilities in convolutional neural networks, and the attention mechanism fails to effectively improve the recognition accuracy, resulting in overfitting and gradient disappearance problems.

Method used

The method based on convolutional recursive neural network and multi-head self-attention is adopted. The spatial characteristics of EEG signals are extracted through convolutional neural networks, combined with the bidirectional long and short-term memory network to learn the dynamic characteristics of the time series, and the multi-head self-attention mechanism is used to redistribute the key information weight.

Benefits of technology

It improves the accuracy of EEG emotional recognition, reduces training time and avoids overfitting, and enhances the dynamic time feature extraction and recognition ability of EEG signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115238731B_ABST
    Figure CN115238731B_ABST
Patent Text Reader

Abstract

The present invention claims protection for an emotion recognition method based on a convolutional recurrent neural network and a multi-head self-attention mechanism, which uses one-dimensional convolution (CNN) and bidirectional long short-term memory network (BiLSTM) to extract the spatial and dynamic time features of electroencephalogram (EEG) signals, and uses a fully connected layer to fuse these features and clone them to the multi-head self-attention mechanism (Multi-Head Self-Attention) to redistribute the weights of the emotional key information, so as to obtain an accurate recognition of the emotional state. The designed model was verified on the Database for Emotion Analysis and Physiological Signals (DEAP), and the designed model was compared with other emotion recognition models. The experimental results show that the convolutional smoothing signal of the convolutional network can greatly improve the recognition ability of the LSTM network. When the input is an EEG time series, BiLSTM can effectively learn the key information of the past and future of the time series, and at the same time, the multi-head self-attention mechanism can redistribute the weights to improve the accuracy. Compared with the methods in recent years, the proposed model still achieves remarkable results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electroencephalogram emotion recognition, and is an emotion recognition method based on a convolutional recurrent neural network and a multi-head self-attention mechanism. Background Art

[0002] Emotion plays an important role in human daily life. It affects all aspects of people, including human decision-making, speech, sleep patterns, health, communication, and various other characteristics. Emotion recognition is often implemented based on facial expressions, speech, and physiological signals, etc. Physiological signals can more accurately reflect the fluctuations of human emotional states, and have better robustness and noise resistance. Therefore, emotion recognition based on multi-channel electroencephalogram time series signals has become a research hotspot in the field of emotion computing.

[0003] Electroencephalogram is the electrical signal on the epidermis of the human brain, which has the characteristics of non-linearity and non-stationarity. Feature extraction and classification of such signals have always been difficult problems for researchers. As a branch of machine learning, deep learning has shown remarkable effects in processing EEG feature information and classification. Compared with various existing deep learning technologies, the use of convolutional neural networks has begun to increase because they can automatically extract discriminative features and classify. And in order to solve the problem that convolutional neural networks lack time series, the LSTM model is added.

[0004] CNN is a very effective image processing and classification model. The architecture uses convolutional operations to extract various features of data, and then passes the features to the next layer. The convolutional smoothed signal processed by CNN has a significant improvement on the LSTM network. Many previous networks did not consider the improvement of the attention mechanism on recognition, and lacked the convolutional smoothed signal of CNN, which made it difficult for LSTM to learn time series. Moreover, in the end-to-end method, the key information of the past and future states of the EEG time series cannot be obtained, and the weight reallocation of the key information of emotion is not achieved.

[0005] CN113724732A, a convolutional recurrent neural network model based on the fusion of multi-head attention mechanisms. First, a fully convolutional network is proposed to extract speech spectrogram emotion features. This network is based on the Alexnet network, and by adding branches after the pooling layer of the Alexnet network, the loss of emotion information is prevented; a 2-layer BiLSTM network is used to extract speech frame-level emotion features, and the BiLSTM network is connected in parallel with the fully convolutional network to form a hybrid network for extracting speech emotion features. Second, a feature fusion algorithm based on multi-head attention mechanism is proposed. This method uses the multi-head attention mechanism to achieve the adaptive fusion of features between the Alexnet network and the BiLSTM network. At the same time, to suppress the divergence of network gradients, the features extracted by the hybrid network and the multi-head attention fusion features are connected through shortcut connection to form features for emotion recognition. Finally, the features are fed into a softmax classifier to achieve emotion classification. The hybrid network of Alexnet and BiLSTM in this patent is prone to too deep layers and serious overfitting. The convolutional network in the hybrid structure of CNN and BiLSTM in the present invention is a single-layer one-dimensional convolution, which greatly reduces the network depth, speeds up the training and inference speed and avoids overfitting. This patent uses multi-head attention without self-attention, and the effect of weight allocation for key information is not good. In the present invention, the results output by BiLSTM are cloned to Q, K, V of multi-head self-attention, which can accurately allocate the weights of its own key information.

[0006] CN113450830A, a speech emotion recognition method for a convolutional recurrent neural network with multiple attention mechanisms, includes: Step 1, extract spectrogram features and frame-level features. Step 2, the spectrogram features are fed into the CNN module to learn the time-frequency related information in the features. Step 3, the multi-head self-attention layer acts on the CNN module to calculate the weights of different frames under different scales of global features and fuse the features of different depths in the CNN. Step 4, a multi-dimensional attention layer acts on the frame-level features input to the LSTM to comprehensively consider the relationship between local features and global features. Step 5, the processed frame-level features are fed into the LSTM model to obtain the time information in the features. Step 6, a fusion layer summarizes the outputs of different modules to enhance the performance of the model. Step 7, use the Softmax classifier to classify different emotions. The present invention combines a deep learning network, and a parallel connection structure is adopted inside the module to process features simultaneously, which can effectively improve the performance of speech emotion recognition. This patent uses LSTM to perform frame-level processing on the input sequence, which can greatly improve the recognition, but it does not have a deep understanding of the front and back time relationships in the sequence and lacks the learning of front and back features. The present invention uses a BiLSTM network to deeply learn the dynamic time features of the past and the future, and uses multi-head self-attention to re-allocate the weights of key emotion features to improve the recognition accuracy. Summary of the Invention

[0007] The present invention aims to solve the problems of the above prior art. A sentiment recognition method based on a convolutional recurrent neural network and multi-head self-attention is proposed. The technical solution of the present invention is as follows:

[0008] A sentiment recognition method based on a convolutional recurrent neural network and multi-head self-attention, comprising the following steps:

[0009] S1. Use a convolutional neural network CNN to extract the spatial features of electroencephalogram (EEG) signals, and preprocess the spatial features of EEG signals using a 4.0 - 45.0 Hz band-pass filter;

[0010] S2. Regularize and pool the EEG signals with spatial features;

[0011] S3. Input the convolution-smoothed signals into a BiLSTM (Bidirectional Long Short-Term Memory) network to learn the dynamic time features of the EEG time series, and obtain the past and future key sentiment information of the EEG signals;

[0012] S4. Finally, use the multi-head self-attention mechanism to redistribute the weights of the key EEG sentiment information.

[0013] Further, in step S1, a feature extractor composed of a convolutional neural network CNN extracts the spatial features of multi-channel EEG signals, and then transfers the spatial features of the EEG signals to the regularization and pooling layers. The convolution-smoothed signals after being processed by CNN also introduce an LSTM network to increase the time series.

[0014] Further, in step S2, first use the regularization of the Relu activation function to avoid overfitting and accelerate learning, and then transfer the feature signals to the next layer of the BiLSTM (Bidirectional Long Short-Term Memory) network through pooling.

[0015] Further, the LSTM network shares weights, and the weights between the hidden layer and the output layer can be recycled at any time. The LSTM network is a chain model for processing time series, which can effectively compensate for the problem of vanishing gradients. The bidirectional EEG signal extraction method can simultaneously extract the dynamic information of the earlier and later segments in the EEG signal sequence; an LSTM unit consists of three gate control units: a forget gate, a memory gate, and an output gate.

[0016] Further, the LSTM unit consists of three gate control units: a forget gate, a memory gate, and an output gate. The calculation formula is as follows:

[0017] f t = σ(W f ·[h t-1 ,xt ) + b f ),

[0018] i t = σ(W i · [h t-1 , x t ) + b i ),

[0019]

[0020]

[0021] O t = σ(W O · [h t-1 , x t ) + b O ),

[0022] h t = tanh(C t ) × O t ,

[0023] Among them, x t is the time series at time t, C t represents the cell state, is the temporary cell state, σ is the sigmoid function, W is the weight matrix, b is the bias vector corresponding to the weight, h t is the hidden state, f t is the forget gate, i t is the memory gate, O t is the output gate; the forget gate selects the retained features, inputs the information of the previous state and the current state into the sigmoid function at the same time, the memory gate is responsible for updating the state of the LSTM unit, and then the input gate controls the output value to the next LSTM unit.

[0024] Furthermore, in step S4, finally, the multi-head self-attention mechanism is used to reallocate the weights of the key EEG emotion information, specifically including:

[0025] The Transformer model is an autoregressive generative model that uses self-attention mechanism and sinusoidal position information. Each layer includes a time self-attention sublayer, a feed-forward network sublayer, a residual network sublayer, and a dropout layer; the time self-attention sublayer is used to capture key information, the feed-forward network sublayer is used to normalize the hidden layer in the neural network to the standard normal distribution to accelerate convergence, the residual network sublayer alleviates the problem of gradient disappearance, enabling the network to be deeper, and the dropout layer is used to reduce the effect of overfitting.

[0026] Attention essentially assigns a weight coefficient to each element in the EEG sequence. If each element is stored, attention can calculate the similarity between Q and K; the similarity calculated from Q and K reflects the importance of the extracted V value, that is, the weight, and then the weighted sum is used to obtain the attention value. The special point of the self-attention mechanism in the K, Q, V model is that Q = K = V. Q refers to the vector group composed of the sequence at the output end, K refers to the various weights corresponding to each vector of the input sequence, and V refers to the vector group composed of the input sequence. The scaled dot-product attention formula is as follows:

[0027]

[0028] Furthermore, it also includes the step of training using a training model. The training model is a structure of a bidirectional double-layer LSTM, a multi-head self-attention mechanism, and a single-layer convolutional neural network, simply referred to as the CNN1D_BiLSTM_MHSA network. The input data passes through a one-dimensional convolutional network with a convolutional kernel of 128, and then through Batch Normalisation regularization and the Relu activation function, followed by one-dimensional pooling with a kernel of 3 on the data. The obtained data sequence is fed into the BiLSTM deep network. Each LSTM network layer has 256 hidden units, and the output structure is concatenated using a fully connected layer. The obtained 512 data sequences are respectively cloned to the Q, K, V of the multi-head self-attention, and then weights are assigned through scaled dot-product attention. Finally, the softmax activation function is used to obtain the classification result.

[0029] The advantages and beneficial effects of the present invention are as follows:

[0030] The present invention verifies the designed model on the publicly available dataset DEAP and compares the designed model with other EEG emotion recognition models. The experimental results show that the introduction of a convolutional recurrent neural network and a multi-head self-attention mechanism can further improve the accuracy; when the input is EEG time series information, combining CNN and BiLSTM can effectively extract spatial and dynamic time features. And the proposed multi-head self-attention mechanism can further enhance the recognition accuracy on this basis. By comparing with the advanced methods of other EEG emotion recognition models in recent years, the model proposed by the present invention still achieves superior performance

[0031] The innovation of the present invention is mainly the combination of steps 1, 3, and 4. For the processing of EEG signals, the amount of general data is insufficient to support the training of a relatively deep number of layers, and overfitting is likely to occur. Using a single-layer one-dimensional convolution to reduce the number of training layers can accelerate the speed of network training and inference and extract the spatial information in EEG. There is a lack of research on the dynamic time characteristics existing in EEG signals, and BiLSTM is used to learn the key emotional information in the past and future. Under the above steps, a multi-head self-attention mechanism is added to the output result of BiLSTM to redistribute the weights of its key emotional features, supplementing the research on EEG in the attention mechanism. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is the preferred embodiment CNN1D-BiLSTM-MHSA training framework provided by the present invention;

[0033] Figure 2 is the structural diagram of BiLSTM;

[0034] Figure 3 is the structural diagram of Multi-Head Self-Attention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.

[0036] The technical solution for the present invention to solve the above technical problems is:

[0037] The estimation method of the present invention includes the following steps:

[0038] S1, using a convolutional network CNN to extract the spatial features of electroencephalogram (EEG) signals

[0039] S2, regularizing and pooling the processed EEG signals;

[0040] S3, inputting the convolution-smoothed signal into BiLSTM to learn the dynamic time features of the EEG time series and the key emotional information in the past and future of the EEG signal;

[0041] S4, finally using a multi-head self-attention mechanism to redistribute the weights of the key EEG emotional information;

[0042] S5, constructing the parameters of the model and performing model training.

[0043] Further, in step S1, the architecture uses convolutional operations to extract various features of the data and then passes the features to the next layer. A feature extractor composed of CNN is used to extract the spatial features of multi-channel EEG signals. The convolutional smoothed signals processed by CNN have a significant improvement on the LSTM network.

[0044] Further, in step S2, first, the regularization of the Relu activation function is used to avoid overfitting and accelerate learning, and then the feature signals are passed to the next layer through pooling.

[0045] Further, in step S3, the weights between the hidden layer and the output layer of LSTM can be recycled at any time. It is a chain model used to process time series and can effectively compensate for the vanishing gradient problem. Compared with the classical unidirectional EEG signal extraction method, the bidirectional EEG signal extraction method can extract the dynamic information of the earlier and later segments in the EEG signal sequence at the same time.

[0046] Further, in step S4, the multi-head self-attention mechanism obtains different representations of h (i.e., each head) of (Q, K, V), calculates the self-attention of each representation, and connects the results. The weight of key emotional information can be effectively allocated during the training process.

[0047] Further, in step S5, network training is carried out based on the cross-entropy function optimization and stochastic gradient descent (SGD) with backpropagation. The weight sharing of CNN usually leads to different gradient changes in different layers. For this reason, a single-layer convolutional neural network is adopted and a small learning rate is used. Since the double-layer bidirectional LSTM has a large depth, a small number of iteration times (epoch = 50) are sufficient for convergence. To eliminate the influence of the overfitting problem, dropout = 0.5 is adopted in the training network. The convolutional layer uses a single-layer CNN with a kernel of 128 to extract the frequency-domain features of the EEG information sequence, and then the LSTM unit analyzes the time-domain features. Finally, the multi-head self-attention mechanism is used to improve the classification accuracy. In this method, the bidirectional LSTM layer has 256 units (a total of 512), and the cases with 128 and 512 units in the hidden layer are also studied, and the best-performing 256 units are selected. And the best-performing 8 heads are selected from 2 / 4 / 6 / 8 for the multi-head self-attention mechanism. The following table lists the hyperparameters of the training network.

[0048] Table 1 Hyperparameters of CNN

[0049]

[0050] Table 2 Hyperparameters of BiLSTM

[0051]

[0052]

[0053] Table 3 Hyperparameters of the Multi-Head Self Attention Structure

[0054]

[0055] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0056] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.

[0057] The above embodiments should be understood as being only for illustrative purposes of the present invention and not for limiting the protection scope of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A sentiment recognition method based on convolutional recurrent neural network and multi-head self-attention, characterized in that Including the following steps: S1. Extract the spatial features of the electroencephalogram (EEG) signals using a convolutional neural network (CNN), and preprocess the spatial features of the EEG signals using a 4.0 - 45.0 Hz band-pass filter; S2. Regularize and pool the preprocessed EEG signals with spatial features; S3. Input the convolutional smoothed signals into a bidirectional long short-term memory (BiLSTM) network to learn the dynamic time features of the EEG time series, and obtain the past and future key emotional information of the EEG signals; S4. Finally, use the multi-head self-attention mechanism of the Transformer to reallocate the weights of the key EEG emotional information; it also includes the step of training using a training model. The training model is a structure of a bidirectional double-layer LSTM, a multi-head self-attention mechanism, and a single-layer convolutional neural network, abbreviated as the CNN1D_BiLSTM_MHSA network. The input data passes through a one-dimensional convolutional network with a convolutional kernel of 128, then undergoes Batch Normalisation regularization and a Relu activation function, and then one-dimensional pooling with a kernel of 3 is performed on the data. The obtained data sequence is fed into the BiLSTM deep network. Each layer of the LSTM network has 256 hidden units, and the output structure is concatenated using a fully connected layer. The obtained 512 data sequences are respectively cloned to the Q, K, and V of the multi-head self-attention, and then weights are allocated through scaled dot-product attention. Finally, a softmax activation function is used to obtain the classification result.

2. The emotional recognition method based on convolutional recurrent neural network and multi-head self-attention according to claim 1, characterized in that In step S1, a feature extractor composed of a CNN extracts the spatial features of multi-channel EEG signals, and then transfers the spatial features of the EEG signals to the regularization and pooling layers. The convolutional smoothed signals processed by the CNN are also processed by introducing an LSTM network to handle the time series.

3. The emotional recognition method based on convolutional recurrent neural network and multi-head self-attention according to claim 2, characterized in that In step S2, first, the regularization of the Relu activation function is used to avoid overfitting and accelerate learning, and then the feature signals are transferred to the BiLSTM bidirectional long short-term memory network through pooling.

4. The emotional recognition method based on convolutional recurrent neural network and multi-head self-attention according to claim 2, characterized in that The LSTM network can recycle the weights between the hidden layer and the output layer through shared weights. The LSTM network is a chain model for processing time series, which can effectively compensate for the vanishing gradient problem. The bidirectional EEG signal extraction method can simultaneously extract the dynamic information of the earlier and later segments in the EEG signal sequence; an LSTM unit consists of three gate control units: a forget gate, a memory gate, and an output gate.

5. A sentiment recognition method based on a convolutional recurrent neural network and multi-head self-attention according to claim 4, characterized in that, The LSTM unit consists of three gate control units: a forget gate, a memory gate, and an output gate, and the calculation formulas are as follows: f t = σ(W f · [h t-1 , x t ) + b f ), i t = σ(W i · [h t-1 , x t ) + b i ), O t = σ(W O · [h t-1 , x t ) + b O ), h t = tanh(C t ) × O t , where x t is the time series at time t, C t represents the cell state, is the temporary cell state, σ is the sigmoid function, W is the weight matrix, b is the bias vector corresponding to the weight, h t is the hidden state, f t is the forget gate, i t is the input gate, O t is the output gate; the forget gate selects the retained features, inputs the information of the previous state and the current state into the sigmoid function at the same time, the input gate is responsible for updating the state of the LSTM unit, and then the input gate controls the output value to the next LSTM unit.

6. The emotional recognition method based on convolutional recurrent neural network and multi-head self-attention according to claim 5, wherein In step S4, finally, the multi-head self-attention mechanism is used to reallocate the weights of the key EEG emotional information, specifically including: The Transformer model is an autoregressive generative model that uses self-attention mechanisms and sinusoidal positional information. Each layer includes a temporal self-attention sublayer, a feed-forward network sublayer, a residual network sublayer, and a dropout layer. The temporal self-attention sublayer is used to capture key information. The feed-forward network sublayer is used to normalize the hidden layer in the neural network to the standard normal distribution to accelerate convergence. The residual network sublayer alleviates the problem of vanishing gradients, enabling the network to be deeper. The dropout layer is used to reduce the effect of overfitting. Attention essentially assigns a weight coefficient to each element in the EEG sequence. If each element is stored, attention can calculate the similarity between Q and K. The similarity calculated from Q and K reflects the importance of the extracted V value, that is, the weight, and then the weighted sum is used to obtain the attention value. The special point of the self-attention mechanism in the K, Q, V model is that Q = K = V. Q refers to the vector group composed of the sequence at the output end. K refers to the various weights corresponding to each vector in the input sequence. V refers to the vector group composed of the input sequence. The scaled dot-product attention formula is as follows:

Citation Information

Patent Citations

  • Voice emotion recognition method of convolutional recurrent neural network with multiple attention mechanisms

    CN113450830A

  • Convolutional recurrent neural network model based on multi-head attention mechanism fusion

    CN113724732A