A network emotion recognition method based on CCNN and stacked-BiLSTM

CN115444420BActive Publication Date: 2026-09-25KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211106511.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-12
Publication Date
2026-09-25
Estimated Expiration
2042-09-12

AI Technical Summary

Technical Problem

但是在实验研究中发现:1、网络训练时间长、计算机计算成本大;2、无法从高维度的脑电信号中获取到丰富的特征;3、基于脑电的情绪识别数据本身具有的稀疏性导致数据在输入网络之后一些特征无法提取,导致训练效果并不能达到所期待的效果

Benefits of technology

[0032]本发明的有益效果是:与现有技术相比,本发明利用微分熵特征和大脑地形图得到丰富的空频特征作为输入,然后采用多并行连续卷积神经网络-堆叠双向长短时记忆获取更丰富且全面的特征信息,从而在识别分类任务中达到一个高的精确度,为更好的识别情绪提供了一种新的思路

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115444420B_ABST
    Figure CN115444420B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of network emotion recognition method based on CCNN and stacked-BiLSTM, belong to electroencephalogram emotion recognition technical field.The present application is first to data pre-processing: using Butterworth filter decoding out 4 frequency bands theta, alpha, beta, gamma relevant to emotion recognition, using 0.5s non-overlapping sliding window to the 60s physiological signal data relevant to emotion recognition under audio stimulation adopts moving average method to eliminate the fluctuation on emotion caused by artifact, noise and other factors in the process of experimental paradigm, then by extracting 4 frequency band frequency domain feature differential entropy and brain plane topographic map mapping out the electroencephalogram spatial structure of relevant electroencephalogram channel of emotion recognition, a large number of frequency domain and spatial features after sliding window are input to each continuous convolutional neural network in parallel to extract higher semantic space-frequency features, finally using stacked bidirectional long short-term memory to learn information on past and future time slice fully.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a network emotion recognition method based on CCNN and stacked-BiLSTM, belonging to the field of EEG emotion recognition technology. Background Technology

[0002] Humans, as advanced beings, possess complex psychological activities and can mask their true emotional state through facial expressions, vocal information, and body movements. Neuroscience, a recognized cutting-edge technology, recognizes that brainwaves generated by brain regions possess objectivity, reflecting relatively objective facts under objective conditions. Therefore, emotion recognition based on EEG is one of the current research tasks related to EEG information and has significant practical implications. It currently has wide applications in industrial control, medical assistance, and gaming. EEG signals for emotion recognition require external stimuli to trigger the brain to generate corresponding information. However, because EEG waves are relatively small amplitudes compared to external noise, their generation is always accompanied by noise and artifacts such as electrooculography (EOG) and electrocardiogram (ECG). Furthermore, the amount of EEG data related to emotion recognition is very limited, and feature loss or insufficient learning often occurs during network learning. Therefore, researchers have conducted extensive experimental studies to design models with high recognition accuracy and robustness. Thus, extracting high-quality, relatively abundant emotion recognition features from EEG is crucial. However, experimental studies have revealed the following: 1. Long network training time and high computational costs; 2. Inability to extract rich features from high-dimensional EEG signals; 3. The inherent sparsity of EEG-based emotion recognition data means that some features cannot be extracted after the data is input into the network, resulting in training effects that do not meet expectations. Therefore, reducing external interference and enabling neural networks to learn more and richer EEG features remains a challenge. Summary of the Invention

[0003] The technical problem to be solved by this invention is to provide a network emotion recognition method based on CCNN and stacked-BiLSTM to solve the above-mentioned problems, thereby enabling the extraction of richer and more comprehensive feature information in human-computer interaction systems.

[0004] The technical solution of this invention is: a network emotion recognition method based on CCNN (multiple parallel continuous convolutional neural network) and stacked-BiLSTM (stacked bidirectional long short-term memory), the specific steps of which are as follows:

[0005] Step 1: Collect raw EEG signals and preprocess them.

[0006] The preprocessing specifically involves using a Butterworth bandpass filter to decode four EEG bands: theta, alpha, beta, and gamma, and then using a 0.5s non-overlapping sliding window to increase the number of samples.

[0007] In this experiment, 63 seconds of data from each participant were decoded using a Butterworth bandpass filter to obtain the EEG bands of four frequency bands related to emotion recognition: theta, alpha, beta, and gamma. At the same time, a 0.5-second non-overlapping sliding window was used to increase the number of samples, allowing the network to learn fully.

[0008] Step 2: Obtain the frequency domain feature differential entropy (DE) on the four frequency bands. Use the moving average method to reduce the brain electrical fluctuations caused by noise, artifacts and other factors during the emotion recognition process. Calculate the average differential entropy features of the baseline of the first 3 seconds of each of the four frequency bands. Then, map the frequency domain feature differential entropy features of each frequency band after feature smoothing onto a brain planar topographic map related to emotion recognition to obtain its spatial frequency features.

[0009] Step 3: Input the acquired spatial frequency features into a multi-parallel continuous convolutional neural network to fully learn more and higher semantic spatial frequency features. Since EEG is a dynamic temporal signal, there may be some hidden information in the time process that is useful for emotion classification. Therefore, stacked bidirectional long short-term memory network (stacked-BiLSTM) is used to enable it to fully learn the feature information of the past and future in different time slices, thereby achieving accurate decoding of emotion recognition information based on EEG.

[0010] Step 2 specifically includes:

[0011] The frequency domain feature differential entropy feature of each frequency band is obtained by formula (1):

[0012]

[0013] Where p(x) represents the probability density function of continuous information, and [a,b] represents the information value interval, which approximately follows a Gaussian distribution for a specific length. EEG.

[0014] Then, the average differential entropy characteristics of the baselines in the first 3 seconds of the four frequency bands are calculated. The average differential entropy of the baselines in the first 3 seconds of the corresponding frequency band is subtracted from each 0.5-second sliding window data using the moving average method, as shown in formula (3):

[0015]

[0016]

[0017] In the formula, k represents the portion of the sliding window each time under musical stimulation, j represents the four frequency bands, and i represents the portion of the sliding window each time under the baseline signal.

[0018] To minimize the impact of noise, artifacts, and other factors on the acquired brainwaves, this invention utilizes the DEAP database, which collects 32-lead brainwave signals. Since brainwaves are high-dimensional signals with rich spatial information, this invention employs an 8x9 brain topographic map associated with emotion recognition to capture this spatial information.

[0019] Step 3 specifically refers to:

[0020] Spatial frequency features are incorporated into a multi-parallel, continuous convolutional neural network architecture. Each convolutional neural network contains three different types of convolutional layers: 64 kernels of size 5x5, 128 kernels of size 4x4, and 256 kernels of size 4x4. These three types of convolutional layers are combined to form four such convolutional blocks. To prevent the data distribution from slowing down convergence as the network depth increases, which tends to approach the upper and lower limits of a non-linear data distribution, a batch normalization (BN) layer is added after each convolutional layer. This places the input values ​​of each neural network into a standard normal distribution with a mean of 0 and a variance of 1. This results in a smaller output while still yielding a larger gradient, thus accelerating training. After four convolutional blocks, a 64-kernel 1x1 two-dimensional convolutional layer is used to interact with channel information and reduce the network parameters. Then, a 2x2 max pooling layer is used for feature compression, a Flatten layer is used to flatten the multi-dimensional input, and a Dense layer is placed at the end of the convolutional neural network for refitting, minimizing the loss of feature information.

[0021] The spatial frequency features output from multiple parallel continuous neural networks are concatenated using the concatenate function. Utilizing the LSTM principles in formulas (4)-(9), these features are then fed into a stacked bidirectional long short-term memory network. This allows for more comprehensive learning of information from past and future time slices. Specifically:

[0022] f t =σ(W f ·[h t-1 ,x t ]+b f (4)

[0023] i t =σ(W i ·[h t-1 ,x t ]+b i (5)

[0024]

[0025]

[0026] o t =σ(W o [h t-1 ,x t ]+b o (8)

[0027] h t =o t *tanh(C t (9)

[0028] In the formula, f t h represents the value of the forget gate. t-1 Given the hidden state of x at the previous time step, t Let i be the input value at the current moment. t Indicates the value of the memory gate. Indicates a temporary cell state, C t C represents the current state of the cell. t-1 Indicates the cell state at the previous moment, o t h represents the value of the output gate. t This represents the hidden state.

[0029] Finally, classify and predict its classification:

[0030]

[0031] In the formula, Z i Let be the output value of the i-th node, and n be the number of output nodes, i.e., the number of categories. The softmax function can be used to convert the output values ​​of multi-class classification into a probability distribution ranging from [0,1] to 1.

[0032] The beneficial effects of this invention are as follows: Compared with the prior art, this invention utilizes differential entropy features and brain topography to obtain rich spatial frequency features as input, and then employs multi-parallel continuous convolutional neural networks-stacked bidirectional long short-term memory to acquire richer and more comprehensive feature information, thereby achieving high accuracy in recognition and classification tasks, and providing a new approach for better emotion recognition. Attached Figure Description

[0033] Figure 1 This is a diagram of the recognition neural network framework in an embodiment of the present invention;

[0034] Figure 2 This is a planar topographic map of the brain in an embodiment of the present invention;

[0035] Figure 3 This is a space-frequency diagram of the four frequency bands in this embodiment of the invention;

[0036] Figure 4 This is a diagram of the convolutional block structure in an embodiment of the present invention. Detailed Implementation

[0037] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0038] Example 1: As Figure 1-4 As shown, a network-based emotion recognition method based on CCNN and stacked-BiLSTM is described, with the following specific steps:

[0039] Step 1: Acquire raw EEG signals and preprocess them;

[0040] The preprocessing specifically involves: using a Butterworth bandpass filter to decode four EEG bands: theta, alpha, beta, and gamma; and then using a 0.5s non-overlapping sliding window to increase the number of samples.

[0041] Step 2: Obtain the frequency domain feature differential entropy on the four frequency bands, use the moving average method to reduce the EEG fluctuations that occur during emotion recognition, calculate the average differential entropy features of the baseline of the first 3 seconds of the four frequency bands, and then map the frequency domain feature differential entropy features of each frequency band after feature smoothing to a brain planar topography map related to emotion recognition to obtain its spatial frequency features.

[0042] Step 3: Input the obtained spatial frequency features into a multi-parallel continuous convolutional neural network to learn higher semantic spatial frequency features, and then use a stacked bidirectional long short-term memory network (stacked-BiLSTM) to learn past and future feature information on different time slices to achieve accurate decoding of EEG-based emotion recognition information.

[0043] 2. The network emotion recognition method based on CCNN and stacked-BiLSTM according to claim 1, wherein Step 2 specifically comprises:

[0044] The frequency domain feature differential entropy feature of each frequency band is obtained by formula (1):

[0045]

[0046] Where p(x) represents the probability density function of continuous information, and [a,b] represents the information value interval, which approximately follows a Gaussian distribution for a specific length. EEG;

[0047] Then, the average differential entropy characteristics of the baselines in the first 3 seconds of the four frequency bands are calculated. The average differential entropy of the baselines in the first 3 seconds of the corresponding frequency band is subtracted from each 0.5-second sliding window data using the moving average method, as shown in formula (3):

[0048]

[0049]

[0050] In the formula, k represents the portion of the sliding window each time under musical stimulation, j represents the four frequency bands, and i represents the portion of the sliding window each time under the baseline signal.

[0051] 3. The network emotion recognition method based on CCNN and stacked-BiLSTM according to claim 1, wherein Step 3 specifically comprises:

[0052] The spatial frequency features output from multiple parallel continuous neural networks are concatenated together using the concatenate function. Then, using formulas (4)-(9), the spatial frequency features are fed into a stacked bidirectional long short-term memory network, specifically:

[0053] f t =σ(W f ·[h t-1 ,x t ]+b f (4)

[0054] i t =σ(W i ·[h t-1 ,x t ]+b i (5)

[0055]

[0056]

[0057] o t =σ(W o [h t-1 ,x t ]+b o (8)

[0058] h t =o t *tanh(C t (9)

[0059] In the formula, f t h represents the value of the forget gate. t-1 Given the hidden state of x at the previous time step, t Let i be the input value at the current moment.t Indicates the value of the memory gate. Indicates a temporary cell state, C t C represents the current state of the cell. t-1 Indicates the cell state at the previous moment, o t h represents the value of the output gate. t Indicates the hidden state;

[0060] Finally, classify and predict its classification:

[0061]

[0062] In the formula, Z i Let be the output value of the i-th node, and n be the number of output nodes, i.e., the number of categories. The softmax function can be used to convert the output values ​​of multi-class classification into a probability distribution ranging from [0,1] to 1.

[0063] The following experiments and evaluations will be conducted:

[0064] 1. Experimental data and hyperparameter settings

[0065] The DEAP public dataset was used. The experimental paradigm involved 32 healthy participants (16 males and 16 females). Each participant was asked to watch a one-minute video. Each participant's data consisted of 40*40*8064 physiological signal data points (40 music videos, 40 electrode channels, 8064 sampling points), labeled 40*4 (40 channels, each containing four dimensions: valence, arousal, dominance, and liking). After watching, participants scored on valence, arousal, dominance, and liking. A total of 40 experiments were conducted, collecting information from 40 channels. The first 32 channels were EEG channels, and the last 8 were other channels. The sampling rate was 512 Hz. In this experiment, we used the official dataset that had been downsampled (downsampled to 128 Hz) and had noise removed, including from EEG signals. The valence and arousal dimensions were tested. The experimental environment consisted of Python 3.7, an NVIDIA GeForce GTX 1660Ti GPU, and an 11th Gen Intel(R) Core(TM) i5-11400F CPU. The entire neural network was implemented using the TensorFlow framework.

[0066] The dataset uses 5x cross-validation, is trained for 100 epochs, is divided into 6 parts, and is trained simultaneously using 6 parallel continuous convolutional neural networks with a learning rate of 0.001.

[0067] 2. Comparative Analysis of Experimental Results

[0068] To verify the advantages of this invention, comparisons were conducted on different models. Experiments demonstrate that this invention is of significant value.

[0069]

[0070] Table 1: Accuracy of Classification Models

[0071] As shown in Table 1, compared to traditional machine learning models such as SVM and ANN, the accuracy of the model in this embodiment is improved by about 20%, which proves that deep learning models can extract more detailed emotional features and spatial structure features. This invention also investigated some advanced deep learning models, such as Bi-LSTM and CNN-LSTM. Comparing the single Bi-LSTM with the related fusion model CNN-LSTM, the model of this invention shows a fundamental improvement. The hybrid model of this invention not only fully extracts spatial frequency features but also extracts the dynamic temporal features of EEG.

[0072] This invention utilizes a multi-parallel branch structure to reduce computational costs and training time. A scalp map of the brain maps spatial information structures, followed by the use of a continuous convolutional neural network to acquire more comprehensive and semantically rich features. A stacked bidirectional long short-term memory network learns information from past and future time slices. Experimental results show that the proposed multi-parallel continuous convolutional neural network and stacked bidirectional long short-term memory network, tested on the DEAP public database, achieve an average recognition accuracy of 94.8% and 95% in valence and arousal dimensions, respectively. While reducing training time, it also learns relatively comprehensive feature information. This not only avoids the loss of some features and acquires more complete and comprehensive spatial-frequency feature information, but also reduces network runtime even with a large amount of sample data.

[0073] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A network emotion recognition method based on CCNN and stacked-BiLSTM, characterized in that: Step 1: Acquire raw EEG signals and preprocess them; The preprocessing specifically involves: using a Butterworth bandpass filter to decode four EEG bands: theta, alpha, beta, and gamma; and then using a 0.5s non-overlapping sliding window to increase the number of samples. Step 2: Obtain the frequency domain feature differential entropy on the four frequency bands, use the moving average method to reduce the EEG fluctuations that occur during emotion recognition, calculate the average differential entropy features of the baseline of the first 3 seconds of the four frequency bands, and then map the frequency domain feature differential entropy features of each frequency band after feature smoothing to a brain planar topography map related to emotion recognition to obtain its spatial frequency features. Step 3: Input the obtained spatial frequency features into a multi-parallel continuous convolutional neural network to learn higher semantic spatial frequency features, and then use a stacked bidirectional long short-term memory network to learn past and future feature information on different time slices to achieve accurate decoding of EEG-based emotion recognition information. Step 3 specifically refers to: The spatial frequency features output from multiple parallel continuous neural networks are concatenated together using the concatenate function. Then, using formulas (4)-(9), the spatial frequency features are placed into a stacked bidirectional long short-term memory network, specifically: (4); (5); (6); (7); (8); (9); In the formula, Indicates the value of the forget gate. This represents the hidden state at the previous moment. This is the input value at the current moment. Indicates the value of the memory gate. Indicates a temporary cell state. This indicates the current state of the cell. This indicates the cell state at the previous moment. Indicates the value of the output gate. Indicates the hidden state; Finally, classify and predict its classification: (10); In the formula, For the first The output value of each node To output the number of nodes, i.e. the number of categories, we use... The function can convert the output values ​​of multi-class classification into a range of... The probability distribution with a sum of 1.

2. The network emotion recognition method based on CCNN and stacked-BiLSTM according to claim 1, characterized in that, Step 2 specifically includes: The frequency domain feature differential entropy feature of each frequency band is obtained by formula (1): (1); in, f ( x ) represents the probability density function of continuous information. The information represents a range of values, which for a given length approximately follows a Gaussian distribution. EEG; Then, the average differential entropy characteristics of the baselines in the first 3 seconds of the four frequency bands are calculated. The average differential entropy of the baselines in the first 3 seconds of the corresponding frequency band is subtracted from each 0.5-second sliding window data using the moving average method (Formula (3)). (2); (3); In the formula, This represents the portion of the window that slides each time it is stimulated by music. Indicates 4 frequency bands, This represents the portion of the baseline signal that is slid through each window.

Citation Information

Patent Citations

  • Epilepsy electroencephalogram signal identification method and system

    CN113288172A

  • Emotion analysis method and system for multi-view deep learning based on electroencephalogram signals

    CN114129163A