Emotion Recognition Method Based on the Fusion of EEG Spatiotemporal Frequency Features

Through the emotional recognition method based on the fusion of space-time frequency features of EEG, the problem of insufficient extraction of EEG signal features is solved. Time-frequency diagrams and dynamic brain functional network conversion are adopted, and BILSTM combined with channel attention mechanism BILSTM is improved.

CN119279611BActive Publication Date: 2025-07-08NORTHEAST DIANLI UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311536434.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-07-08
Estimated Expiration
2043-11-17

AI Technical Summary

Technical Problem

现有技术难以充分提取脑电信号的多维特征,导致情感识别准确性和鲁棒性不足。

Method used

The emotion recognition method based on EEG space-time frequency feature fusion is adopted. After baseline correction and standardization, time-frequency graph and dynamic brain function network conversion are performed. Combined with PLI theory and Hamming window function, the time-frequency and spatial characteristics of the EEG signal are extracted, and BILSTM with added channel attention mechanism is used for deep characteristics fusion.

Benefits of technology

The accuracy and robustness of emotion recognition are improved, especially in the arousal and valence dimensions, the highest recognition accuracy is achieved, verifying the effectiveness of the method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119279611B_ABST
    Figure CN119279611B_ABST
Patent Text Reader

Abstract

The present invention discloses an emotion recognition method based on the fusion of EEG spatio-temporal-frequency features, including: S1, performing a series of preprocessing steps on the original EEG signals, including baseline correction, normalization, and adding time windows; S2, performing a form conversion on the EEG signals processed in S1 above to obtain the spatial information and time-frequency information of the EEG signals; S3, extracting features in multiple aspects; S4, fusing the extracted emotion features and performing multi-feature emotion recognition; S5, using BILSTM with added channel attention mechanism to further extract the time features of the EEG signals and perform deep fusion of the features; S6, obtaining the emotion classification result. By adopting the above emotion recognition method based on the fusion of EEG spatio-temporal-frequency features, the present invention fully considers the feature information of the three dimensions of time, frequency, and space of the EEG signals, achieves the highest recognition accuracy in the arousal and valence dimensions, improves the accuracy and robustness of emotion recognition, and verifies the effectiveness of the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of emotion recognition, and in particular to an emotion recognition method based on the fusion of EEG spatio-temporal-frequency features. Background Art

[0002] As a comprehensive state generated by people under external stimulus conditions, emotion recognition plays an important role in the process of expressing and perceiving information in daily life. Given the accuracy and non-invasiveness of electroencephalogram (EEG) signals, more and more research focuses on using EEG for emotion recognition. In recent years, the rapid development of Internet technology has promoted the arrival of the artificial intelligence era, and deep learning technology has also flourished. Although various deep learning frameworks have been applied to the field of emotion recognition, how to fully extract the multi-dimensional features of EEG signals and design a model that can not only reduce information loss but also improve the accuracy and robustness of emotion recognition is still challenging. Summary of the Invention

[0003] The object of the present invention is to provide an emotion recognition method based on the fusion of EEG spatio-temporal-frequency features, which fully considers the feature information of the three dimensions of time, frequency, and space of EEG signals, achieves the highest recognition accuracy in the arousal and valence dimensions, improves the accuracy and robustness of emotion recognition, and verifies the effectiveness of the present method.

[0004] To achieve the above object, the present invention provides an emotion recognition method based on the fusion of EEG spatio-temporal-frequency features, including:

[0005] S1. A series of preprocessing steps are performed on the original EEG signal, including baseline correction, normalization, and adding a time window;

[0006] S2. The EEG signal processed in S1 above is subjected to a form conversion to obtain the spatial information and time-frequency information of the EEG signal;

[0007] S3. Multi-faceted feature extraction is performed;

[0008] S4. The extracted emotion features are fused, and multi-feature emotion recognition is performed;

[0009] S5. A BILSTM with a channel attention mechanism added is used to further extract the time features of the EEG signal and perform deep fusion of the features;

[0010] S6. The emotion classification result is obtained.

[0011] Preferably, in the signal form conversion in step S2:

[0012] S21. Construct a dynamic brain functional network based on the PLI theory, capture the spatial features of signals by extracting the topological features of the brain network, segment the EEG signals using a sliding window method, and calculate the phase synchrony between nodes using PLI to generate a relationship matrix reflecting the connection strength between nodes. Its original expression is:

[0013]

[0014] where E<·> represents the expectation operation on the data within the time window; N represents the time points; represents the phase difference between two signals at time t n ; sign is a sign function; the value of PLI ranges from 0 to 1. If the PLI is closer to 1, it indicates a stronger connection strength, and vice versa, it indicates a weaker connection strength;

[0015] The relationship matrix constructed based on the PLI theory needs to be further thresholded. An adaptive threshold method is used to threshold the relationship matrix for each time window. According to the characteristics of the data, the adaptive threshold is expressed as:

[0016] threshold=mean_pli-std_pli / 5 (2)

[0017] where mean_pli represents the mean of the relationship matrix, and std_pli represents the variance of the relationship matrix;

[0018] S22. Use the short-time Fourier transform of the Hamming window function to convert the EEG signal into a time-frequency diagram. Its expression is:

[0019] S(n,ω)=|X(n,ω)| 2 (3)

[0020]

[0021] where S(n,ω) represents the time-frequency diagram; X(n,ω) is the short-time Fourier transform, which is a two-dimensional function that relates time and frequency; x(m) is the input signal; w(n - m) is the window function, and a Hanning window with a window length of 1 s and an overlap of 0.75 s is used.

[0022] Preferably, in the feature extraction in step S3:

[0023] S31. Based on the construction of the dynamic brain functional network, eight graph theory features of the brain functional network are extracted, including clustering coefficient, shortest path length, global efficiency, local efficiency, betweenness centrality, closeness, degree centrality, and eigenvector centrality;

[0024] S32. Based on the ResNet-18 model, fine-tune it with the time-frequency graph as the input to extract deep time-frequency features;

[0025] S33. Extract the power spectral density and differential entropy features of the EEG data within each time window.

[0026] Preferably, in step S4, using the method of feature splicing, combine the time-frequency-spatial domain features of the collected EEG signals, splice the last dimension of the feature data in each domain to obtain the information of the fused multi-feature data, and use the features in different domains to represent the feature information of the EEG signal. The calculation is as follows:

[0027] F = concatenate(f deep , f brain , f de , f psd ) (5)

[0028] Among them, F is the fused feature; f deep is the deep time-frequency feature extracted by fine-tuning the pre-trained ResNet-18; f brain is the spatial domain feature extracted by constructing a dynamic brain functional network; f de is the extracted differential entropy; f psd is the extracted power spectral density;

[0029] Preferably, in step S5, use a bidirectional long short-term memory network structure with a channel attention mechanism to further extract the time features of the EEG signal and perform deep fusion of the features. The expression of the channel attention mechanism is as follows:

[0030] M c (F) = softmax(MLP(AvgPool(F)) + MLP(MaxPool(F))) (6)

[0031] Among them, softmax is the activation function used to map the output probability; AvgPool represents average pooling; MaxPool represents max pooling; M c (F) ∈ R C×1 , represents the attention weight; C represents the number of channels;

[0032] Preferably, control the flow of information through a bidirectional long short-term memory network structure, where:

[0033] (1) The forget gate is used to control the update of the current time cell state; using the short-term memory output h at the previous moment t-1 and the sequence data x to be input at the current moment t as the input, determine the probability f of the long-term memory to be forgotten through the activation function t, the calculation process is as shown in Equation 7:

[0034] f (t) = σ(W f · [h (t-1) , x (t) + b f ) (7)

[0035] where σ is the sigmoid activation function used to map the calculation result to the interval [0, 1]; W f is the weight matrix; b f is the bias;

[0036] (2) The input gate is used to determine which information in the current time input should be added to the cell state; it consists of two parts. The first part is used to calculate the probability of retaining information, and the calculation process equation is as follows;

[0037] i t = σ(W i · [h t-1 , x t + b i ) (8)

[0038] The second part is used to calculate the output value at the current moment, and the calculation process equation is as follows:

[0039]

[0040] where i t represents the probability value of retaining information; represents the output value at the current moment; σ is the sigmoid activation function used to map the calculation result to the interval [0, 1]; W is the weight matrix; b is the bias; then represents the information to be updated at the current moment;

[0041] (3) The output gate is used to control the amount of short-term memory information output at the current time; the new information obtained by using the input gate and the forget gate is updated to the cell state C t ; the cell state C t and the output gate jointly determine the output h t of this layer, and the calculation process is as follows:

[0042]

[0043] o t = σ(W o · [h t-1 , x t + b o ) (11)

[0044] h t = ot *tanh(C t )。 (12)

[0045] Therefore, the present invention adopts the above-mentioned emotion recognition method based on EEG spatio-temporal frequency feature fusion, fully considers the feature information of the three dimensions of time, frequency, and space of EEG signals, achieves the highest recognition accuracy in the arousal and valence dimensions, improves the accuracy and robustness of emotion recognition, and verifies the effectiveness of this method.

[0046] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0047] Figure 1 is a framework diagram of an emotion recognition method based on EEG spatio-temporal frequency feature fusion; among them, (a) is a 32-channel EEG signal; (b) converts the EEG signal into a time-frequency diagram and a dynamic brain functional network respectively; (c) extracts DE features and PSD features of the EEG signal; fine-tunes the pre-trained ResNet-18 to extract the deep time-frequency features of the EEG signal; manually extracts the spatial features of the EEG signal based on the dynamic brain functional network; (d) simply fuses the extracted emotion features; (e) uses BILSTM with a channel attention mechanism to further extract the temporal features of the EEG signal and perform deep fusion of the features; (f) obtains the emotion classification result;

[0048] Figure 2 is an architecture diagram of the fine-tuned pre-trained ResNet-18 model of the emotion recognition method based on EEG spatio-temporal frequency feature fusion;

[0049] Figure 3 is an architecture diagram of the BILSTM of the emotion recognition method based on EEG spatio-temporal frequency feature fusion;

[0050] Figure 4 is a structural diagram of the LSTM unit of the emotion recognition method based on EEG spatio-temporal frequency feature fusion;

[0051] Figure 5 is a confusion matrix diagram of the arousal dimension of the emotion recognition method based on EEG spatio-temporal frequency feature fusion;

[0052] Figure 6 is a confusion matrix diagram of the valence dimension of the emotion recognition method based on EEG spatio-temporal frequency feature fusion;

[0053] Figure 7 is the single-subject emotion recognition accuracy of the arousal and valence dimensions of the emotion recognition method based on EEG spatio-temporal frequency feature fusion;

[0054] Figure 8Are the recognition accuracies of five models of the emotion recognition method based on the fusion of EEG spatio-temporal frequency features in the arousal dimension;

[0055] Figure 9 Are the recognition accuracies of five models of the emotion recognition method based on the fusion of EEG spatio-temporal frequency features in the valence dimension. Specific implementation manners

[0056] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0057] As Figure 1 shown, first, a series of preprocessing steps are performed on the original EEG signals, including baseline correction, normalization, and adding time windows. Next, two key data conversion stages are carried out to obtain the time-frequency information and spatial information of the EEG signals. In the first stage, the EEG signals are converted into time-frequency maps. In order to better process the time-frequency maps and extract relevant time-frequency features, the architecture of the pre-trained ResNet-18 model is adjusted and successfully migrated to the research. The second stage involves constructing a dynamic brain functional network based on the PLI theory, and capturing the spatial features of the signals by extracting the topological features of the brain network. In addition, the power spectral density and differential entropy features are also extracted to supplement the comprehensive feature information. Considering that the EEG signals are continuously collected according to each electrode channel and have a time sequence, in order to better capture the time information of the EEG signals of each channel, we introduce the BILSTM model that performs well on sequence data. At the same time, we also introduce a channel attention mechanism, as part of the classification model, for predicting the final emotion results. The implementation of this comprehensive method aims to improve the emotion recognition performance for small sample data sets.

[0058] Embodiment 1

[0059] S1. Preprocess the original EEG signals;

[0060] (1) Baseline processing: Aiming at the problem that the baseline data of the first 3 s of the signal data in the DEAP data set may affect the experimental accuracy, the present invention uses the method of baseline averaging to process the baseline data.

[0061] First, the baseline data of the first 3 s of the EEG signals of each subject are taken out and divided into three 1-s signal segments. The baseline average value is obtained by averaging the signal values of the three segments using formula (13).

[0062]

[0063] Then, the 60s EEG signals after the baseline time of each subject were divided into 60 signal segments of 1s, and the signal values after baseline processing were obtained by subtracting the baseline average value from the signal value of each segment. The calculation relationship is shown in Equation (14).

[0064] data_re i = raw_data i - b_mean (14)

[0065] Finally, the 60 signal segments after baseline processing were concatenated to obtain the complete signal data. The calculation relationship is shown in Equation (15).

[0066] data = concat(data_re i ) (15)

[0067] where b1, b2, b3 respectively represent the 1s baseline signals, b_mean ∈ R E×S , raw_data ∈ R E×S , data_re i ∈ R E×S , data ∈ R E×S×N , E represents the number of electrode channels, S represents the length of the 1s signal segment, N represents 60 signal segments, i = 1, 2, 3..., 60, and concat represents the concatenation operation. After baseline processing, the signal data form of each subject is (40, 7680), where 40 represents the number of electrode channels and 7680 represents the data length.

[0068] Among them, the summary of the DEAP dataset is shown in Table 1.

[0069] Table 1 Detailed information of the DEAP dataset

[0070]

[0071] (2) Sliding time window setting and sample division: Human emotions usually change within a time period of 0.5 - 4 seconds. Selecting an appropriate time is crucial for emotion recognition. Through a large number of experiments, it is proved that setting the sliding time window with a window length of 3s and a stride size of 1s can achieve the best recognition performance. Therefore, in the present invention, the window size is set to 3s (including 384 sampling points) and the stride size is set to 2s.

[0072] According to the setting of the time window, each electrode channel signal is divided into 58 time windows. Therefore, based on the characteristic that each subject in the DEAP dataset conducts 40 experiments, each subject can collect approximately 2380 samples. The data of each sample is represented as 32×384. Since only the face video data of 18 subjects in the dataset is complete, 18 subjects with complete video data are selected for the emotion recognition experiment in this study.

[0073] According to the data processing method, the arousal and valence labels are classified. The label threshold is set to 5. When the experimental score is greater than or equal to 5, the label is classified as high arousal / valence; otherwise, the label is classified as low arousal / valence. After differentiating the categories, classification is performed for each time window.

[0074] (3) Data normalization processing: Without changing the distribution information of the original data, in order to eliminate the dimensional difference between EEG signals and thus improve the accuracy of emotion recognition. The present invention uses the min-max normalization method to normalize the EEG signals within each time window, mapping the data uniformly to the range [0, 1]. This process can be described as:

[0075]

[0076] where data represents each electrode data in each time window, min_data represents the minimum data point of each electrode in each time window, and max_data represents the maximum data point of each electrode in each time window.

[0077] S2. Signal form conversion

[0078] S21. Construction of dynamic brain functional network

[0079] The brain is a dynamic system for complex information interaction in different emotional states. By analyzing the brain network, the collaborative working mode between different brain regions can be studied, providing a new way to deeply understand the principle of brain operation. By recording electroencephalogram (EEG) signals, constructing a dynamic brain functional network, and analyzing network characteristics to capture the spatio-temporal relationship between multiple brain regions, more comprehensive and multi-dimensional brain spatial information can be obtained, which is helpful for emotion recognition. When constructing the brain functional network, the selection of nodes and the construction of edges are very important. Usually, scalp electrodes are used as network nodes, and network edges are defined according to the correlation of scalp electrode time series. The phase lag index (PLI) is a phase-based functional connectivity analysis method mainly used to study the functional connectivity between multi-channel EEG signals. In the present invention, a sliding window method is used to segment the EEG signals, and the PLI is used to calculate the phase synchrony between nodes, generating a relationship matrix reflecting the connection strength of nodes. Its original expression is:

[0080]

[0081] Among them, E<·> represents the expectation operation on the data within the time window; N represents the time point; represents the phase difference between two signals at time t n ; sign is a sign function. The value of PLI ranges between 0 and 1. If the PLI is closer to 1, it indicates a stronger connection strength; conversely, it indicates a weaker connection strength.

[0082] The relationship matrix constructed based on the PLI theory needs further thresholding. To avoid information loss to the greatest extent, the present invention selects an adaptive threshold method to perform thresholding on the relationship matrix of each time window. According to the characteristics of the data, the adaptive threshold is expressed as:

[0083] threshold=mean_pli - std_pli / 5 (2)

[0084] where mean_pli represents the mean of the relationship matrix, and std_pli represents the variance of the relationship matrix. After selecting the appropriate threshold, we set the values in the relationship matrix that are less than the threshold to 0, thus obtaining the final brain functional network.

[0085] S22. Time-frequency diagram conversion

[0086] Research has found that different types of spectrograms can capture different characteristics of signals, and these features can be automatically extracted through deep learning methods and generate the required target presentation. Among them, the time-frequency diagram performs excellently in describing the time-frequency changes of signals. Therefore, the short-time Fourier transform using the Hamming window function is used to convert the EEG signal into a time-frequency diagram. Its expression is:

[0087] S(n,ω)=|X(n,ω)| 2 (3)

[0088]

[0089] where S(n,ω) represents the time-frequency diagram; X(n,ω) is the short-time Fourier transform, which is a two-dimensional function that relates time and frequency; x(m) is the input signal; w(n - m) is the window function, and here we use a Hanning window with a window length of 1s and an overlap of 0.75s.

[0090] S3. Multi-faceted feature extraction

[0091] The main objective of feature extraction is to distill the most valuable information from the data, aiming to enhance the efficiency of data processing and improve the model performance. In this invention, feature extraction is carried out in three aspects: First, spatial feature extraction based on the brain functional network; Second, deep time-frequency feature extraction using the time-frequency diagram as the input; Finally, to supplement the comprehensive feature information, power spectrum and differential entropy feature extraction are performed.

[0092] S31. Based on the construction of the dynamic brain functional network in step S21, eight graph theory features of the brain functional network are extracted, including clustering coefficient, shortest path length, global efficiency, local efficiency, betweenness centrality, closeness, degree centrality, and eigenvector centrality. By extracting these graph theory features, detailed information about the structure and characteristics of the brain functional network can be obtained, thus providing more comprehensive brain spatial information for emotion recognition.

[0093] S32. During the entire emotion recognition process, in order to achieve the optimal emotion recognition performance, the present invention performs parameter settings on the proposed model. The pre-trained ResNet-18 model used for extracting deep features is fine-tuned. To prevent information loss to the greatest extent, the 7×7 downsampling convolution is replaced with a 3×3 downsampling convolution, and the stride and padding size of this convolution layer are reduced. At the same time, the max pooling layer is removed. As shown in Table 2, the detailed information of the ResNet-18 model architecture.

[0094] Table 2 ResNet-18 model architecture

[0095]

[0096] ResNet-18 is selected as the model because of its simple network structure, fewer parameters, and low computational complexity, which is particularly suitable for relatively small datasets. In addition, since ResNet-18 has learned rich deep features during pre-training, it is fine-tuned with the time-frequency diagram as the input to extract deep time-frequency features, which can effectively accelerate the model training process and improve the accuracy. The advantage of this strategy is that it not only fully utilizes the image features learned by ResNet-18 during pre-training but also can further adapt to our dataset, capture the correlation between the time domain and the frequency domain, thereby improving the recognition ability of time-frequency features. Through this method, the feature extraction potential of ResNet-18 can be maximally exerted, and the model performance can be further optimized. As Figure 2 shown, the architecture of the fine-tuned pre-trained ResNet-18 model.

[0097] S33. Existing research has shown that power spectral density (PSD) and differential entropy (DE) are widely used in emotion recognition tasks. Differential entropy (DE) is used to measure the uncertainty or information content of continuous random variables; power spectral density (PSD) is used to analyze the frequency components and energy distribution of signals, and it provides information about the energy distribution of signals in different frequency ranges. To supplement the comprehensive feature information, this paper extracts the power spectral density and differential entropy features from the EEG data within each time window to capture the frequency-domain features related to emotions.

[0098] S4. Fuse the extracted emotion features and perform multi-feature emotion recognition;

[0099] The sufficiency of features has an important impact on the performance of emotion recognition. To fully consider the information in each domain and improve the emotion recognition performance, the present invention adopts a feature splicing method to combine the time-frequency-spatial domain features of the collected EEG signals. Specifically, the last dimension of the feature data in each domain is spliced to obtain the information of the fused multi-feature data. In this way, the features in different domains can be comprehensively utilized to more comprehensively represent the feature information of the EEG signals, further improving the accuracy and performance of emotion recognition. The calculation is as follows:

[0100] F = concatenate(f deep , f brain , f de , f psd ) (5)

[0101] where F is the fused feature; f deep is the deep time-frequency feature extracted by fine-tuning the pre-trained ResNet-18; f brain is the spatial domain feature extracted by constructing a dynamic brain functional network; f de is the extracted differential entropy; f psd is the extracted power spectral density.

[0102] S5. Use a BILSTM with a channel attention mechanism to further extract the time features of the EEG signals and perform deep feature fusion;

[0103] Since emotion expression is related to specific brain regions, the channel information collected in the brain regions related to emotion expression should be of higher importance. However, simply splicing this information does not highlight the more important channel information. The channel attention mechanism can effectively compress the spatial information of multi-channel EEG signals and generate statistical information about the channels by paying attention to the importance of different channels.

[0104] To achieve the best classification accuracy, multiple experiments were conducted to adjust the architecture and parameters of the classification model BILSTM. As shown in Table 3, the detailed information of the model architecture.

[0105] Table 3 BILSTM Model Structure

[0106]

[0107] The LSTM layer is used to model and extract features from the input data. To improve the expressive power and robustness of the feature representation, a channel attention layer is used to increase the weights of important features. The weighted data is linearly transformed through a fully connected layer to learn the complex non-linear relationships in the features. Then, a batch normalization layer is used to make each feature dimension have a similar distribution, thereby accelerating the training convergence and preventing the problems of gradient vanishing or explosion. Next, the output of the batch normalization layer is non-linearly mapped using the ReLU activation function to increase the expressive power of the model, and a part of the neurons are randomly discarded with a dropout probability to prevent the model from overfitting. Finally, a fully connected layer and the ReLU activation function are used to generate the final prediction results.

[0108] Based on the entire model architecture, as shown in Table 4, the detailed information of the hyperparameters is given:

[0109] (1) The number of nodes in BILSTM is 128 * 3; (2) In the Dropout layer, a part of the neurons are discarded with a probability of dropout = 0.4; (3) The model is trained using the Adam optimizer with an initial learning rate of 0.001, and then the learning rate is reduced by 50% every 20 epochs until the learning rate reaches 1e-6; (4) An L2 regularization term with a weight of 0.0001 is added to prevent overfitting; (5) The batch size is set to 32; (6) The number of training epochs is set to 200 epochs.

[0110] Table 4 BILSTM Model Hyperparameters

[0111]

[0112] Therefore, in the present invention, a channel attention mechanism is introduced. By adaptively adjusting the weights of each channel, it is possible to better focus on the channels with important information, thereby improving the recognition performance of the model. The expression of the channel attention mechanism is as follows:

[0113] M c (F) = softmax(MLP(AvgPool(F)) + MLP(MaxPool(F))) (6)

[0114] Among them, softmax is an activation function used to map output probabilities; AvgPool represents average pooling; MaxPool represents max pooling; M c (F) ∈ R C×1 , representing the attention weight; C represents the number of channels.

[0115] The bidirectional long short-term memory network (BILSTM) represents an improved recurrent neural network structure. As Figure 3 shown, the architecture of BILSTM integrates two LSTM layers, a forward LSTM layer and a backward LSTM layer, and each LSTM layer contains multiple LSTM units. In the forward LSTM, the input sequence is gradually input in chronological order, and the hidden state at each time step is passed to the next time step; while in the backward LSTM, the input sequence is input in the reverse chronological order, and the hidden state at each time step is passed forward. Finally, the outputs of the forward and backward LSTMs are fused together to form the final output result. Through this bidirectional structure, the network can consider both past and future information simultaneously, thereby improving the ability to understand each moment in the sequence. The core idea of this architecture is to use a series of structures called "gates" to control the flow of information. As Figure 4 shown, the structure of an LSTM unit includes three key gates: the forget gate, the input gate, and the output gate. The functions of these three gates are specifically explained below:

[0116] (1) The forget gate is used to control the update of the cell state at the current time. It takes the short-term memory output h t-1 at the previous moment and the sequence data x t to be input at the current moment as inputs, and determines the probability f t of the long-term memory to be forgotten through the activation function. The calculation process is as shown in Equation 7:

[0117] f (t) = σ(W f · [h (t-1) , x (t) + b f ) (7)

[0118] Among them, σ is the sigmoid activation function used to map the calculation result to the interval [0, 1]; W f is the weight matrix; b f is the bias.

[0119] (2) The input gate is used to determine which information in the current input should be added to the cell state. It consists of two parts. The first part is used to calculate the probability of retaining information, and the calculation process is as follows;

[0120] i t = σ(W i · [h t-1, x t + b i ) (8)

[0121] The second part is used to calculate the output value at the current moment, and the calculation process equation is as follows:

[0122]

[0123] where i t represents the probability value of the retained information; represents the output value at the current moment; σ is the sigmoid activation function used to map the calculation result to the interval [0, 1]; W is the weight matrix; b is the bias. Then represents the information to be updated at the current moment.

[0124] (3) The output gate is used to control the amount of information of the short-term memory output at the current time. The new information obtained by using the input gate and the forget gate is updated to the cell state C at the current moment t . The cell state C t and the output gate jointly determine the output h of this layer t . The calculation process is as follows:

[0125]

[0126] o t = σ(W o · [h t-1 , x t + b o ) (11)

[0127] h t = o t * tanh(C t ) (12)

[0128] S6. Obtain the emotion classification result;

[0129] The electroencephalogram signal is a time-varying sequence. In order to further capture the time-domain characteristics of the signal to achieve better recognition performance, the present invention uses the fused features as the input and uses a bidirectional long short-term memory neural network containing 128 * 3 nodes to obtain the context time information of the electroencephalogram signal and obtain the final emotion recognition result.

[0130] To evaluate the performance of the emotion recognition method proposed in the present invention, the dataset is divided into a training set, a validation set, and a test set according to a specific ratio. The samples in the dataset are randomized and the test set is obtained by dividing them at a ratio of 20% to evaluate the generalization ability of the model. For the remaining data, five-fold cross-validation is used to divide the training set and the validation set, which are used for model training and model tuning respectively. The summary of the DEAP dataset is shown in Table 1

[0131] To verify the effectiveness and feasibility of the proposed BRPD-SE-BILSTM electroencephalogram (EEG) emotion recognition method, on the premise of using the same dataset, the method of the present invention is compared with the methods of related research. The experimental results of these methods are given in their original texts, as shown in Table 5, and the results of the comparative experiment are given. The emotion recognition method proposed in the present invention fully considers the feature information of the three dimensions of time, frequency, and space of EEG signals, and achieves the highest recognition accuracies of 97.01% and 93.92% on the arousal and valence dimensions, respectively, which are 1.97% and 0.03% higher than the optimal recognition method in terms of accuracy, verifying the effectiveness of the proposed method.

[0132] Table 5 Results of different emotion recognition methods

[0133]

[0134] To evaluate the performance of the model used in the present invention, precision, recall rate, and F1 score are used as the evaluation criteria for the model, and the evaluation criteria are as follows:

[0135] Precision: Measures the ability of the classification model to predict true classes in the results.

[0136]

[0137] Recall rate: Measures the ability of the model to correctly capture positive class samples.

[0138]

[0139] F1 score: Measures the balance of the model on positive and negative class samples.

[0140] Among them, TP is true positive, FN is false negative, and FP is false positive. As shown in Table 6, the three performance indicators of precision, recall rate, and F1 score of the model in the arousal and valence dimensions. It can be seen from the table that the model of the present invention shows good performance indicators of precision, recall rate, and F1 score in the arousal and valence dimensions, indicating that the model has good ability to perform emotion classification based on EEG data.

[0141] Table 6 Results of performance indicators in the arousal and valence dimensions

[0142]

[0143] Although the three evaluation metrics of precision, recall, and F1-score can better understand and evaluate the performance of the model in different tasks, their limitation is that they only provide the performance evaluation of the model at a certain point and do not provide more detailed class information. Therefore, a confusion matrix is drawn to show the comparison between the prediction results of the classifier and the actual classes, so as to provide the detailed performance of the classifier on different classes. Figure 5 The confusion matrix representing the arousal dimension; Figure 6 The confusion matrix representing the valence dimension. The horizontal and vertical coordinates of the confusion matrix represent the predicted labels and the true labels respectively, and the values on the main diagonal represent the number of samples predicted correctly. From the numbers of true positives, false positives, true negatives, and false negatives in the confusion matrix, it can be clearly seen that the model of the present invention shows excellent performance in the arousal and valence dimensions.

[0144] To better understand the emotion classification ability of each subject, evaluate the stability and consistency of the model on different subjects, and further improve the generalization ability of the model, an emotion recognition experiment was conducted separately for each subject, and the recognition accuracies of all subjects in the arousal and valence dimensions were plotted as bar charts, as Figure 7 shown. It can be seen from the figure that Subject 7, Subject 8, and Subject 15 achieved the highest recognition accuracy of 100% in the valence dimension and the arousal dimension respectively. In addition, the recognition accuracies of Subject 6, Subject 9, Subject 17, and Subject 21 in the valence dimension were 88.65%, 89.15%, 88.54%, and 88.29% respectively, and the recognition accuracy of Subject 17 in the arousal dimension was 89.71%. The recognition accuracies of the remaining subjects in the arousal and valence dimensions were all higher than 90%, which fully reflects the generalization ability and stability of our model.

[0145] Example 2

[0146] To solve the problem of insufficient feature extraction, this paper proposes a method that comprehensively uses deep learning models to automatically extract features and manual extraction of features, and has achieved remarkable results. To verify the important impact of feature sufficiency on the effect of emotion recognition, a control experiment was conducted. Tables 7 and 8 show the experimental results on the feature problem, comparing the classification accuracies of common features (such as DE and PSD) for emotion recognition based on electroencephalogram with the fused depth-time-frequency features (DEEP) and spatial domain features (Brain) proposed in this invention. It can be seen from the table that the accuracy of any combined features is higher than that of emotion recognition using only the spatial domain features extracted from the brain functional network and the depth-time-frequency features extracted from the deep neural network model. The average accuracies of the fused depth-time-frequency feature DEEP and spatial domain feature Brain in the arousal and valence dimensions are 86.94% and 88.41% respectively. At the same time, the deep-time-frequency feature DEEP, spatial domain feature Brain, manually extracted frequency domain features PSD and DE in the combined features show the best classification results in the arousal and valence dimensions, with average accuracies of 97.01% and 93.92% respectively. These classification results fully illustrate that the sufficiency of feature extraction has a serious impact on the accuracy of emotion classification.

[0147] Table 7 Results of the combination of brain functional network features and deep features with manually extracted frequency domain features in the arousal dimension

[0148]

[0149] Table 8 Results of the combination of brain functional network features and deep features with manually extracted frequency domain features in the valence dimension

[0150]

[0151] To further discuss the impact of different model architectures in the feature extraction stage and classification stage on the overall research, this invention constructs a set of comparison models, details are as follows:

[0152] Model 1: The pre-trained ResNet-18 model in the deep feature extraction process was adjusted. The convolutional kernel, stride, and padding of the first convolutional layer were adjusted from 3×3, 1, and 1 to 7×7, 2, and 3 respectively, and a max pooling layer was added, and the rest of the model architecture remained unchanged. The purpose of doing this is to compare the impact of information loss on the experimental results.

[0153] Model 2: Remove the channel attention mechanism in the BILSTM model. The purpose of doing this is to compare the impact of the attention mechanism on the experimental results in the classification stage.

[0154] Considering the influence of the number of layers of the LSTM network on the experimental results, two models were constructed respectively to compare the differences between using a 1-layer and a 2-layer LSTM network and the 3-layer LSTM network used in the present invention.

[0155] Model 3: It includes a 2-layer LSTM network and a channel attention module.

[0156] Model 4: It includes a 1-layer LSTM network and a channel attention module.

[0157] The above are the model combinations constructed in the present invention for comparison. The model of the present invention was respectively compared with the 4 constructed comparison models in terms of performance in the arousal and valence dimensions. As Figure 8 and Figure 9 shown, the recognition accuracies of the five models in the arousal and valence dimensions are respectively shown. It can be seen that our model shows the highest recognition accuracy in both the arousal and valence dimensions. Compared with Model 1, the accuracies of our model in the arousal and valence dimensions are respectively increased by 5.15% and 4.7%. This is mainly because Model 1 adopted a larger convolution kernel and max pooling operation, which may lead to information loss and thus reduce the accuracy of emotion recognition. In the emotion recognition task, the electroencephalogram (EEG) information of different EEG channels was collected, and it was found through constructing a brain network that different EEG channels have different degrees of influence on emotions. Therefore, channel attention is used to assign different weights to the EEG channels. Compared with Model 2, the accuracies of the model of the present invention in the arousal and valence dimensions are respectively increased by 3.43% and 2.87%. This shows that the attention mechanism has an important influence on the emotion recognition performance. In the comparison with Model 3 and Model 4, the number of layers of the LSTM network was mainly concerned, and it was found that as the number of network layers increases, the emotion recognition performance shows a gradually increasing trend. In the arousal dimension, our model is respectively increased by 0.03% and 6.49% compared with Model 3 and Model 4; in the valence dimension, the model of the present invention is respectively increased by 1.47% and 2.36% compared with Model 3 and Model 4. The experimental results show that different numbers of layers of the LSTM network consider the context time information to different degrees. And the model of the present invention fully considers the context time information of the EEG signals and achieves the optimal emotion recognition performance in the arousal and valence dimensions.

[0158] Therefore, the present invention adopts the above-mentioned emotion recognition method based on the fusion of EEG spatio-temporal frequency features, solves the problem of insufficient extraction of EEG signal features in the emotion recognition task, designs a model that reduces information loss, improves accuracy and robustness under small sample data, fully considers the feature information of the three dimensions of time, frequency and space of the EEG signal, achieves the highest recognition accuracy in the arousal and valence dimensions, and verifies the effectiveness of this method.

[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. An emotion recognition method based on the fusion of EEG spatio-temporal frequency features, characterized in that: Including: S1. A series of preprocessing steps are performed on the original EEG signals, including baseline correction, normalization, and adding a time window; S2. The EEG signals processed in S1 above are subjected to form conversion to obtain the spatial information and time-frequency information of the EEG signals; S3. Feature extraction is carried out in multiple aspects; S4. The extracted emotion features are fused; Using the method of feature splicing, the time-frequency-spatial domain features of the collected EEG signals are combined, and the last dimension of the feature data in each domain is spliced to obtain the information of the fused multi-feature data. The feature information of the EEG signals is characterized by using the features in different domains, and the calculation is as follows: (5); Among them, is the fused feature; is the deep time-frequency feature extracted by fine-tuning the pre-trained ResNet-18; is the spatial domain feature extracted by constructing a dynamic brain functional network; is the extracted differential entropy; is the extracted power spectral density; S5. The BILSTM with channel attention mechanism added is used to further extract the time features of the EEG signals and perform deep fusion of the features; S6. Obtain the emotion classification result.

2. The emotional recognition method based on EEG spatio-temporal frequency feature fusion according to claim 1, wherein In the signal form conversion in step S2: S21. Based on the PLI theory, a dynamic brain functional network is constructed. The spatial features of the signals are captured by extracting the topological features of the brain network. The sliding window method is used to segment the EEG signals, and the PLI is used to calculate the phase synchrony between nodes, generating a relationship matrix reflecting the connection strength between nodes. Its original expression is: (1); Among them, represents the expectation operation on the data within the time window; represents the time point; represents the phase difference between two signals at time ; is a sign function; The value of is between 0 and 1. If is closer to 1, it means the stronger connection strength. On the contrary, it means the weaker connection strength; The relationship matrix constructed based on the PLI theory needs further thresholding processing. The adaptive threshold method is used to threshold the relationship matrix of each time window. According to the characteristics of the data, the expression of the adaptive threshold is: (2); Among them, represents the mean of the relationship matrix, represents the variance of the relationship matrix; S22. The short-time Fourier transform using the Hamming window function is used to convert the EEG signals into time-frequency diagrams, and its expression is: (3); (4); Among them, represents a time-frequency diagram; is the short-time Fourier transform, which is a two-dimensional function that relates time and frequency; is the input signal; is the window function, and a Hamming window with a window length of 1 s and an overlap of 0.75 s is used.

3. The emotional recognition method based on EEG spatio-temporal frequency feature fusion according to claim 1, wherein, In the feature extraction in step S3: S31. Based on the construction of the dynamic brain functional network, 8 graph theory features of the brain functional network are extracted, including clustering coefficient, shortest path length, global efficiency, local efficiency, betweenness centrality, closeness, degree centrality, and eigenvector centrality; S32. Based on the ResNet-18 model, it is fine-tuned with the time-frequency diagram as the input to extract deep time-frequency features; S33. The power spectral density and differential entropy features of the EEG data within each time window are extracted.

Citation Information

Patent Citations

  • Multi-modal data multi-view sleep staging method

    CN116070168A

  • Driver road accident emotion recognition method based on electroencephalogram signals in simulation environment

    CN116942181A