Hybrid Neural Network EEG Emotion Recognition Method and System Based on Multidimensional Features

By adopting a hybrid neural network method with multi-dimensional features in EEG emotional recognition, combined with channel-by-channel convolution, dot product and BiLSTM network, the problem of long recognition time, underutilization of frequency dimensions and lack of front-end feature learning is solved, and more efficient and real-time EEG emotional recognition is achieved.

CN116467654BActive Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310405083.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2025-05-27
Estimated Expiration
2043-04-14

AI Technical Summary

Technical Problem

The prior art has problems in the recognition of EEG signals that the emotional information that is too long, the frequency dimension is not fully considered, and the lack of deep learning of the front and back characteristics.

Method used

Using a hybrid neural network method based on multi-dimensional features, by calculating the difference between the signal and the baseline signal under the subject's emotional stimulation, the band features are not overlapped, divided into bands, divided into frames, and extracted, and converted into a 4D matrix sequence. Then, a residual network is constructed using channel-by-channel convolution and dot product, fused the frequency domain channel attention network, and finally input the convolution smoothing signal to the BiLSTM network for dynamic time feature learning.

Benefits of technology

It improves the accuracy and real-time nature of EEG emotional recognition, makes full use of the three-dimensional features of time, frequency and space, reduces feature redundancy, and enhances the complexity and recognition rate of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116467654B_ABST
    Figure CN116467654B_ABST
Patent Text Reader

Abstract

The present invention claims protection for a hybrid neural network emotion recognition method and system based on multi-dimensional features, constructs 4D features containing time-frequency-space information as the input of the recognition model, retains the spatial information between electrodes, the time correlation of the EEG sequence, and makes full use of various sub-band information corresponding to different emotions. Then, a residual network based on depthwise separable convolution is proposed, which not only extracts the spatio-frequency features in the input signal but also further reduces the training parameters, and applies the Fca-Block to suppress the information irrelevant to emotion recognition in the features. Finally, we use Bi-LSTM to learn the time information in the samples bidirectionally, and in order to highlight the time importance of the frame window in the samples, the weighted sum of the hidden layer states at all frame moments is selected as the input of Softmax. Compared with the methods in recent years, the proposed model still achieves remarkable results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of EEG emotion recognition and is a hybrid neural network emotion recognition method based on multi-dimensional features. Background Art

[0002] Emotion is a psychological phenomenon that integrates a person's feelings, thoughts, and behaviors. It is a person's psychological reaction to external or self-stimulation, and also includes the physiological reaction that accompanies this psychological reaction. With the development of science and technology, people's lives are becoming more and more closely related to intelligent technology, and accurate identification of emotions is becoming increasingly important in the field of human-computer interaction. Research in neurophysiology and social psychology has proved that EEG signals are highly correlated with many cognitive processes, including emotional processes, and can objectively reflect the true emotions of the subjects, which has become a hot research direction.

[0003] Due to the nonlinearity, non-smoothness, low signal-to-noise ratio and multi-channel correlation of EEG signals, it is essentially a complex chaotic data. How to extract features that are highly related to emotions and choose a suitable classification model have always been two stumbling blocks that hinder researchers. It has been proven that EEG signals contain feature components related to emotions in the three dimensions of time, frequency and space. However, most literature only considers one or two of these three dimensions and performs pattern recognition through simple combinations, completely ignoring the spatial information interaction between channels, the prior knowledge between frequency bands and the information complementarity between different features. This not only has no effect on improving the accuracy of the model, but also leads to feature redundancy and increased complexity. Therefore, by integrating information from different fields, adaptively capturing important time, frequency and space features in subsequent classification models is of great significance for improving the emotion recognition rate.

[0004] In recent years, deep learning, as a new development direction of machine learning, has achieved fruitful results in the fields of CV, NLP, etc., and many deep models have been developed. Many researchers have successfully introduced deep learning methods into the field of EEG-based emotion recognition, such as GNN (graph neural network), GCN (graph convolutional neural network), Transformer, etc. While excellent results have been obtained, challenges still exist. How to build a network suitable for EEG signal feature extraction is a problem that needs to be solved urgently.

[0005] This patent integrates multiple electrophysiological signals with large dimensions, and the recognition time is too long, making it difficult to meet real-time requirements; moreover, this patent does not fully consider the emotional information contained in the frequency dimension of the electrical signal; this patent uses LSTM to process the input sequence at the frame level, which can greatly improve the recognition, but does not have a deep understanding of the before and after time relationship in the sequence, and lacks the learning of the before and after features. The present invention uses a BiLSTM network to conduct in-depth learning of past and future dynamic time features. Summary of the invention

[0006] The present invention aims to solve the above problems of the prior art. A hybrid neural network EEG emotion recognition method and system based on multi-dimensional features is proposed. The technical solution of the present invention is as follows:

[0007] A hybrid neural network EEG emotion recognition method based on multi-dimensional features, comprising the following steps:

[0008] S1, calculates the difference between the signal of the subject under emotional stimulation and the baseline signal to represent the emotional state data of the segment;

[0009] S2, perform non-overlapping windowing, frequency band division, frame division, frequency band feature extraction, matrix mapping on the signal after baseline removal, and convert the entire window segment into a 4D matrix sequence;

[0010] S3, by using DC (channel-by-channel convolution) and PC (dot product) to build a residual network and fuse the frequency domain channel attention network FcaNet to obtain the spatial-frequency information in the features;

[0011] S4, input the convolution smoothed signal into the BiLSTM bidirectional long short-term memory network, learn the dynamic time characteristics of the EEG time series, obtain the past and future key emotional information of the EEG signal, and assign weights to the hidden layer state of the memory at each frame time and sum them as the input of Softmax;

[0012] Step S5, network training is performed based on cross entropy function optimization and stochastic gradient descent SGD with back propagation.

[0013] Furthermore, the step S1 calculates the difference between the signal of the subject under emotional stimulation and the baseline signal to represent the emotional state data of the segment, and the specific steps are as follows:

[0014] The sampling frequency of the dataset is S, given the entire baseline segment Where M and N1 represent the number of electrodes and the number of sampling points of the segment respectively; R represents a data correspondence rule; First, X base Evenly divided into 1-second segments, we get represents the i-th baseline segment, Then, the mean of the baseline data was calculated The formula is as follows:

[0015]

[0016] Same as above method, follow The length standard divides the test data into L segments Subtract the baseline mean from each experimental data Get the baseline removal data, the formula is as follows:

[0017]

[0018] Finally, the L segments of baseline-removed data are spliced ​​into complete data.

[0019] Furthermore, the step S2 performs non-overlapping windowing, frequency band division, frame division, frequency band feature extraction, and matrix mapping on the signal after baseline removal, and converts the entire window segment into a 4D matrix sequence, specifically including:

[0020] After removing the baseline, the signal X trial.rmov Perform non-overlapping windowing with a time length of u seconds to obtain the divided data The signal is divided into four frequency bands: θ, α, β, and γ through FIR filter. Considering that human emotion changes are time-dynamic, the segmented signal The image is divided into frames of equal length of 0.5s; the DE (differential entropy) and PSD (power spectral density) features of all channels in each frame window are extracted; the single-band frame vector is mapped to a 2D matrix using sensitive transformation space mapping; from the perspective of a single frequency band, the window fragment will be transformed into a 2D matrix sequence, and the two feature matrices of different frequency bands are fused to obtain the final 4D data S = {S 1 ,S 2 ,...,S j}∈R 9x9x8x2u .

[0021] Furthermore, the single-band frame vector is mapped to a two-dimensional matrix by using sensitive transformation space mapping, specifically including: mapping 62 electrode channels to a 9×9 matrix according to the electrode placement of the "10 / 20" system, filling the corresponding DE (differential entropy) and PSD (power spectral density) eigenvalues ​​for the matrix elements with position mapping, and replacing the remaining positions with 0.

[0022] Furthermore, step S3 constructs a residual network by using DC (channel-by-channel convolution) and PC (dot product) and fuses the frequency domain channel attention network FcaNet to obtain the space-frequency information in the feature, specifically including:

[0023] First, the 2u frames S in each segment jThe data is sent to the CNN module in chronological order, and the 3D data structure of each frame is 9×9×8. When entering the convolutional coding layer, two convolutional layers are set up first. The convolution kernel of the first layer is 1×1, and the number of convolution kernels is 64. The second layer uses a 3×3 convolution kernel, and the number of convolution kernels is 128. Different convolution kernels are used to extract deep information in the 3D data. Then DC (channel-by-channel convolution) and PC (dot product) are used to construct the residual network. DC is used to extract the internal features of the expanded single feature map, and PC is used to express the relationship between feature maps. After each convolution layer, RELU is used as the activation function and BatchNorm processing is performed. Since the residual network needs to be cyclically operated and the output size needs to be kept unchanged, padding operation is added to DC (channel-by-channel convolution). Fca-Block will assign weights to different channel features. After the cycle, a 2×2 maximum pooling layer is used for dimensionality reduction, and then the data is transformed into one-dimensional data through the straightening layer. Finally, each frame of data is convolutionally encoded to obtain the vector S' j ∈R 1152 .

[0024] Furthermore, the S4 inputs the convolution smoothed signal into the BiLSTM bidirectional long short-term memory network to learn the dynamic time characteristics of the EEG time series and obtain the past and future key emotional information of the EEG signal, specifically including:

[0025] The convolution smoothed signal is input into the BiLSTM bidirectional long short-term memory network to learn the dynamic time characteristics of the EEG time series, obtain the past and future key emotional information of the EEG signal, and assign weights to the hidden layer state of the memory at each frame time and sum them as the input of Softmax.

[0026] Furthermore, the calculation formula of an LSTM unit is as follows:

[0027] f t =σ(W f ·[h t-1 ,x t ])+b f ),

[0028] i t =σ(W i ·[h t-1 ,x t ])+b i ),

[0029] C t =tanh(W C ·[h t-1 ,x t ])+b C ),

[0030] C t =f t ×C t-1 +i t ×C t ,

[0031] O t =σ(W O ·[h t-1 ,x t ])+b O ),

[0032] h t =tanh(C t )×O t ,

[0033] Among them, x t is the time series at time t, C t Represented as cell state, C t is the temporary cell state, σ is the sigmoid function, W is the weight matrix, b is the bias vector of the corresponding weight, and h t is the hidden state, f t For the forget gate, i t For the memory gate, O t The output gate is the forget gate; the forget gate selects the retained features, and inputs the information of the previous state and the current state into the sigmoid function at the same time. The memory gate is responsible for updating the state of the LSTM unit, and then the input gate controls the output value to the next LSTM unit. The output formula of the bidirectional LSTM is as follows:

[0034] y t =σ(W h ·[h t ,h′ t ])+b h )

[0035] First, perform a nonlinear transformation on the hidden layer state at each frame time. The formula is as follows:

[0036] H temp,j =tanh(W temp,j h j +b temp,j )

[0037] The number of memories in each LSTM layer is The hidden layer is processed in the form of connection to obtain h j ∈R q , H temp,j That is, the nonlinear expression of the hidden layer, W temp,j ∈R d×q and b temp,j ∈Rd is the weight and offset of the tanh function, and the value of d is set to 512; H temp,j ∈R d After that, the Softmax function is used to calculate the weight for each frame moment to obtain A temp,j , the specific formula is as follows:

[0038]

[0039] where u temp,j ∈R d is a trainable parameter, A temp,j The larger the value, the more important the corresponding frame is in the time series. Multiply all frame data by the weight and sum them up. The formula is as follows:

[0040]

[0041] Finally, using Z temp The prediction results are obtained through the softmax classifier.

[0042] Furthermore, the step S5 performs network training based on cross entropy function optimization and back-propagation stochastic gradient descent SGD, specifically including:

[0043] In the training network, dropout=0.5 is used, and two convolution layers are set first. The convolution kernel of the first layer is 1×1, and the number of convolution kernels is 64. The second layer uses a 3×3 convolution kernel, and the number of convolution kernels is 128. Different convolution kernels are used to extract deep information in 3D data, and then DC (channel-by-channel convolution) and PC (dot product) are used to construct the residual network. Padding operation is added to DC, and 256 units are selected.

[0044] A hybrid neural network EEG emotion recognition system based on multi-dimensional features, comprising:

[0045] Calculation module: used to calculate the difference between the signal of the subject under emotional stimulation and the baseline signal, so as to represent the emotional state data of the segment;

[0046] Preprocessing module: used to perform non-overlapping windowing, frequency band division, frame division, frequency band feature extraction, matrix mapping on the signal after baseline removal, and convert the entire window segment into a 4D matrix sequence;

[0047] Fusion module: used to construct a residual network by using DC (channel-by-channel convolution) and PC (dot product) and fuse the frequency domain channel attention network FcaNet to obtain the spatial-frequency information in the features;

[0048] Training module: used to input the convolution smoothed signal into the BiLSTM bidirectional long short-term memory network, learn the dynamic time characteristics of the EEG time series, obtain the past and future key emotional information of the EEG signal, and assign weights to the memory hidden layer state at each frame time and sum them as the input of Softmax; network training is performed based on cross entropy function optimization and stochastic gradient descent SGD with back propagation.

[0049] The advantages and beneficial effects of the present invention are as follows:

[0050] The present invention provides a hybrid neural network based on time-frequency-space three-dimensional features. Under the same experimental conditions, the proposed recognition model can improve the problem of low EEG emotion recognition rate. The difference between the signal of the subject under emotional stimulation and the baseline signal is calculated to represent the emotional state data of the segment. In order to retain the spatial information between electrodes, the temporal correlation of EEG sequences, and make full use of various sub-band information corresponding to different emotions, this paper constructs a 4D feature containing time-frequency-space information as the input of the recognition model. Specifically, the signal after baseline removal is subjected to non-overlapping windowing, frequency band division, frame division, frequency band feature extraction (DE, PSD), and matrix mapping, and the entire window segment is converted into a 4D matrix sequence. Then, a residual network based on deep separable convolution is proposed. In addition to extracting the space-frequency features in the input signal, it also further reduces the training parameters, and applies Fca-Block to suppress information in the features that is not related to emotion recognition. Finally, we use Bi-LSTM to learn the time information in the sample bidirectionally. In order to highlight the time importance of the frame window in the sample, the weighted sum of the hidden layer states at all frame moments is selected as the input of Softmax.

[0051] The present invention uses the difference data between the stimulated condition and the baseline average state as the signal of EEG emotion for feature extraction, and the extracted features can better reflect the emotional state; DC (channel-by-channel convolution) and PC (dot product) are used to construct the residual network, which not only meets the real-time performance of the system, but also avoids overfitting; the weighted sum of the hidden layer states at all frame moments is selected as the input of Softmax, which can better highlight the information of important frame moments. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is an overall block diagram of a hybrid neural network EEG emotion recognition method based on multi-dimensional features according to a preferred embodiment of the present invention;

[0053] Figure 2 It is a flow chart of feature extraction including time-frequency-space three-dimensional information of the present invention;

[0054] Figure 3 It is a space-frequency feature extraction module;

[0055] Figure 4It is the temporal feature extraction module. DETAILED DESCRIPTION

[0056] The following will describe the technical solutions in the embodiments of the present invention in detail in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention.

[0057] The technical solution of the present invention to solve the above technical problems is:

[0058] A hybrid neural network EEG emotion recognition method based on multi-dimensional features comprises the following steps:

[0059] S1, calculates the difference between the signal of the subject under emotional stimulation and the baseline signal to represent the emotional state data of the segment;

[0060] S2, perform non-overlapping windowing, frequency band division, frame division, extract frequency band features (DE, PSD), and matrix mapping on the signal after baseline removal, and convert the entire window segment into a 4D matrix sequence;

[0061] S3, by using DC (channel-by-channel convolution) and PC (dot product) to build a residual network and fuse the frequency domain channel attention network FcaNet to obtain the spatial-frequency information in the features;

[0062] S4, inputs the convolution smoothed signal into the BiLSTM bidirectional long short-term memory network, learns the dynamic time characteristics of the EEG time series, obtains the past and future key emotional information of the EEG signal, and assigns weights to the hidden layer state of the memory at each frame time and sums them as the input of Softmax.

[0063] S5, construct model parameters and perform model training.

[0064] Furthermore, the step S1 calculates the difference between the signal of the subject under emotional stimulation and the baseline signal, thereby representing the emotional state data of the segment. The specific steps are as follows:

[0065] The sampling frequency of the dataset is S, given the entire baseline segment Where M and N1 represent the number of electrodes and the number of sampling points in the segment respectively. First, X base Evenly divided into 1-second segments, we get represents the i-th baseline segment, Then, the mean of the baseline data was calculated The formula is as follows:

[0066]

[0067] Same as above method, follow The length standard divides the test data into L segments Subtract the baseline mean from each experimental data Get the baseline removal data, the formula is as follows:

[0068]

[0069] Finally, the L segments of baseline-removed data are spliced ​​into complete data.

[0070] Furthermore, the step S2 performs non-overlapping windowing, frequency band division, frame division, frequency band feature extraction (DE, PSD), and matrix mapping on the signal after baseline removal, and converts the entire window segment into a 4D matrix sequence.

[0071] Specifically include:

[0072] After removing the baseline, the signal X trial.rmov Perform non-overlapping windowing with a time length of u seconds to obtain the divided data The FIR filter is used to divide the frequency into four frequency bands: θ, α, β, and γ. Considering that human emotion changes are time-dynamic, the segmented signal X i S Divide into 0.5s frames of equal length. Extract DE and PSD features of all channels in each frame window. Use sensitive transformation space mapping to map the single-band frame vector to a 2D matrix. From the perspective of a single frequency band, the window fragment will be transformed into a 2D matrix sequence, and the two feature matrices of different frequency bands will be fused to obtain the final 4D data S = {S 1 ,S 2 ,...,S j}∈R 9x9x8x2u .

[0073] Furthermore, the residual network is constructed by using DC and dot product PC and the frequency domain channel attention network FcaNet is integrated to obtain the spatial-frequency information in the features, including:

[0074] First, the 2u frames S in each segment jThe data is sent to the CNN module in chronological order, and the 3D data structure of each frame is 9×9×8. Entering the convolutional coding layer, since the 3D structure of the data is relatively small, in order to better retain the information in the data, we first set up two convolutional layers. The convolution kernel of the first layer is 1×1, and the number of convolution kernels is 64. The second layer uses a 3×3 convolution kernel, and the number of convolution kernels is 128. Different convolution kernels are used to extract deep information in 3D data. Then DC (channel-by-channel convolution) and PC (dot product) are used to construct a residual network. Compared with traditional convolution, this combination not only reduces the number of training parameters; but also can use DC to extract the internal features of a single feature map after expansion, and then use PC to express the relationship between feature maps; the existence of the residual structure effectively avoids the problem of network degradation. After each convolution layer, RELU is used as the activation function and BatchNorm processing is performed. Since the residual network needs to be cyclically operated, the output size needs to be kept unchanged, so the DC is Padding operation is added. Fca-Block will assign weights to different channel features to improve the accuracy of the classification model. After the loop is completed, a 2×2 maximum pooling layer is used for dimensionality reduction, and then the data is transformed into one-dimensional data through a straightening layer. Finally, each frame of data is convolutionally encoded to obtain a vector S' j ∈R 1152 .

[0075] Furthermore, the convolution smoothed signal is input into the BiLSTM bidirectional long short-term memory network to learn the dynamic time characteristics of the EEG time series, obtain the past and future key emotional information of the EEG signal, and assign weights to the memory hidden layer state at each frame time and sum it up as the input of Softmax. Specifically including:

[0076] The calculation formula of an LSTM unit is shown as follows:

[0077] f t =σ(W f ·[h t-1 ,x t ])+b f ),

[0078] i t =σ(W i ·[h t-1 ,x t ])+b i ),

[0079] C t =tanh(W C ·[h t-1 ,x t ])+b C ),

[0080] Ct =f t ×C t-1 +i t ×C t ,

[0081] O t =σ(W O ·[h t-1 ,x t ])+b O ),

[0082] h t =tanh(C t )×O t ,

[0083] Among them, x t is the time series at time t, C t Represented as cell state, C t is the temporary cell state, σ is the sigmoid function, W is the weight matrix, b is the bias vector of the corresponding weight, and h t is the hidden state, f t For the forget gate, i t For the memory gate, O t is the output gate; the forget gate selects the retained features, inputs the previous state information and the current state information into the sigmoid function at the same time, and the memory gate is responsible for updating the state of the LSTM unit, and then the input gate controls the output value to the next LSTM unit. Different from the above-mentioned unidirectional LSTM, the output formula of the bidirectional LSTM is as follows:

[0084] y t =σ(W h ·[h t ,h′ t ])+b h )

[0085] In order to explore the importance of different frame window time segments, this paper does not use the output of Bi-LSTM as the result of the entire temporal feature learning module. Instead, a nonlinear transformation is first performed on the hidden layer state at each frame moment, and the formula is as follows:

[0086] H temp,j =tanh(W temp,j h j +b temp,j )

[0087] The number of memories in each LSTM layer is The hidden layer is processed in the form of connection to obtain h j ∈R q , H temp,jThat is, the nonlinear expression of the hidden layer, W temp,j ∈R d×q and b temp,j ∈R d are the weights and biases of the tanh function. The value of d is set to 512. temp,j ∈R d After that, the Softmax function is used to calculate the weight for each frame moment to obtain A temp,j , the specific formula is as follows:

[0088]

[0089] where u temp,j ∈R d is a trainable parameter, A temp,j The larger the value, the more important the corresponding frame is in the time series. Multiply all frame data by the weight and sum them up. The formula is as follows:

[0090]

[0091] Finally, using Z temp The prediction results are obtained through the softmax classifier.

[0092] Further, in step S5, network training is performed based on stochastic gradient descent (SGD) of cross entropy function optimization and back propagation. The weight sharing of CNN usually causes different gradient changes in different layers. For this reason, a single-layer convolutional neural network is used and a smaller learning rate is used. Since the two-layer bidirectional LSTM has a large depth, a small number of iterations (epoch=50) is sufficient for convergence. In order to eliminate the influence of the overfitting problem, dropout=0.5 is used in the training network. Since the three-dimensional structure of the data is relatively small, in order to better retain the information in the data, we first set up two layers of convolution. The convolution kernel of the first layer is 1×1, and the number of convolution kernels is 64. The second layer uses a 3×3 convolution kernel, and the number of convolution kernels is 128. Different convolution kernels are used to extract deep information in 3D data. Then DC (channel-by-channel convolution) and PC (dot product) are used to construct a residual network. Since the residual network needs to be cyclically operated and the output size needs to be kept unchanged, a padding operation is added to DC. In this method, the bidirectional LSTM layer has 256 units (512 in total), and the cases where the hidden layer has 128 and 512 units are also studied, and the 256 units with the best effect are selected.

[0093] A hybrid neural network EEG emotion recognition system based on multi-dimensional features, comprising:

[0094] Calculation module: used to calculate the difference between the signal of the subject under emotional stimulation and the baseline signal, so as to represent the emotional state data of the segment;

[0095] Preprocessing module: used to perform non-overlapping windowing, frequency band division, frame division, frequency band feature extraction, matrix mapping on the signal after baseline removal, and convert the entire window segment into a 4D matrix sequence;

[0096] Fusion module: used to construct a residual network by using DC (channel-by-channel convolution) and PC (dot product) and fuse the frequency domain channel attention network FcaNet to obtain the spatial-frequency information in the features;

[0097] Training module: used to input the convolution smoothed signal into the BiLSTM bidirectional long short-term memory network, learn the dynamic time characteristics of the EEG time series, obtain the past and future key emotional information of the EEG signal, and assign weights to the memory hidden layer state at each frame time and sum them as the input of Softmax; network training is performed based on cross entropy function optimization and stochastic gradient descent SGD with back propagation.

[0098] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0099] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0100] The above embodiments should be understood to be only used to illustrate the present invention and not to limit the protection scope of the present invention. After reading the contents of the present invention, technicians can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A method for electroencephalogram (EEG) emotion recognition based on a hybrid neural network with multi-dimensional features, characterized in that, it includes the following steps: S1. Calculate the difference between the signal of the subject under emotional stimulation and the baseline signal, and use this to represent the emotional state data of this segment; S2. Perform non-overlapping windowing, frequency band division, frame division, extraction of frequency band features, and matrix mapping on the signal after removing the baseline, and convert the entire window segment into a 4D matrix sequence; S3. Construct a residual network and fuse the frequency domain channel attention network FcaNet by using DC per-channel convolution and PC dot product to obtain the spatio-frequency information in the features, specifically including: First, 2u frames S in each segment are sequentially fed into the CNN module in chronological order. The 3D data structure of each frame is 9×9×8. Entering the convolutional encoding layer, two convolutional layers are first set up. The convolutional kernel of the first layer is 1×1, and the number of convolutional kernels is 64. The second layer uses a 3×3 convolutional kernel, and the number of convolutional kernels is 128. Different convolutional kernels are used to extract deep information in the 3D data. Then, DC per-channel convolution and PC dot product are used to construct a residual network. The internal features of the expanded single feature map are extracted through DC, and the relationship between cross-feature maps is expressed using PC. After each convolutional layer, RELU is used as the activation function and BatchNorm processing is performed. Since the residual network requires cyclic operations and needs to keep the output size unchanged, padding operations are performed on the DC per-channel convolution. The Fca-Block will assign weights to different channel features. After the loop ends, a 2×2 max pooling layer is used for dimensionality reduction, and then the data is transformed into one-dimensional data through a flattening layer. Finally, each frame of data is encoded through convolution to obtain the vector S' j ; j ∈R 1152 ; S4. Input the convolution-smoothed signal into a BiLSTM (Bidirectional Long Short-Term Memory) network to learn the dynamic time features of the EEG time series, obtain the past and future key emotional information of the EEG signal, and sum the weighted memory hidden layer states at each frame moment as the input to Softmax; Step S5. Perform network training based on the cross-entropy function optimization and Stochastic Gradient Descent (SGD) with backpropagation.

2. The method for EEG emotion recognition based on a hybrid neural network with multi-dimensional features according to claim 1, characterized in that, in step S1, the difference between the signal of the subject under emotional stimulation and the baseline signal is calculated, and this is used to represent the emotional state data of this segment. The specific steps are as follows: The sampling frequency of the dataset is S, and the entire baseline segment is given where M and N1 represent the number of electrodes and the number of sampling points in this segment respectively; R represents a data correspondence rule; first, divide X base evenly into time periods of 1 second in length to obtain denoting the i-th baseline segment, Then, calculate the average value of the baseline data The formula is as follows: In the same way as the above method, according to the length standard of divide the experimental data into L segments Subtract the baseline average value from each segment of experimental data to obtain baseline-removed data, and the formula is as follows: Finally, splice the L segments of data after removing the baseline into complete data.

3. The method for EEG emotion recognition based on a hybrid neural network with multi-dimensional features according to claim 1, characterized in that, in step S2, non-overlapping windowing, frequency band division, frame division, extraction of frequency band features, and matrix mapping are performed on the signal after removing the baseline, and the entire window segment is converted into a 4D matrix sequence. Specifically, it includes: For the signal X after baseline removal trial.rmov Perform non-overlapping windowing with a time length of u seconds to obtain the segmented data Divide it into four frequency bands of θ, α, β, and γ through an FIR filter; considering that human emotional changes are time-varying, the segmented signal is segmented into equal-length frames of 0.5 s; extract the differential entropy of DE and the power spectral density feature of PSD for all channels within each frame window; use the sensitive transformation space mapping method to map the single-band frame vector to a 2D matrix; from the perspective of a single frequency band, the window segment will be transformed into a sequence of 2D matrices, and fuse the two feature matrices of different frequency bands to obtain the final 4D data S = {S 1 , S 2 ,..., S j} ∈ R 9x9x8x2u .

4. The method for EEG emotion recognition based on a hybrid neural network with multi-dimensional features according to claim 3, characterized in that, the method of using sensitive transformation space mapping to map the single-frequency band frame vector to a 2D matrix specifically includes: according to the electrode placement positions of the "10 / 20" system, map the 62 electrode channels to a 9×9 matrix, and fill the matrix elements with position mapping with the corresponding differential entropy (DE) and power spectral density (PSD) eigenvalue features, and replace the remaining positions with 0.

5. The method for EEG emotion recognition based on a hybrid neural network with multi-dimensional features according to claim 1, characterized in that, in S4, the convolution-smoothed signal is input into a BiLSTM network to learn the dynamic time features of the EEG time series, and obtain the past and future key emotional information of the EEG signal. Specifically, it includes: Input the convolution-smoothed signal into a BiLSTM network to learn the dynamic time features of the EEG time series, obtain the past and future key emotional information of the EEG signal, and sum the weighted memory hidden layer states at each frame moment as the input to Softmax.

6. The method for EEG emotion recognition based on a hybrid neural network with multi-dimensional features according to claim 5, characterized in that, the calculation formula of an LSTM cell is as shown: f t = σ(W f · [h t-1 , x t ) + b f ), i t = σ(W i · [h t-1 , x t ) + b i ), O t = σ(W O · [h t-1 , x t ) + b O ), h t = tanh(C t ) × O t , where x t is the time series at time t, C t represents the cell state, is the temporary cell state, σ is the sigmoid function, W is the weight matrix, b is the bias vector corresponding to the weight, h t is the hidden state, f t is the forget gate, i t is the input gate, and O t is the output gate; the forget gate selects the features to be retained, inputs the information of the previous state and the current state into the sigmoid function at the same time, the input gate is responsible for updating the state of the LSTM unit, and then the output gate controls the output value to the next LSTM unit. The output formula of the bidirectional LSTM is shown as follows: y t = σ(W h · [h t , h′ t ) + b h ) First, perform a non-linear transformation on the hidden layer state at each frame time, and the formula is as follows: H temp,j = tanh(W temp,j h j + b temp,j ) where the number of memories in each LSTM layer is The hidden layer is processed in a concatenated form to obtain h j ∈R q , H temp,j is the non-linear expression of the hidden layer, W temp,j ∈R d×q and b temp,j ∈R d are the weights and offsets of the tanh function, and the value of d is set to 512; obtaining H temp,j ∈R d After that, the Softmax function is used to calculate the weights for each frame moment to obtain A temp,j , and the specific formula is as follows: where u temp,j ∈R d is a trainable parameter. The larger the value of A temp,j , the more prominent the importance of the corresponding frame in the time series. After multiplying all frame data by the weights and summing them up, the formula is as follows: Finally, utilize Z temp The prediction result is obtained through the softmax classifier.

7. The method for electroencephalogram emotion recognition based on a hybrid neural network with multi-dimensional features according to claim 6, wherein, in step S5, network training is performed based on cross-entropy function optimization and stochastic gradient descent SGD with backpropagation, specifically including: In the training network, dropout = 0.5 is adopted, and two convolutional layers are first set; the convolutional kernel of the first layer is 1×1, and the number of convolutional kernels is 64; the second layer uses a 3×3 convolutional kernel, and the number of convolutional kernels is 128; different convolutional kernels are used to extract deep information in 3D data, and then DC channel-by-channel convolution and PC dot product are used to construct a residual network. For the DC plus Padding operation, 256 units are selected.

8. A system for electroencephalogram emotion recognition based on a hybrid neural network with multi-dimensional features, wherein, it includes: A calculation module: used to calculate the difference between the signal of the subject under emotional stimulation and the baseline signal, so as to represent the emotional state data of this segment; A preprocessing module: used to perform non-overlapping windowing, frequency band division, frame division, extraction of frequency band features, and matrix mapping on the signal after baseline removal, and convert the entire window segment into a 4D matrix sequence; A fusion module: used to construct a residual network by using DC channel-by-channel convolution and PC dot product and fuse the spatio-temporal frequency information in the features of the frequency domain channel attention network FcaNet; specifically including: First, 2u frames S in each segment are sequentially fed into the CNN module in chronological order. The 3D data structure of each frame is 9×9×8. Entering the convolutional encoding layer, two convolutional layers are first set. The convolutional kernel of the first layer is 1×1, and the number of convolutional kernels is 64. The second layer uses a 3×3 convolutional kernel, and the number of convolutional kernels is 128. Different convolutional kernels are used to extract deep information in the 3D data. Then, DC channel-by-channel convolution and PC dot product are used to construct a residual network. The internal features of the expanded single feature map are extracted through DC, and the relationship between cross-feature maps is expressed using PC. After each convolutional layer, RELU is used as the activation function and BatchNorm processing is performed. Since the residual network requires cyclic operations and needs to keep the output size unchanged, padding operations are performed on the DC channel-by-channel convolution. The Fca-Block will assign weights to different channel features. After the loop ends, a 2×2 max pooling layer is used for dimensionality reduction, and then the data is transformed into one-dimensional data through a flattening layer. Finally, each frame of data is encoded through convolution to obtain the vector S' j j ∈R 1152 ;​ A training module: used to input the convolution-smoothed signal into a BiLSTM bidirectional long short-term memory network, learn the dynamic time features of the EEG time series, obtain the past and future key emotional information of the EEG signal, and sum the weighted memory hidden layer states at each frame time as the input of Softmax; network training is performed based on cross-entropy function optimization and stochastic gradient descent SGD with backpropagation.