Unsupervised semantic perception domain adaptive EEG emotion recognition method, medium and equipment
Through the unsupervised semantic perception domain adaptation method, combined with the self-attention mechanism and GCN/BiLSTM, the problems of insufficient information capture in frequency domain and unconsidered topological structure dynamic characteristics are solved, and the accuracy and robustness of cross-domain emotion recognition are improved.
Patent Information
- Application Number
- CN202510472208.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-04
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
AI Technical Summary
Existing EEG emotion recognition methods are difficult to fully capture the frequency domain information of signals, and fail to effectively consider the topological structure between EEG leads and the dynamic characteristics of the brain over time.
Unsupervised semantic perception domain adaptation method is adopted to process EEG signals through denoising, downsampling and slicing, and dual-view frequency domain features and spatiotemporal features are extracted, combined with self-attention mechanism and GCN/BiLSTM for feature fusion, and a semantic alignment mechanism is introduced for cross-domain adaptation.
It improves the accuracy and robustness of emotion recognition, and can realize effective migration and generalization of different domains without relying on target domain tag information.
Smart Images

Figure CN120372402A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of emotion recognition, and particularly to an unsupervised semantic-aware domain adaptation EEG emotion recognition method, medium and device. Background Art
[0002] Emotions play a crucial role in human cognition and behavior. Emotions affect our cognition, decision-making, and interaction with others, influencing social interaction and mental health. Accurately recognizing emotions is crucial for both personal mental health and interpersonal communication. Traditional emotion recognition methods often rely on subjective feedback and are easily interfered by individual differences and environmental factors. Electroencephalogram (EEG), as a non-invasive technology, can capture brain activities in real time and provide an objective basis for emotion recognition.
[0003] EEG signals are weak physiological signals and are extremely vulnerable to interference from external environments, muscle movements, and other factors. Effective data preprocessing techniques are crucial for improving the accuracy and stability of emotion recognition. Currently, EEG feature extraction methods are mainly divided into two categories: one is the manual feature extraction method, and the other is the deep learning-based feature extraction method.
[0004] In recent years, graph convolutional neural networks in deep learning methods have been widely applied in many fields. GCN can learn the complex relationships between nodes, which brings new ideas to the research of multi-channel EEG emotion recognition. Some researchers have used graph models to represent multi-channel EEG features and dynamically learned the functional relationships between EEG channels through neural network training, thereby improving the performance of EEG feature extraction and verifying the effectiveness of the proposed method on public datasets. Other researchers have proposed a regularized graph neural network model to capture the local and global relationships between different EEG channels and verified the performance of the model on public data.
[0005] Although effective results have been achieved, there are still some limitations. Specifically:
[0006] (1) Most studies are based on a single frequency-domain feature extraction method and it is difficult to comprehensively capture the frequency-domain information of signals, which often leads to the model ignoring some important frequency components when recognizing and analyzing EEG signals.
[0007] (2) The topological structure between EEG leads and the dynamic characteristics of the brain changing over time are not considered simultaneously.
[0008] Therefore, an unsupervised semantic-aware domain adaptation EEG emotion recognition method, medium and device are proposed to solve the problems raised above. Summary of the Invention
[0009] The purpose of the present invention is to provide an unsupervised semantic-aware domain adaptation EEG emotion recognition method, medium and device to solve the problems in the current market proposed in the above background technology.
[0010] To achieve the above object, the present invention provides the following technical solutions:
[0011] An unsupervised semantic-aware domain adaptation EEG emotion recognition method, comprising the following steps:
[0012] Step 1: Denoise, downsample and slice the EEG data to split the continuous EEG signal into multiple time segments;
[0013] Among them, the data resampling rate is 128Hz, and an eye muscle electrogram, eye movement and power supply noise in the electroencephalogram data are removed using a 4 - 45Hz band-pass filter. The EEG data is sliced in seconds, and continuous T seconds are selected to form a data sample;
[0014] Step 2: Extract dual-view frequency domain features and spatio-temporal features from the preprocessed EEG data. First, calculate the differential entropy of five frequency bands, namely theta, slow_alpha, alpha, beta, and gamma, in each time segment to obtain a DE feature matrix. Use the Welch method to calculate the power spectral density of the five frequency bands in each time segment to obtain a PSD feature matrix. Based on the self-attention mechanism, fuse the DE feature matrix and the PSD feature matrix into a dual-view frequency domain feature matrix;
[0015] Step 3: For spatial feature extraction, input the DE feature matrix of each time segment into the GCN to obtain the spatial feature representation of each time segment; for temporal feature extraction, input the spatial feature representations of the T time segments output by the GCN into the BiLSTM to obtain the temporal feature representation of each time segment;
[0016] Step 4: Train a classifier, flatten and concatenate the output of the spatio-temporal feature extractor into a feature representation vector h, and output the classification result by calculating the probability distribution of the predicted class; train a discriminator, further encode the features output by the target feature extractor, and act as a classifier to distinguish whether the encoded features come from the source domain or the target domain;
[0017] Step 5: Use the cross-entropy loss as the classification loss function to measure the difference between the two probability distributions, thereby optimizing the feature extractor, and calculate the global domain alignment loss from the domain true label and the predicted label generated by the domain discriminator;
[0018] Step 6: Introduce a semantic alignment mechanism. During the training process, use the trained source domain feature extractor and classifier to output the pseudo-labels of the target domain samples. During the adversarial domain adaptation learning process, the accuracy of the pseudo-labels continuously improves with training.
[0019] As a further optimization scheme of the present invention, in Step 2, the specific steps for calculating the differential entropy of five frequency bands in each time segment to obtain the DE feature matrix are as follows:
[0020] First, extract the DE features of five frequency bands per second. The calculation formula of DE is as follows:
[0021]
[0022] where σ is the variance of x, and e is the Euler's constant;
[0023] After DE feature extraction, further transform the segmented EEG segments into a feature matrix M represented by DE segments i ∈R T ×N×K (i = 1, 2,..., n), where T represents the time window length, N represents the number of EEG channels, K is the number of frequency bands, and n is the total number of samples.
[0024] As a further optimization scheme of the present invention, in Step 2, the specific steps for calculating the power spectral density of five frequency bands in each time segment to obtain the PSD feature matrix are as follows:
[0025] First, divide the original signal x(n) into K sub-segments of length N, and apply a window function w(n) to each sub-segment to reduce spectral leakage. Then, apply the fast Fourier transform to each segment of the signal to obtain the spectral information. The calculation formula is as follows:
[0026]
[0027] where x k (n) represents the signal of the k-th sub-segment, w(n) is related to the Hanning window, f is the frequency variable, and X i (k) represents the spectral information obtained after FFT transformation;
[0028] When the frequency f is set to 128 Hz, the window size is set to 1 second, and the power spectral density of each sub-segment is calculated by taking the square of the modulus of the FFT result X k (f) of each window:
[0029]
[0030] where P k (f) is the power spectral density estimate within the k-th window, N is the length of each window, and Xk (f) is the FFT result of the k-th segment of the signal;
[0031] Finally, the PSDs of all sub-segments are averaged to obtain the PSD of the entire signal, and the calculation formula is as follows:
[0032]
[0033] where K represents the total number of sub-segments, and P(f) is the estimated value of the PSD.
[0034] As a further optimization scheme of the present invention, in step two, based on the self-attention mechanism, the specific steps of fusing the DE feature matrix and the PSD feature matrix into a dual-view frequency-domain feature matrix are as follows:
[0035] The formulas used are defined as follows:
[0036]
[0037] where Q, K, and V represent Query, Key, and Value respectively, and d k is the size of the last dimension of the query;
[0038] Regarding DE and PSD as the key vector, query vector, and value vector respectively, and realizing the interaction and fusion between them through scaled dot-product attention. Then, the results of the dual-view frequency-domain features calculated based on the self-attention mechanism are fused through the following formula to obtain the new feature F dual , and its feature fusion formula is as follows:
[0039] F dual (f1,f2) = Attention(f1,f2,f2) + Attention(f2,f1,f1)
[0040] where f1 represents the feature DE, and f2 represents the feature PSD.
[0041] As a further optimization scheme of the present invention, in step three, the specific steps when extracting spatial features are as follows:
[0042] First, model the EEG signals of each channel and the connection relationship between channels as a graph G = {V, ε, A}, where V is a feature set containing N nodes, ε is the set of edges, and A is the adjacency matrix, representing the tightness of the connection between nodes;
[0043] Based on the Pearson correlation coefficient, evaluate the linear correlation degree between two signals according to the time-domain amplitude of the signals, and its calculation formula is as follows:
[0044]
[0045] Among them, i and j represent two different nodes, and μ i and μ j represent the means of nodes i and j, and E is the expected value. σ i and σ j represent the standard deviations of i and j respectively;
[0046] In addition, according to the following formula, the threshold of the correlation coefficient is set to 0.5 to obtain the adjacency matrix A, and its calculation formula is as follows:
[0047]
[0048] As a further optimization scheme of the present invention, in step three, the specific steps for extracting time features are:
[0049] Use BiLSTM as the time feature extractor to receive the calculation results obtained by T GCNs. Each LSTM unit has an input gate, a forget gate, an output gate, and a memory unit inside, which are used to learn the long-term dependencies in the sequence data. The forget gate determines how much of the cell state information from the previous time step needs to be discarded at the current time step through a sigmoid layer. The input gate determines how much of the existing input information is added to the cell state. The cell state is updated according to the forget gate and the input gate, retaining the long-term dependence on past information. Finally, the output gate determines the hidden state output at this time step based on the current cell state, thereby maintaining sensitivity to time-related information, so as to better capture and characterize the important features on the long time scale in EEG data. Given the input sequence x = (x1, x2,..., x T ), its calculation process can be expressed as:
[0050] f t = σ(W f ·[h t-1 , xt] + b f )
[0051] i t = σ(W i ·[h t-1 , x t + b i )
[0052] o t = σ(W o ·[h t-1 , x t + b o )
[0053]
[0054] h t = o t⊙tanh(C t )
[0055] where f t , i t , o t represent the calculation results of the forget gate, input gate, and output gate respectively, C t is the cell state at time t, h t is the output at time t, x t is the input at the current time, and W and b are the weight matrix and bias term respectively.
[0056] As a further optimization scheme of the present invention, in step five, the discriminator consists of three fully connected layers and corresponding activation functions. The first two fully connected layers are configured with BN and ReLu layers, and the last fully connected layer and Sigmoid are combined as a classifier. The definition of Sigmoid is as follows:
[0057]
[0058] where x i is the output of the i-th neuron in the previous layer, and f sigmoid is the output probability of the i-th class, ranging from [0, 1].
[0059] The classification loss is used to update the feature extractor θ E and the classifier parameter θ C , and the global domain alignment loss is used to assist in the update of the target domain feature extractor and the discriminator θ D parameters. The formula principles of both the domain alignment loss and the classification loss are:
[0060]
[0061] where y and p(x) represent the true label and the predicted label respectively, and N represents the number of samples.
[0062] As a further optimization scheme of the present invention, in step six, semantic alignment is achieved by calculating the distance between the feature representations of the same category data in the source domain and the target domain. The formula definition is as follows:
[0063]
[0064] where K represents the total number of categories, and represent the feature centers of the k-th category in the source domain and the target domain respectively. φ represents the cosine similarity, which is used to calculate the distance between the corresponding category feature centers.
[0065] As a further optimization solution of the present invention, in Steps Five and Six, the classification loss, domain alignment loss, and semantic alignment loss constitute the total loss of the adversarial domain adaptation process, which is specifically defined as follows:
[0066] L T =L C +w d L D +w s L S
[0067] Wherein, L C 、L D and L S respectively represent the classification loss, domain alignment loss, and semantic alignment loss. w d and w s are the weight coefficients of the corresponding losses respectively.
[0068] A computer-readable storage medium, on which instructions are stored, and when the instructions are executed on a computer, the computer is caused to execute the unsupervised semantic-aware domain adaptation EEG emotion recognition method.
[0069] An electronic device, comprising: one or more processors; one or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device is caused to execute the unsupervised semantic-aware domain adaptation EEG emotion recognition method.
[0070] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0071] The unsupervised semantic-aware domain adaptation method proposed by the present invention can be used for cross-domain EEG emotion recognition, enabling the model to achieve effective transfer and generalization across different domains without relying on target domain label information. Based on the dual-perspective frequency-domain feature extraction mechanism of the self-attention mechanism, frequency-domain features are extracted from two perspectives of differential entropy and power spectral density, and the features are fused based on the self-attention module, thereby providing more comprehensive and richer frequency-domain information for the model. At the same time, the spatio-temporal feature extraction formed by GCN and BiLSTM can comprehensively consider the topological structure between EEG leads and the dynamic characteristics of the brain changing over time to extract complex topological structure and temporal dynamic features.
[0072] The present invention introduces a semantic alignment mechanism in the unsupervised domain adaptation framework, enabling the model to achieve semantic consistency between different domains and further improving the accuracy and robustness of emotion recognition. The proposed method is evaluated on the publicly available EEG emotion recognition dataset in Experimental Example 1, Experimental Example 2, and Experimental Example 3. The experimental results show that the proposed unsupervised semantic-aware domain adaptation EEG emotion recognition method of the present invention is significantly superior to previous methods in terms of accuracy and F1 score.
[0073] The above summary is for the purpose of the specification only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present invention will become apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 It is an architecture diagram of the unsupervised semantic-aware domain adaptation EEG emotion recognition method of the present invention;
[0075] Figure 2 It is a generation diagram of the EEG time slice in the present invention;
[0076] Figure 3 It is a diagram of the dual-view frequency domain feature fusion module based on the self-attention mechanism in the present invention;
[0077] Figure 4 It is a diagram of the spatio-temporal feature extraction module in the present invention;
[0078] Figure 5 It is a diagram of the backpropagation process in the pre-training and domain adaptation phases of the present invention; wherein, X S and Y S respectively refer to the data and true labels of the source domain, and X T and respectively refer to the data and pseudo-labels of the target domain.
[0079] Figure 6 It is a comparison diagram of the valence classification accuracy of each subject in the present invention; wherein, in the within-subject paradigm, the classification accuracy of each subject by the compared model before and after domain adaptation.
[0080] Figure 7 It is a comparison diagram of the arousal classification accuracy of each subject in the present invention; wherein, in the within-subject paradigm, the classification accuracy of each subject by the compared model before and after domain adaptation.
[0081] Figure 8 It is a comparison diagram of the valence classification accuracy of each subject in the present invention; wherein, in the between-subject paradigm, the classification accuracy of each subject by the comparison model before and after domain adaptation.
[0082] Figure 9 A comparison graph of the arousal classification accuracy for each subject in the present invention; wherein, under the inter-subject paradigm, the classification accuracy of the comparison model for each subject before and after domain adaptation is compared.
[0083] Figure 10 A comparison graph of the value accuracy for each theme in the present invention; wherein, 1: GCN; 2: BiLSTM; 3: Dual-view frequency domain feature fusion module (DFDFFM); 4: Unsupervised adversarial domain adaptation (UADA); 5: Semantic alignment mechanism).
[0084] Figure 11 A graph of the arousal accuracy for each subject in the present invention; wherein, 1: GCN; 2: BiLSTM; 3: Dual-view frequency domain feature fusion module (DFDFFM); 4: Unsupervised adversarial domain adaptation (UADA); 5: Semantic alignment mechanism).
[0085] Figure 12 A graph of the feature visualization results of the model with different arousal key modules in the present invention; wherein, (a)-(e) correspond to each row in Table X, blue represents the feature data corresponding to high arousal (positive), and red represents the characteristic data corresponding to low arousal (negative).
[0086] Figure 13 A graph of the feature visualization results of the valence model with different key modules in the present invention; wherein, (a)-(e) correspond to each row in Table X, blue represents the feature data corresponding to high valence (positive), and red represents the characteristic data corresponding to low valence (negative).
[0087] Figure 14 A graph of the accuracy based on different weight loss parameter combinations in the present invention; wherein, (a) and (b) respectively represent the average accuracies of Valence and Arousal under the intra-subject paradigm, and (c) and (d) respectively represent the average accuracies of Valence and Arousal under the inter-subject paradigm. Detailed implementation manners
[0088] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0089] Embodiment 1
[0090] Please refer to Figure 1, An unsupervised semantic-aware domain adaptation EEG emotion recognition method. First, the EEG dataset is preprocessed to improve the signal quality, which is beneficial for feature extraction. Second, an unsupervised adversarial semantic-aware domain adaptation framework is used to solve the individual difference problem in EEG emotion recognition. The feature extractor of this framework includes a dual-view frequency-domain feature extraction module based on the self-attention mechanism and a spatio-temporal feature extraction module based on GCN and LSTM.
[0091] Specifically, the preprocessed labeled training data D S ={X S ,Y S} is defined as the source domain, where X S and Y S represent the source domain data and labels respectively. And the unlabeled test data D T ={X T} is defined as the target domain, where X T represents the target domain data. Assume that the source domain label Y S and the target domain label Y T share the same label space, but X S and X T have different but related distributions. Use the labeled source domain samples D S and the unlabeled samples of the target domain D T to train the model to predict the labels of the target domain samples D T .
[0092] Figure 2 is the training and testing process diagram of the proposed method, which consists of three stages:
[0093] (1) Pre-training stage: Use the labeled source domain data D S to train the source domain feature extractor E S and the classifier C. Through the pre-training of this stage, the model obtains preliminary emotion recognition ability.
[0094] (2) Unsupervised domain adaptation stage: Use the unlabeled target domain data D T to improve the recognition performance of the pre-trained model on the target domain. Take the trained source domain feature extractor E S as the initialized target domain feature extractor E T to avoid the training imbalance problem caused by the discriminator D being too strong in the early stage of adversarial domain adaptation learning. In this process, both the source domain and target domain data are used to perform adversarial training on the target feature extractor E T to reduce the difference in data distributions between the source domain and the target domain. In addition, a semantic alignment mechanism is introduced to prevent the model from aligning the features of different domains to the wrong categories while reducing the domain difference.
[0095] (3) Testing phase: Use the feature extractor E trained with unsupervised domain adaptation T and the classifier C to perform emotion recognition on the target domain data, and verify the recognition effect of the model on the actual target domain data.
[0096] Specifically, first, denoise, downsample, and slice the EEG data to split the continuous EEG signal into multiple time segments;
[0097] In the preprocessing process, first, to reduce the model complexity, resample the data to 128 Hz. Second, use a 4 - 45 Hz band - pass filter to remove electro - oculogram, eye movement, and power noise from the electroencephalogram data. Slice the EEG data in seconds, and select continuous T seconds to form a data sample, as Figure 2 shown. The i - th data sample is labeled as x i ∈R T×N×f (i = 1, 2, …n), N represents the number of EEG conductance channels, f is the sampling frequency, and n is the total number of samples.
[0098] In the feature extractor:
[0099] Regarding dual - view frequency - domain feature fusion: Frequency - domain features can reveal the electro - activity patterns of the brain in different states. According to the characteristics of EEG, divide the EEG data into five frequency bands: theta (4 - 7 Hz), slow_alpha (8 - 9 Hz), alpha (8 - 11 Hz), beta (12 - 29 Hz), and gamma (30 - 44 Hz), and introduce a dual - view frequency - domain feature fusion module in the feature extractor. This module applies the attention mechanism to fuse the DE and PSD features extracted from different frequency bands to enhance the complementarity of frequency - domain features.
[0100] Regarding DE feature extraction: DE mainly focuses on the complexity and uncertainty of the EEG signal in the frequency domain. By calculating the differential entropy of the EEG signal, evaluate the energy distribution and information entropy of the signal at different frequency components, so as to capture the frequency - domain features related to the emotional state and extract the DE features of five frequency bands per second. The calculation formula of DE is as follows:
[0101]
[0102] where σ is the variance of x, and e is the Euler's constant.
[0103] After DE feature extraction, further transform the segmented EEG segments into a feature matrix M represented by DE segments i ∈R T ×N×K(i = 1, 2, … n), where T represents the time window length, N represents the number of EEG channels, K is the number of frequency bands, and n is the total number of samples.
[0104] Regarding PSD feature extraction: PSD describes the power intensity of a signal at different frequencies and can capture temporal information such as the trend and periodicity of EEG changes over time. The Welch method is used to extract PSD. First, the original signal x(n) is segmented into K sub-segments of length N, and a window function w(n) is applied to each sub-segment to reduce spectral leakage. Then, the fast Fourier transform is applied to each segment of the signal to obtain spectral information, and the calculation formula is as follows:
[0105]
[0106] where x k (n) represents the signal of the k-th sub-segment, w(n) is related to the Hann window, f is the frequency variable, and X i (k) represents the spectral information obtained after the FFT transformation. The frequency f is set to 128 Hz, and the window size is set to 1 second. The power spectral density of each sub-segment is calculated by taking the square of the modulus of the FFT result X k (f) of each window:
[0107]
[0108] where P k (f) is the estimated power spectral density within the k-th window, N is the length of each window, and X k (f) is the FFT result of the k-th segment of the signal. Finally, the PSDs of all sub-segments are averaged to obtain the PSD of the entire signal, and the calculation formula is as follows:
[0109]
[0110] where K represents the total number of sub-segments, and P(f) is the estimated value of PSD.
[0111] Regarding dual-view frequency-domain feature fusion based on the self-attention mechanism: The core of this mechanism is to use the attention mechanism to achieve effective fusion of DE features and PSD features, and its formula is defined as follows:
[0112]
[0113] where Q, K, and V represent Query, Key, and Value respectively. d k is the size of the last dimension of the query. As Figure 3As shown in the figure, DE and PSD are regarded as key vectors, query vectors, and value vectors respectively, and their interaction and fusion are achieved through scaled dot - product attention. Then, the results of the dual - perspective frequency - domain features calculated based on the self - attention mechanism are fused through the following formula to obtain the new feature F dual . The feature fusion formula is as follows:
[0114] F dual (f1,f2)=Attention(f1,f2,f2)+Attention(f2,f1,f1)
[0115] where f1 represents the feature DE, and f2 represents the feature PSD.
[0116] Regarding spatio - temporal feature extraction, based on the characteristics of the electrode distribution structure and the temporal characteristics in the EEC signal, the respective advantages of GCN and BiLSTM are utilized and combined skillfully to better capture the spatial and temporal features of the emotional state.
[0117] Spatio - temporal feature extraction includes a spatial feature extractor and a temporal feature extractor, and the overall structure is as Figure 4 shown. In the spatial feature extractor, T parallel GCNs are designed to receive the DE feature input matrix of T seconds in chronological order and output the calculation results to the temporal feature extractor in chronological order. In the temporal feature extractor, a bidirectional long - short - term memory network module with T time steps is designed to further mine the EEG temporal - dimension features.
[0118] Regarding spatial feature extraction, based on the application of the graph convolutional neural network model in image processing and the graph - modeling method, the problem of multi - lead EEG spatial feature extraction is studied. The EEG signals of each channel and the connection relationship between channels are modeled as a graph G={V,ε,A}, where V is a feature set containing N nodes, ε is the set of edges, and A is the adjacency matrix, representing the tightness of the connection between nodes. Based on the Pearson correlation coefficient, the linear correlation degree between two signals is evaluated according to the time - domain amplitude of the signals, and its calculation formula is:
[0119]
[0120] where i and j represent two different nodes, μ i and μ j represent the means of nodes i and j, and E is the expected value. σ i and σ j represent the standard deviations of i and j respectively. In addition, according to the following formula, the threshold of the correlation coefficient is set to 0.5 to obtain the adjacency matrix A, and the formula is as follows:
[0121]
[0122] The propagation rule of graph convolution is defined as follows:
[0123]
[0124] where represents the node feature representation of the (l + 1)-th layer in the t-th time window, which depends on the feature representation of the l-th layer is the adjacency matrix A plus the self-connection term, and I is the identity matrix; is the weight matrix of the l-th layer in the t-th time window, and σ(·) is the activation function.
[0125] Regarding time feature extraction, the time feature extractor is used to extract and characterize the time-dependent relationships between EEG channels within T seconds, so as to provide high-quality emotion feature representations for subsequent recognition tasks. A BiLSTM is used as the time feature extractor to receive the calculation results obtained from T GCNs. The BiLSTM consists of two LSTM networks, one responsible for forward propagation and the other for backward propagation. This design combines the advantages of long short-term memory and bidirectional data streams, and can utilize the context information of EEG simultaneously, providing a more comprehensive perspective on the data at each time point, as Figure 4 shown
[0126] Inside each LSTM cell, there are an input gate, a forget gate, an output gate, and a memory unit, which are used to learn the long-term dependencies in sequential data. The forget gate determines how much of the cell state information from the previous time step needs to be discarded at the current time step through a sigmoid layer. The input gate determines how much of the existing input information is added to the cell state. The cell state is updated according to the forget gate and the input gate, retaining the long-term dependence on past information. Finally, the output gate determines the hidden state output at this time step based on the current cell state. This enables the entire time feature encoder to maintain sensitivity to time-related information, thus better capturing and characterizing the important features on long time scales in EEG data. Given the input sequence x = (x1, x2, …, x T ), its calculation process can be expressed as:
[0127] f t = σ(W f · [h t-1 , x t + b f )
[0128] i t = σ(W i · [h t-1 , x t + b i )
[0129] o t= σ(W o · [h t-1 , x t + b o )
[0130]
[0131]
[0132] h t = o t ⊙ tanh(C t )
[0133] Among them, f t , i t , o t respectively represent the calculation results of the forgetting gate, input gate, and output gate. C t is the cell state at time t, h t is the output at time t, x t is the input at the current time, and W and b are the weight matrix and bias term respectively.
[0134] Regarding the classifier, its role is to flatten and concatenate the output of the spatio-temporal feature extractor into a feature representation vector h, and output the classification result by calculating the probability distribution of the predicted class. The specific definition is as follows:
[0135] f(h) = Wh + b
[0136] Among them, W represents the transformation matrix, and b represents the bias.
[0137] The activation function softmax is mainly used to provide probability values for each emotional state category. The probability values are between 0 and 1, and the maximum value is used as the final emotional state. The specific calculation formula is expressed as:
[0138]
[0139] S = argmax j P(j|X n )
[0140] Among them, P(j|X n ) represents the probability that the sample X n belongs to the emotional state category j, and f(h) is the output of the fully connected layer. c represents the number of emotional state categories. S represents the emotional state category with the maximum probability.
[0141] Regarding the discriminator, the main functions of the discriminator D are two:
[0142] (1) Further encode the features output by the target feature extractor E T ;
[0143] (2) Act as a classifier to distinguish whether the encoded features come from the source domain or the target domain.
[0144] Therefore, the discriminator consists of three fully connected layers and corresponding activation functions. The first two fully connected layers are configured with BN and ReLu layers, and the last fully connected layer is combined with Sigmoid as a classifier. The definition of Sigmoid is as follows:
[0145]
[0146] where x i is the output of the i-th neuron in the previous layer, and f sigmoid is the output probability of the i-th class, ranging from [0, 1].
[0147] The optimization training is divided into two stages: pre-training and adversarial domain adaptation. The gradient transfer process of the optimization process is as Figure 5 shown. In the pre-training stage, the parameters of the feature extractor are mainly updated through the classification loss. The goal of the adversarial domain adaptation stage is to prompt the feature extractor to obtain shared features between the source domain and the target domain while extracting valuable classification features. In addition, semantic alignment between the source domain and the target domain also needs to be ensured in this stage. Therefore, the adversarial domain adaptation process includes classification loss, domain alignment loss, and semantic alignment loss.
[0148] Specifically, the cross-entropy loss is used as the classification loss function. The cross-entropy loss can measure the difference between two probability distributions, thereby optimizing the feature extractor to learn valuable classification features and improving the classification performance. The main purpose of the classification loss is to update the parameters of the feature extractor θ E and the classifier parameter θ C , and its definition is as follows:
[0149]
[0150] where y and p(x) represent the true label and the predicted label respectively, and N represents the number of samples.
[0151] The global domain alignment loss is calculated from the domain true label and the predicted label generated by the domain discriminator. Since the domain discriminator D is also a binary classification task, the formula principle of the domain alignment loss is the same as that of the cross-entropy loss. However, the main purpose of the global domain alignment loss is to help update the parameters of the target domain feature extractor and the discriminator θ D , and its definition is as follows:
[0152]
[0153] The global domain alignment loss mainly has two functions:
[0154] (1) Improve the ability of the domain discriminator D to identify whether the features come from the source domain or the target domain;
[0155] (2) Prompt the feature extractor E T to learn the shared features of the source domain and the target domain, so that the discriminator D cannot identify the source of the features.
[0156] Therefore, in order to achieve the above two purposes simultaneously, a gradient reversal layer is added between the domain discriminator D and the feature extractor E. T This layer reverses the sign of the gradient during backpropagation, enabling the feature extractor to maximize the loss of the domain discriminator during the optimization process, thus forming an adversarial relationship. Therefore, after adding the gradient reversal layer, the optimization objective of the feature extractor will be opposite to that of the task discriminator.
[0157] Regarding the semantic alignment loss, semantic alignment is achieved by calculating the distance between the feature representations of the same category data in the source domain and the target domain. The formula is defined as follows:
[0158]
[0159] where K represents the total number of categories, and represent the feature centers of the k-th category in the source domain and the target domain respectively. φ represents the cosine similarity, which is used to calculate the distance between the corresponding category feature centers. Therefore, during the training process, the semantic alignment loss can ensure that the feature extractor extracts features of the same semantic category from different domains and maps them to adjacent regions.
[0160] During the training process, the trained source domain feature extractor and classifier are used to output the pseudo-labels of the target domain samples During the process of adversarial domain adaptation learning, the accuracy of the pseudo-labels will continuously improve with training. Therefore, the pseudo-labels of the target domain are continuously updated to ensure the effect of semantic alignment.
[0161] Regarding the total loss, according to the above, the total loss of the adversarial domain adaptation process consists of three parts: classification loss, domain alignment loss, and semantic alignment loss. The specific definition is as follows:
[0162] L T = L C + w d L D + w s L S
[0163] where L C , L D and L S represent the classification loss, domain alignment loss, and semantic alignment loss respectively. w dand w s They are the weight coefficients corresponding to the respective losses. During the adversarial domain adaptation process, these three loss functions jointly guide the feature extractor to learn domain-shared features and distinguishable semantic representations, and prevent incorrect semantic alignment.
[0164] Experimental Example 1
[0165] DEAP is a multi-modal emotion database. This database contains the physiological signals of 32 subjects when watching 40 one-minute music emotion videos. The physiological signals are respectively composed of 32-lead EEG data and 8-lead peripheral physiological signals. Only EEG signals are used. The duration of each EEG data is 63s. The first 3s is the experimental preparation stage, and the subsequent 60s is the time when the subjects watch the MV, that is, the first 3 seconds of data are not used in this experiment.
[0166] After each subject watches the music video, they need to give an emotional score of 1-9 for each video clip, including Valence, Arousal, Dominance, and Liking. Emotional recognition is performed on the two labels of Valence and Arousal. Less than 5 is low price, and the rest are classified as high price.
[0167] Experimental platform and settings:
[0168] This invention is implemented based on python and pytorch. Adam is used as the optimizer during the training process. The initial learning rate is 0.001. After each iteration, the learning rate is gradually decreased by multiplying by a decay factor of 0.997. The maximum number of epochs is 1000, and the model with the best effect is selected as the final model. The Batch size is 128.
[0169] Evaluation metrics:
[0170] When evaluating the performance of the EEG emotion recognition model, the evaluation metrics include accuracy and F-Score.
[0171] Accuracy refers to the ratio of the number of samples correctly classified by the classifier to the total number of samples. The formula is expressed as:
[0172]
[0173] Among them, TP represents the number of true positive cases, TN represents the number of true negative cases, FP represents the number of false positive cases, and FN represents the number of false negative cases. Accuracy provides the overall classification accuracy of the model.
[0174] F1 is an index that comprehensively considers precision and recall, and has good evaluation performance for unbalanced data sets. The calculation formula of F1 is:
[0175]
[0176] Among them, Pre represents the proportion of samples predicted as positive by the classifier that are actually positive, and the calculation formula is:
[0177]
[0178] Rec represents the proportion of samples that are actually positive and are predicted as positive by the classifier, and the calculation formula is:
[0179]
[0180] Experimental process:
[0181] The proposed emotion recognition method is evaluated through two experimental paradigms, within-subject and between-subject. In the experiment based on the between-subject paradigm, the EEG data of each subject is divided into a training set and a test set. Modeling is performed for each subject respectively, and the final average result is calculated. In the experiment based on the between-subject paradigm, a leave-one-subject-out cross-validation strategy is adopted to evaluate the performance of the proposed method, that is, the EEG signals of 31 subjects are used as training data (source domain data), and the EEG signals of the remaining 1 subject are used as test data (target domain data) for the experiment. This process is repeated to ensure that the EEG signal of each subject is used as test data once.
[0182] Within-subject: Although there is no individual difference in the experiment based on the within-subject paradigm, there is non-stationarity between the training set and the test set. Table 1 lists the emotion recognition results of the proposed method for each subject in the within-subject paradigm.
[0183] The experimental results show that in the binary classification of Valence, the lowest, highest, and average classification accuracies of the proposed method for 32 subjects are 84.57%, 99.85%, and 97.41% respectively. In the binary classification of Arousal, the lowest, highest, and average classification accuracies of 32 subjects are 85.74%, 100%, and 97.10% respectively. The differences in accuracy between the best result and the worst result are 15.28% and 14.26% respectively. This shows that the non-stationarity existing in the data within each subject is different. On the other hand, in these two classification tasks, the average value of the F1 score reaches more than 96.5%. This shows that the emotion recognition method proposed by the present invention has the ability to resist sample imbalance. Figure 6 And Figure 7 Shows the changes in the classification accuracies of valence and arousal before and after domain adaptation. It can be seen from the figure that very few subjects are negatively affected by domain adaptation, but the accuracies of most subjects are effectively improved after domain adaptation.
[0184]
[0185]
[0186] Table 1 Emotional recognition results of the proposed method under the within-subject paradigm
[0187] Within-subject: The main purpose of the experiment based on the within-subject paradigm is to verify the performance of the method in overcoming individual differences. Table 2 lists the emotion recognition results of the proposed method for each subject under the between-subject paradigm.
[0188] The experimental results show that in the binary classification of Valence, the lowest, highest, and average classification accuracies of the proposed method for 32 subjects are 57.42%, 79.03%, and 66.04% respectively. In the binary classification of Arousal, the lowest, highest, and average classification accuracies of 32 subjects are 50.64%, 87.09%, and 68.75% respectively. According to the experimental results, the experiment of the between-subject paradigm is more difficult than that of the within-subject paradigm. The accuracy and F1 of each subject are lower than those of the within-subject paradigm, and the difference between the best and worst results is also larger.
[0189] Therefore, it can be concluded that: (1) The negative impact of individual differences on the model is greater than that of non-stationarity. (2) Individual differences have a certain impact on the robustness of the model.
[0190] Figure 8 And Figure 9 Shows the changes in the classification accuracy of the model for each subject before and after domain adaptation. It can be seen that domain adaptation has a certain degree of improvement in the accuracy of almost each subject, especially the Valence of subject 29 and the Arousal of subject 3 are the most obvious.
[0191]
[0192]
[0193] Table 2 Emotional recognition results of the proposed method under the within-subject paradigm
[0194] Experimental Example 2
[0195] By designing ablation experiments to verify the effectiveness of the key modules of the method. Among them, the key modules include GCN, BiLSTM, dual-view frequency domain feature fusion module, spatio-temporal feature extraction module, unsupervised adversarial domain adaptation, and semantic alignment mechanism.
[0196] By gradually increasing the key modules to analyze the changes in the model effect and prove the role of the modules. The experiment is carried out based on the between-subject paradigm.
[0197] Table 3 shows the average performance of models with five different combinations of modules in terms of arousal and valence. The experimental environments, datasets, and preprocessing for the five groups of models are the same. It can be seen that as different key modules are added, the recognition ability of the model gradually improves. In particular, the dual-view frequency-domain feature fusion and unsupervised adversarial domain adaptation modules have the most obvious improvement on the method. The unsupervised adversarial domain adaptation module helps to improve the accuracy of Valence and Arousal by 5.65% and 4.79% respectively, and the F1 by 7.42% and 4.75% respectively. Figure 10 And Figure 11 shows the accuracy changes of each subject in the five groups of ablation experiments. It can be seen that different key modules have improved the recognition effect of each subject to a certain extent. For example, the dual-view frequency-domain feature fusion module has significantly improved the Valence of subject 16. In addition, it can also be seen that the performance of a small number of subjects has decreased after adding semantic alignment. This is because the individual differences are too large, resulting in the model being unable to provide accurate pseudo-labels.
[0198] In addition, to further prove the role of different key modules, t-SNE technology is used for visual analysis of the model. T-SNE can map high-dimensional feature data to a low-dimensional space and generate corresponding scatter plots to clearly reflect the feature distribution.
[0199] Figure 12 and Figure 13 are the visualization results of Valence and Arousal corresponding to the five different ablation experiments respectively, and the target domain data comes from subject 29. According to the distribution of feature points, the model can gradually extract valuable classification features after adding different key modules. In particular, after adding the dual-view frequency-domain feature fusion and unsupervised adversarial domain adaptation modules, the visualization effect of the model has been significantly improved, and the distance between points in the same category cluster is more compact.
[0200] Generally speaking, each key module has a significant effect on improving the emotion recognition effect based on EEG signals, and the mutual combination between different modules is appropriate.
[0201]
[0202] Average performance of the five groups of ablation experiments in Table 3
[0203] Among them, 1: GCN; 2: BiLSTM; 3: Dual-view frequency-domain feature fusion module (DFDFFM); 4: Unsupervised adversarial domain adaptation (UADA); 5: Semantic alignment mechanism
[0204] Experimental Example III
[0205] Verify the impact of key hyperparameters on the model's capabilities. The key hyperparameters mainly include the time window length, shift, and loss weights.
[0206] Increase the time window length from 6 seconds to 15 seconds and vary the shift value between 1 second and 2 seconds to determine the appropriate window length and shift.
[0207] According to the results in Table 4, the proposed method performs best when the window length and shift are set to 9 seconds and 1 second, respectively. In the total loss of the domain adaptation process, w d and w s are the weights of the domain alignment loss and semantic alignment loss, respectively, and they have a crucial impact on the results. Use the grid search method to find the appropriate weights w = {w d , w s}.
[0208] Specifically, the candidate values for each weight are 0.001, 0.005, 0.01, 0.05, 0.1. A total of 25 groups of experiments were conducted according to different combinations. Figure 14 Shows the average accuracy of Valence and Arousal for different parameter combinations under two paradigms. It can be seen from the figure that the best results are achieved in the within-subject and between-subjects paradigms with parameter settings of w = {0.05, 0.5} and w = {0.1, 0.01}, respectively.
[0209]
[0210] Table 4 Precision comparison of the proposed method based on different window lengths and offsets
[0211] Those of ordinary skill in the art can understand that all or part of the steps carried out in the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0212] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0213] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. An unsupervised semantic perception domain adaptation EEG emotion recognition method, characterized in that It includes the following steps: Step 1: Denoise, downsample, and slice the EEG data to split the continuous EEG signal into multiple time segments; Among them, the data resampling rate is 128Hz. A band-pass filter with a frequency range of 4 - 45Hz is used to remove electrooculogram, eye movement, and power supply noise from the electroencephalogram data. The EEG data is sliced in seconds, and continuous T seconds are selected to form a data sample; Step 2: Extract dual-view frequency-domain features and spatio-temporal features from the preprocessed EEG data. First, calculate the differential entropy of the five frequency bands of theta, slow_alpha, alpha, beta, and gamma in each time segment to obtain the DE feature matrix. Use the Welch method to calculate the power spectral density of the five frequency bands in each time segment to obtain the PSD feature matrix. Based on the self-attention mechanism, fuse the DE feature matrix and the PSD feature matrix into a dual-view frequency-domain feature matrix; Step 3: For spatial feature extraction, input the DE feature matrix of each time segment into the GCN to obtain the spatial feature representation of each time segment; for temporal feature extraction, input the spatial feature representations of the T time segments output by the GCN into the BiLSTM to obtain the temporal feature representation of each time segment; Step 4: Train the classifier. Flatten and concatenate the output of the spatio-temporal feature extractor into a feature representation vector h, and output the classification result by calculating the probability distribution of the predicted class; train the discriminator to further encode the features output by the target feature extractor and act as a classifier to distinguish whether the encoded features come from the source domain or the target domain; Step 5: Use the cross-entropy loss as the classification loss function to measure the difference between the two probability distributions, thereby optimizing the feature extractor. Calculate the global domain alignment loss from the domain true label and the predicted label generated by the domain discriminator; Step 6: Introduce a semantic alignment mechanism. During the training process, use the trained source domain feature extractor and classifier to output the pseudo-labels of the target domain samples. During the adversarial domain adaptation learning process, the accuracy of the pseudo-labels continuously improves with training.
2. The unsupervised semantic-aware domain adaptation EEG emotion recognition method according to claim 1, characterized in that: In Step 2, the specific steps to calculate the differential entropy of the five frequency bands in each time segment to obtain the DE feature matrix are as follows: First, extract the DE features of the five frequency bands per second. The formula for DE is as follows: where σ is the variance of x, and e is the Euler's constant; After DE feature extraction, the segmented EEG segments are further transformed into a feature matrix M represented by DE segments i ∈R T×N×K (i = 1, 2, … n), where T represents the time window length, N represents the number of EEG channels, K is the number of frequency bands, and n is the total number of samples 3. The unsupervised semantic-aware domain adaptation EEG emotion recognition method according to claim 1, characterized in that: In Step 2, the specific steps to use the Welch method to calculate the power spectral density of the five frequency bands in each time segment to obtain the PSD feature matrix are as follows: First, split the original signal x(n) into K sub-segments with a length of N, and apply a window function w(n) to each sub-segment to reduce spectral leakage. Then, apply the fast Fourier transform to each segment of the signal to obtain the spectral information. The calculation formula is as follows: where x k (n) represents the signal of the k-th sub-segment, w(n) is related to the Hanning window, f is the frequency variable, X i (k) represents the spectrum information obtained after FFT transformation; After the frequency f is set to 128 Hz, the window size is set to 1 second, and the power spectral density of each sub-segment is calculated by taking the square of the modulus of the FFT result X k (f): where P k (f) is the power spectral density estimate within the k-th window, N is the length of each window, and X k (f) is the FFT result of the k-th segment of the signal; Finally, average the PSDs of all sub-segments to obtain the PSD of the entire signal. The calculation formula is as follows: where K represents the total number of sub-segments, and P(f) is the estimated value of the PSD.
4. The unsupervised semantic-aware domain adaptation EEG emotion recognition method according to claim 1, characterized in that: In Step 2, the specific steps to fuse the DE feature matrix and the PSD feature matrix into a dual-view frequency-domain feature matrix based on the self-attention mechanism are as follows: The formula used is defined as follows: Among them, Q, K, and V represent Query, Key, and Value respectively, and d k is the size of the last dimension of the query; Regarding DE and PSD as the key vector, query vector, and value vector respectively, their interaction and fusion are realized through scaled dot-product attention. Then, the results of the dual-view frequency-domain features calculated based on the self-attention mechanism are fused through the following formula to obtain the new feature F dual , and its feature fusion formula is as follows: F dual (f1,f2) = Attention(f1,f2,f2) + Attention(f2,f1,f1) Among them, f1 represents the feature DE, and f2 represents the feature PSD.
5. The unsupervised semantic-aware domain adaptation EEG emotion recognition method according to claim 1, characterized in that: In step three, the specific steps for extracting spatial features are as follows: First, model the EEG signals of each channel and the connection relationship between channels as a graph G = {V, ε, A}, where V is a feature set containing N nodes, ε is the set of edges, and A is the adjacency matrix, representing the tightness of the connection between nodes; Based on the Pearson correlation coefficient, evaluate the linear correlation degree between two signals according to the time-domain amplitude of the signals, and its calculation formula is as follows: where i and j represent two different nodes, μ i and μ j represent the means of nodes i and j, E is the expected value, σ i and σ j represent the standard deviations of i and j respectively; In addition, according to the following formula, set the threshold of the correlation coefficient to 0.5 to obtain the adjacency matrix A, and its calculation formula is as follows:
6. The unsupervised semantic perception domain adaptation EEG emotion recognition method according to claim 1, characterized in that: In step three, the specific steps for extracting temporal features are as follows: Use BiLSTM as the time feature extractor to receive the calculation results obtained from T GCNs. Each LSTM unit has an input gate, a forget gate, an output gate, and a memory unit inside, which are used to learn the long-term dependencies in the sequence data. The forget gate determines how much of the cell state information from the previous time step needs to be discarded at the current time step through a sigmoid layer. The input gate decides how much of the existing input information should be added to the cell state. The cell state is updated according to the forget gate and the input gate, retaining the long-term dependence on past information. Finally, the output gate determines the hidden state output at this time step based on the current cell state, thereby maintaining sensitivity to time-related information, so as to better capture and characterize the important features on the long time scale in EEG data. Given the input sequence x = (x1, x2, …, x T ), its calculation process can be expressed as: f t = σ(W f · [h t-1 , x t + b f ) i t = σ(W i · [h t-1 , x t + b i ) o t = σ(W o · [h t-1 , x t + b o ) h t = o t ⊙tanh(C t ) Among them, f t , i t , o t represent the calculation results of the forget gate, input gate, and output gate respectively. C t is the cell state at time t, h t is the output at time t, x t is the input at the current time, and W and b are the weight matrix and bias term respectively.
7. A computer-readable storage medium, characterized in that, Instructions are stored on the storage medium, and when the instructions are executed on a computer, the computer executes the unsupervised semantic-aware domain adaptation EEG emotion recognition method described in any one of claims 1 to 6.
8. An electronic device, characterized in that, Including: One or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device executes the unsupervised semantic-aware domain adaptation EEG emotion recognition method described in any one of claims 1 to 17.