Electroencephalogram emotion recognition method based on ChannelMix and bidirectional spatial-temporal feature fusion

By adopting ChannelMix and bidirectional spatiotemporal feature fusion method in EEG signal recognition, the problem of poor generalization ability in cross-individual emotion recognition is solved, and higher recognition accuracy and better generalization performance are achieved.

CN120045893APending Publication Date: 2025-05-27BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411940979.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing cross-individual emotion recognition methods based on EEG signals are difficult to effectively generalize to the target subjects, and they fail to fully utilize the complementarity of timing and spatial characteristics, resulting in a low recognition accuracy.

Method used

An EEG emotion recognition method based on ChannelMix and bidirectional spatiotemporal feature fusion is proposed. The spatial and fusion features of the multi-stage bidirectional spatiotemporal feature fusion module is extracted, and the discriminant nature of the target domain data is enhanced by using ChannelMix strategy and soft pseudo-label module to achieve smooth alignment of feature space and tag space.

Benefits of technology

The accuracy of cross-individual EEG emotion recognition is improved, and the generalization performance of the model in the target domain is improved by effectively fusion of spatiotemporal features and reducing interdomain differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045893A_ABST
    Figure CN120045893A_ABST
Patent Text Reader

Abstract

The invention discloses an electroencephalogram emotion recognition method based on ChannelMix and bidirectional spatial-temporal feature fusion, belongs to the field of artificial intelligence and machine learning, and researches an unsupervised domain adaptive electroencephalogram emotion recognition method. The method comprises the following steps: firstly, fully mining space-time complementary information of an electroencephalogram signal by using a multi-stage bidirectional space-time feature fusion module so as to extract rich space-time representation; then forcing a spatial-temporal feature fusion module to extract domain invariant features through a domain discriminator; secondly, generating a pseudo label for a target domain sample through a soft pseudo label module, and taking the pseudo label as an additional supervision signal, thereby improving the discrimination of a target domain; then, through a ChannelMix module, an intermediate domain can be effectively established, and the source domain and the target domain are aligned to the intermediate domain, so that the difference between the source domain and the target domain is reduced; and finally, the generalization performance is further improved by a data enhancement method based on ChannelMix. The emotion recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and machine learning, and particularly relates to emotion recognition of electroencephalogram (EEG) signals using a deep neural network. Background Art

[0002] Emotion is an important part of human mental activities, profoundly affecting our lives and playing an important role in aspects such as physical health, decision-making, learning, and work. At present, emotion recognition technology is becoming increasingly common in various industries, such as human-computer interaction, gaming and entertainment, healthcare, etc. With the rapid development of computer and human-computer interaction technologies, it has become a popular research topic. Emotion recognition technology can be roughly divided into two categories: one is the method based on non-physiological signals, such as facial expressions, voices, and body movements. The other is the method based on physiological signals, such as electroencephalogram (EEG), electrocardiogram (ECG), and electrooculogram (EOG). Compared with non-physiological signals, physiological signals are more difficult to deceive or disguise. Therefore, the emotion recognition method based on physiological signals is often more reliable. In particular, EEG can more directly reflect a person's mental and emotional state by directly recording the activities of nerve cells in the cerebral cortex, and has a high recognition accuracy. In addition, due to its characteristics of high temporal resolution, non-invasiveness, and low cost, emotion recognition using EEG has been increasingly favored by researchers.

[0003] In the past few decades, significant progress has been made in EEG-based emotion recognition technology. Early machine learning-based methods relied on two processes, feature design and classifier learning, to achieve emotion recognition. However, such methods rely on complex feature engineering and feature selection. In recent years, deep learning technology has achieved great success in many fields, including computer vision, natural language processing, etc. Researchers have proposed various frameworks and methods for EEG emotion recognition tasks based on deep learning. For example, some researchers use convolutional neural networks (CNNs) to process emotion recognition tasks, significantly improving the accuracy of EEG emotion recognition. In addition, the application of the Transformer model in EEG emotion recognition has also received attention.

[0004] Benefiting from the powerful automatic feature extraction capabilities of deep learning models, emotion recognition based on EEG signals has made rapid progress. However, individual differences between different subjects still pose a huge challenge to the EEG emotion recognition task, making it difficult for models trained on source subject data to be effectively generalized to target subjects. At the same time, because cross-individual experiments are more practical in real life, more and more researchers have recently tended to study cross-individual methods. Although these EEG signal recognition methods have improved the cross-individual emotion classification performance to a certain extent, the huge data distribution offset of EEG signals causes data from different individuals or new situations of the current individual to still lack generalization capabilities.

[0005] To address this problem, some researchers have used the unsupervised domain adaptation (UDA) method to improve generalization. The UDA method aims to utilize labeled source domain data and unlabeled target domain data and map them into a similar feature space, thereby alleviating individual differences and enabling the model trained on the source subject data to perform well on the target subject.

[0006] Although researchers have done a lot of work on emotion recognition based on unsupervised domain adaptation, there are still some problems that need to be solved. For example, most existing studies only consider the extraction of a single feature, without considering the temporal information and spatial topological information of the EEG signal at the same time. Or they can extract the spatiotemporal features of the EEG signal, but the extraction process of the spatiotemporal features is separated from each other, and the efficient fusion of spatiotemporal features cannot be achieved. Secondly, how to improve the discriminability of the target domain data in the latent space is also an urgent problem to be solved. By improving the discriminability of the target domain data in the latent space, the model can make reliable prediction results on the target domain data. In addition, when the distribution difference between the source domain and the target domain data is too large, it will bring more arduous challenges to the cross-individual EEG emotion recognition task, resulting in the deterioration of the generalization performance of the model in the target domain. In response to these problems, the present invention designs a novel and effective EEG emotion recognition method based on ChannelMix and bidirectional spatiotemporal feature fusion. Summary of the invention

[0007] The present invention proposes an EEG emotion recognition method based on ChannelMix and bidirectional spatio-temporal feature fusion. First, in order to make full use of the spatio-temporal complementary features of EEG signals, the present invention proposes a spatio-temporal feature fusion module, which can capture multi-view spatial features while learning the temporal features of EEG signals. And in the process of extracting spatio-temporal features, through multi-stage bidirectional spatio-temporal feature fusion, the spatio-temporal representation of EEG is fully mined. Second, a soft pseudo-label module is proposed, which can generate credible soft pseudo-labels for target domain data, enabling the model to iteratively train the classifier using both labeled source domain data and target domain data with soft pseudo-labels, so as to enhance the discriminability of target domain samples in the latent space. In addition, the ChannelMix strategy is proposed. It samples data on each channel of source domain and target domain samples by using sampling weights generated by Beta distribution, which can effectively construct an intermediate domain. At the same time, two loss functions are used in the feature space and label space to narrow the differences between the intermediate domain and the source domain and target domain data. Compared with directly aligning the source domain and target domain, ChannelMix can make the alignment process smoother and simpler. Finally, a data augmentation method is proposed based on the ChannelMix strategy. Similar to Mixup, ChannelMix is applied separately to the source domain or target domain, thereby increasing the diversity of source domain and target domain data, and further improving the generalization performance of the model. To verify the effectiveness of the proposed algorithm, a large number of experiments are carried out on two publicly available emotion recognition datasets (SEED, SEED-IV), and the average accuracies of 93.80% (±04.96) and 79.37% (±06.05) are obtained respectively. The experimental results show that the present invention can further improve the emotion recognition accuracy in cross-subject experiments based on EEG. Description of the Drawings

[0008] Figure 1 EEG Emotion Recognition Method Based on ChannelMix and Bidirectional Spatio-Temporal Feature Fusion

[0009] Figure 2 Multi-Stage Bidirectional Spatio-Temporal Feature Fusion Module

[0010] Figure 3 Multi-View Feature Pyramid Module

[0011] Figure 4 Multi-View Feature Extraction Module

[0012] Figure 5 Spatial Feature Incorporating Temporal Feature Module

[0013] Figure 6 Temporal Feature Incorporating Spatial Feature Module

[0014] Figure 7 Soft Pseudo-Label Module Detailed implementation manners

[0015] The present invention provides an EEG emotion recognition method based on ChannelMix and bidirectional spatio-temporal feature fusion. This method regards the cross-subject EEG emotion recognition task as an unsupervised domain adaptation problem, aiming to reduce the domain shift of EEG data between different subjects and generalize the model trained on the labeled source domain data to the unlabeled target domain. Formally, define a labeled source domain dataset representing the i-th EEG signal sample and its corresponding one-hot label and an unlabeled target domain dataset representing the j-th EEG signal sample n s and n / respectively represent the number of samples in the source domain and the target domain. Note that the data in the two domains are sampled from different but related distributions, and it is assumed that the two domains have a common label space Y ∈ R K , where K represents the number of categories. The goal is to solve the significant domain divergence problem, smoothly transfer the knowledge learned from the source domain to the target domain, and improve the generalization performance of the model in the target domain.

[0016] As Figure 1 shown, the framework mainly includes five parts: ChannelMix spatio-temporal feature fusion module classifier domain discriminator and soft pseudo-label module (SPLM) The goal is to predict the labels of the unlabeled samples in the target domain by learning a combined network : This combined network is trained on D s and D t .

[0017] The invention mainly includes the following steps: Step 1, multi-stage bidirectional spatio-temporal fusion feature extraction; Step 2, reducing the inter-domain difference based on the domain discriminator; Step 3, generating pseudo-labels for the target domain samples through the soft pseudo-label module; Step 4, smoothly aligning the source domain and the target domain through ChannelMix; Step 5, data augmentation based on ChannelMix; Step 6, end-to-end joint training.

[0018] Step 1, multi-stage bidirectional spatio-temporal fusion feature extraction

[0019] Since neural activities are usually continuous, it is considered that the context-related representation in the spatio-temporal features of EEG signals is beneficial to the final emotion recognition. Specifically, the spatial dimension reflects the dependencies between different brain regions, while the time dimension reflects the correlations of brain activities between different time slices. Therefore, a spatio-temporal feature fusion module (as shown in Figure 2 ) is proposed to extract spatio-temporal fusion features. The module is divided into two branches. One branch is a naive Transformer with L layers, where L is set to 10 and it is divided into H stages for feature fusion, and H is set to 5, which is used to extract the temporal features of EEG signals. The other branch is a CNN-based multi-view spatial feature extraction module, in which the multi-view feature pyramid (MVFP) is used to provide multi-view spatial features, and the multi-view feature extraction (MVFE) is used to further extract multi-view spatial features for feature fusion at each stage. The spatial feature integrating into temporal feature module (STT) and the temporal feature integrating into spatial feature module (TTS) are used to fuse the spatio-temporal features of the two branches at different stages, helping to enrich the semantic information of the features.

[0020] To extract the temporal features in EEG data, an overlapping sliding window with a window size of T along the time dimension is adopted, and T is set to 16, and the step size of the sliding window is set to 8. Therefore, the EEG signal data input to this module is represented as (C and B represent the number of channels and frequency bands respectively), x represents the EEG signal data segment sampled by the sliding window, where C is 62 and B is 5. To extract the spatio-temporal features of EEG signals simultaneously, the input x of the module is reshaped into as the input of the Transformer branch to extract temporal features, and CB is 310. According to the spatial position information of the electrodes, x is mapped into as the input of the convolutional branch to extract multi-view spatial features, and TB is 80. First, for the Transformer branch, x / is input into the Embedding and added with the positional embedding to obtain features with a shape of T×M. At the same time, for the other branch, x s passes through the MVFP module (as shown in Figure 3 ) to obtain feature pyramids O 1 ×M, L 2 ×M and L 3 ×M, where L 1 , O 2 and O 3 , L 1 , L 3 , L 3And M is 64, 16, 4, 64. The Embedding module and the MVFP module are not only used for feature extraction but also play a role in data dimensionality reduction, which can significantly reduce the number of model parameters and prevent overfitting. Secondly, the two branches undergo feature interaction in H stages. In each stage, first, the MVFE module is used to further extract multi-view spatial features, then the STT module integrates the spatial features into the temporal features, then the Transformer Blocks further strengthen the fused temporal features, and finally, the TTS module integrates the temporal features into the spatial features.

[0021] The MVFE module (as Figure 4 shown) consists of a group of convolutional layers and a fully connected layer. Specifically, the input of this module is a group of multi-view features output by the MVFP module which first passes through different convolutional layers to extract spatial features of different views, and then passes through a shared fully connected layer to obtain S = {S 1 , S 2 , S 3}, where S 1 represents the spatial feature of the first view, S 2 represents the spatial feature of the second view, and S 3 represents the spatial feature of the third view. This process is expressed as:

[0022] S = FC(Conv(O)) (1)

[0023] where FC represents the fully connected layer, and Conv represents a group of convolutional layers with a kernel size of 3×3, 3×3, 1×1, padding of 1, 1, 0, and a stride of 1 for all.

[0024] The STT module (as Figure 5 shown) consists of a multi-head self-attention layer. Passing the output of the MVFE module through a multi-head self-attention layer can promote the fusion between multi-view spatial features, obtaining the fused multi-view spatial features S' = {S′ 1 , S′ 2 , S′ 3}, where S′ 1 represents the spatial feature of the first view after fusion, S′ 2 represents the spatial feature of the second view after fusion, and S′ 3 represents the spatial feature of the third view after fusion. This process is expressed as:

[0025] S' = FFN(MHA(norm(S))) (2)

[0026] "norm" refers to LayerNorm, "MHA" refers to multi-head self-attention, and "FFN" refers to feed-forward neural network. Due to the differences between the Transformer and CNN architectures, in order to fuse the multi-view spatial features extracted by CNN into the temporal features in the Transformer, bilinear interpolation is adopted to align S′ 1 and S′ 3 to the dimension of S′ 2 respectively, obtaining S″ 1 and S″ 3 . And the aligned S″ 1 and S″ 3 are added to S″ 2 to get the final for fusing with the temporal features.

[0027] The TTS module (as shown in Figure 6 ) is similar to the STT module and only contains one multi-head self-attention layer. First, the temporal feature T is added to S 2 to get TS 2 . Similar to (2), the temporal features contained in TS 2 are diffused to the spatial features of other views through a multi-head self-attention layer, thus effectively fusing the multi-view CNN and Transformer features.

[0028] After H stages of feature fusion, the multi-view spatial features are added to the temporal features after passing through an STT module, and then through a max pooling, average pooling, and a fully connected layer to obtain the final spatio-temporal fusion features with rich semantic information.

[0029] Step 2: Reducing the domain difference based on the domain discriminator

[0030] Due to the significant individual differences in emotions, the cross-subject performance of emotion recognition models usually performs poorly. To reduce the domain difference, a domain discriminator (a two-layer fully connected network) is established to identify the attribution of the features extracted by . A gradient reversal layer (GRL) is added before , and the gradient is multiplied by -λ during the backpropagation process. The interaction between the encoder and the domain discriminator strives to reach a Nash equilibrium. Through adversarial training, D can be confused and is encouraged to extract domain-invariant features, thereby reducing the difference between the source domain and the target domain in the feature space. The data labels belonging to D s are set to 1, the data labels belonging to D t are set to 0, and the domain adversarial loss on the source domain samples and target domain samples is defined as:

[0031]

[0032] represents the weighted average of the results calculated for all possible data x randomly sampled from the dataset D s to define the expected value of the loss function and characterize the average loss of the entire dataset. s which is used to define the expected value of the loss function to characterize the average loss of the entire dataset. represents the weighted average of the results calculated for all possible data x randomly sampled from the dataset D t to define the expected value of the loss function and characterize the average loss of the entire dataset. t which is used to define the expected value of the loss function to characterize the average loss of the entire dataset.

[0033] Step 3: Generate pseudo-labels for target domain samples through the soft pseudo-label module

[0034] Since no class information is provided in the target domain, it is impossible to ensure that the target domain samples have obvious discriminability in the latent space. The main goal of the cross-individual EEG emotion recognition task is to improve the recognition performance of the model on target domain samples. Therefore, it is necessary to assume that the classifier (a one-layer linear mapping layer) can also learn confident decision boundaries on the target domain. For this purpose, the soft pseudo-label module (as shown in Figure 7 ) is proposed to provide additional supervision information for the target domain. Different from traditional hard labels, soft pseudo-labels are represented in the form of probability distributions, reflecting the confidence of the model in the sample belonging to each class, effectively reducing the unnecessary bias caused by incorrect hard pseudo-labels, and helping to improve the stability of model training. Based on the clustering hypothesis, it is considered that samples close in the feature space tend to have the same label. For first, according to the feature vector of use the k-NearestNeighbors algorithm with cosine similarity as the metric method in the feature space to find k samples that belong to D t and are the most similar to represents the i-th target domain sample in the feature space that is most similar to Take the average of the predicted probabilities output by the classifier for the k samples as the predicted probability of Finally, obtain the soft pseudo-label through label sharpening The calculation process of the soft pseudo-label is expressed as:

[0035]

[0036] where k represents the number of samples other than itself found for KNN, set to 5, K represents the number of classes in the dataset, and different settings are used for different datasets. θ is a hyperparameter that controls the softness of the labels, set to 0.5. Therefore, the source domain classification loss and the target domain classification loss calculated using the soft pseudo-labels are as follows:

[0037]

[0038] Equation (7) represents the prediction p s of the source domain sample x s and the true label y s of the source domain sample to calculate the loss. Equation (8) represents the prediction p t of the target domain sample x t and the soft pseudo-label of the target domain sample calculated by Equation (6) to calculate the loss, and an indicator function is used to exclude low-confidence soft pseudo-labels, that is, setting the loss term that does not satisfy the condition max(p t )≥τ to 0. Where represents the weighted average of the results calculated from all possible data (x s ,y s ) randomly sampled from the dataset D s to define the expected value of the loss function to characterize the average loss of the entire dataset. l ce (·) represents the cross-entropy loss, represents the indicator function, and τ represents the confidence threshold, set to 0.85.

[0039] Step 4. Smoothly align the source domain and the target domain through ChannelMix

[0040] When the gap between the source domain and the target domain becomes larger, the quality of the pseudo-labels generated based on the soft pseudo-labels may decline, affecting the generalization performance of the model in the target domain. Therefore, to reduce the difference between the source domain and the target domain, the ChannelMix strategy is proposed. It generates an intermediate domain by sampling data from the source domain and the target domain, thereby smoothly transferring the knowledge learned in the source domain to the target domain.

[0041] Define CM λ (·) as a linear interpolation operation on a pair of samples (x a ,y a ) and (x b ,y b ) to reconstruct a mixed sample (x m ,y m ), where x represents the sample data and y represents the sample label. λ n represents a random sample from the beta distribution Beta(α,α) of (xa , y a ) The mixing ratio of the nth channel in the sample, α is set to 2.

[0042]

[0043] Represents x m The nth channel in Represents x a The nth channel in Represents x b The nth channel in, ⊙ represents the multiplication operation, and C represents the total number of channels.

[0044] Through ChannelMix, intermediate domain data x can be created from samples taken on each channel of the source domain and target domain EEG signals Composed of i , x i Represents the created intermediate domain data, Represents x i The nth channel in. Additionally, by aggregating the mixing ratio λ of each source domain sample channel n The sample-level mixing ratio is obtained And it is used for interpolation to obtain the label y i , y i Represents x i The corresponding label. Finally, a new intermediate domain is constructed through ChannelMix Represents the lth sample And its corresponding label n i Represents the number of intermediate domain samples.

[0045] To smoothly reduce the differences between the source domain and the target domain, two loss terms are utilized in the feature space and the label space respectively, and a domain adversarial loss is used to reduce the inter-domain differences.

[0046] 1) Label space: In the label space, a supervised mixing loss is used. Based on the intermediate domain logit and its corresponding mixed label, the domain differences between the intermediate domain and the source domain and the target domain are measured through cross-entropy loss.

[0047]

[0048] Represents the weighted average of the results calculated from all possible data (x i , y i , y i ) randomly sampled from the dataset D, which is used to define the expected value of the loss function to characterize the average loss of the entire dataset.

[0049] 2) Feature space: Since the pseudo-labels in the target domain are not necessarily correct, only constraining in the label space is not sufficient to eliminate domain differences. Therefore, on the feature space without using the pseudo-labels in the target domain as supervision information, minimize the feature differences between the intermediate domain and the source / target domains to further achieve domain alignment.

[0050] Use cosine similarity as the measure of feature similarity. The similarity between the intermediate domain and the source / target domain features is expressed as:

[0051]

[0052] cos represents cosine similarity. In a mini-batch, represents the nth intermediate domain sample, represents the source domain sample used to create i.e., the nth source domain sample in the mini-batch, represents its corresponding target domain sample, i.e., the nth target domain sample in the mini-batch. Define the feature similarity with all source domain samples having labels is set to 1, otherwise set to 0, i.e., represents the feature similarity label between the nth intermediate domain sample and the source domain sample in the mini-batch, represents the label of the nth source domain sample in the mini-batch, y s represents the labels of all source domain samples in the mini-batch. Since there are no reliable ground-truth labels in the target domain, only maximize the feature similarity between and the target domain sample . The hybrid loss in the feature space is expressed as:

[0053]

[0054]

[0055] N represents the size of the mini-batch, set to 8.

[0056] 3) Hybrid domain adversarial: To force to generate domain-invariant features, set the domain label of the source domain sample to 1 and the domain label of the target domain sample to 0. Combining with the gradient reversal layer, a hybrid domain adversarial loss is used on the intermediate domain, expressed as formula (19):

[0057]

[0058] represents from the dataset Di All possible data x randomly sampled from i The weighted average of the calculated results is used to define the expected value of the loss function to characterize the average loss of the entire dataset. Equation (19) shows that the mixed-domain adversarial loss is calculated through the binary cross-entropy loss function and the degree of adversarial learning is controlled by the sample-level mixing ratio λ.

[0059] Step 5. Data augmentation based on ChannelMix

[0060] To improve the generalization performance of the model, ChannelMix is used on the source domain dataset and the target domain dataset respectively to construct the dataset denotes the i-th sample and its corresponding label and denotes the j-th sample and its corresponding label n is and n it respectively represent the number of samples in D is and D it Using the given augmented EEG data, train the classifier and the domain discriminator as follows:

[0061]

[0062] represents the weighted average of the results calculated from all possible data (x is , y is , y is ) randomly sampled from the dataset D, which is used to define the expected value of the loss function. represents the weighted average of the results calculated from all possible data (x it , y it , y it ) randomly sampled from the dataset D, which is used to define the expected value of the loss function. represents the weighted average of the results calculated from all possible data x is randomly sampled from the dataset D, which is used to define the expected value of the loss function. is represents the weighted average of the results calculated from all possible data x i / randomly sampled from the dataset D, which is used to define the expected value of the loss function. it

[0063] ​​Step 6: End-to-end joint training

[0064] The proposed deep network uses the overall loss L proposed by Equation (27) total for optimization:

[0065]

[0066] L dom , L inter_CM , L intra_CM are obtained from Equations (10), (5), (11), (23), and (26), respectively. γ 1 , γ 2 and γ 3 are dynamic weight coefficients, both initialized to 1. End-to-end collaborative training is performed by optimizing this loss function. The base learning rate is set to 5e-4, the batch size is set to 8, and the training model has a total of 50 epochs. We use SGD with momentum of 0.9 and weight decay of 0.001 as the optimizer. A leave-one-out cross-validation strategy is adopted to prove the effectiveness of the method, that is, using one subject as the target domain, and all the remaining subjects together as the source domain and looping this operation until each subject has been used as the target domain subject once. Then, the average accuracy and standard deviation of all target domain subjects are calculated and used to verify the effectiveness of the method.

Claims

1. An EEG emotion recognition method based on ChannelMix and bidirectional spatiotemporal feature fusion, characterized by: Step 1: Multi-stage bidirectional spatiotemporal fusion feature extraction; The module is divided into two branches, one of which is a simple Transformer with L layers, where L is set to 10. It is divided into H stages for feature fusion, where H is set to 5, and is used to extract the temporal features of EEG signals. The other branch is a multi-view spatial feature extraction module based on CNN, where the multi-view feature pyramid (MVFP) is used to provide multi-view spatial features, and the multi-view feature extraction (MVFE) is used to further extract multi-view spatial features for feature fusion at each stage. The spatial feature integration into the temporal feature module (STT) and the temporal feature integration into the spatial feature module (TTS) are used to integrate the spatiotemporal features of the two branches at different stages; In order to extract the temporal features in the EEG data, an overlapping sliding window with a window size of T is used along the time dimension. T is set to 16 and the step size of the sliding window is set to 8. Therefore, the EEG signal data input to this module is represented as C and B represent the number of channels and frequency bands, respectively. x represents the EEG signal data segment sampled by the sliding window, where C is 62 and B is 5. In order to extract the spatiotemporal features of the EEG signal at the same time, the input x of the module is transformed into As the input of the Transformer branch, it is used to extract temporal features, and CB is 310; according to the spatial position information of the electrode, x is mapped to As the input of the convolution branch, it is used to extract the spatial features of multiple views, and TB is 80; first, for the Transformer branch, x t After inputting into Embedding and adding position embedding, the feature of shape T×M is obtained; at the same time, for the other branch, x s After the MVFP module, feature pyramids O1, O2 and O3 with shapes of L1×M, L2×M and L3×M are obtained, where L1, L2, L3 and M are 64, 16, 4 and 64; Embedding module and MVFP module; secondly, the two branches undergo H stages of feature interaction; in each stage, the multi-view spatial features are first further extracted through the MVFE module, and then the spatial features are integrated into the temporal features through the STT module, and then the Transformer Blocks further strengthen the fused temporal features, and finally the temporal features are integrated into the spatial features through the TTS module; The MVFE module consists of a set of convolutional layers and a fully connected layer; specifically, the input of this module is a set of multi-view features output by the MVFP module. It first extracts the spatial features of different views through different convolutional layers, and then passes through a shared fully connected layer to obtain S = {S1, S2, S3}, where S1 represents the spatial features of the first view, S2 represents the spatial features of the second view, and S3 represents the spatial features of the third view; This process is represented as: S=FC(Conv(O)) (1) Among them, FC represents a fully connected layer, Conv represents a set of convolutional layers with kernel sizes of 3×3, 3×3, 1×1, padding of 1, 1, 0, and stride of 1; The STT module consists of a multi-head self-attention layer; the output of the MVFE module After a multi-head self-attention layer, the fusion of multi-view spatial features can be promoted, and the fused multi-view spatial features S'={S'1,S'2,S'3} are obtained, where S'1 represents the spatial features of the first view after fusion, S'2 represents the spatial features of the second view after fusion, and S'3 represents the spatial features of the third view after fusion; this process is expressed as: S'=FFN(MHA(norm(S))) (2) norm is LayerNorm, MHA is multi-head self-attention, and FFN is a feedforward neural network. Bilinear interpolation is used to align S′1 and S′3 to the dimension of S2, obtaining S″1 and S″3 respectively. Then, the aligned S″1 and S″3 are added to S′2 to obtain the final Used for fusion with time series features; The TTS module is similar to the STT module and only contains one multi-head self-attention layer. First, the temporal features T and S2 are added to obtain TS2, and the temporal features contained in TS2 are diffused to the spatial features of other views through a multi-head self-attention layer, thereby effectively fusing the multi-view CNN and Transformer features. After H stages of feature fusion, the spatial features of multiple views are added to the temporal features after passing through an STT module, and then passed through a maximum pooling, average pooling and a fully connected layer to obtain the final spatiotemporal fusion features; Step 3: Generate pseudo labels for target domain samples through the soft pseudo label module; for Firstly based on The eigenvector of In the feature space, the k-Nearest Neighbors algorithm with cosine similarity as the metric is used to find the neighbors that belong to D t The k most similar samples i∈{1,2,…,k}, Indicates that the i-th and The most similar target domain sample; the average of the predicted probabilities of k samples output by the classifier is taken as The predicted probability Finally, the soft pseudo label is obtained by label sharpening The calculation process of soft false labels is expressed as: Where k represents the number of samples other than itself that KNN looks for, which is set to 5. K represents the number of categories in the dataset, and K uses different settings in different datasets. θ is a hyperparameter that controls the softness of the label, which is set to 0.

5. Therefore, the source domain classification loss and the target domain classification loss calculated using soft pseudo labels are: Formula (7) represents the source domain sample x s The prediction p s and the true label y of the source domain sample s Calculate the loss; Formula (8) represents the target domain sample x t The prediction p t The soft pseudo label of the target domain sample calculated by formula (6) Calculate the loss and use an indicator function to exclude soft false labels with low confidence, that is, not satisfying The loss term of the condition is set to 0; Represents the dataset D s All possible data (x s ,y s ) is used to define the expected value of the loss function to characterize the average loss of the entire data set. All have similar meanings; ce (·) represents the cross entropy loss, represents the indicator function, τ represents the confidence threshold, which is set to 0.85; Step 4: Smoothly align the source domain and the target domain through ChannelMix; CM λ (·) is defined as a pair of samples (x a ,y a ) and (x b ,y b ) is used to reconstruct a mixed sample (x m ,y m ), where x represents sample data and y represents sample label; λ n represents a random sample (x) from the beta distribution Beta(α,α) a ,y a ) The mixing ratio of the nth channel in the sample, α is set to 2; Represents x m The nth channel in Represents x a The nth channel in Represents x b The nth channel in , ⊙ represents the multiplication operation, and C represents the total number of channels; ChannelMix can be used to create a channel-wise sampling of the source and target EEG signals. The intermediate domain data x i ,x i Indicates the created intermediate domain data. Represents x i In addition, by aggregating the mixing ratio λ of each source domain sample channel n Get the sample-level mixing ratio And use it to interpolate the label y i ,y i Represents x i The corresponding label; Finally, a new intermediate domain is constructed through ChannelMix represents the lth sample And its corresponding label n i represents the number of samples in the middle domain; Using two loss terms in feature space and label space and using a domain adversarial loss to reduce the difference between domains; 1) Label space: A supervised hybrid loss is used in the label space to measure the domain difference between the intermediate domain and the source and target domains through cross entropy loss based on the intermediate domain logit and its corresponding hybrid label; 2) Feature space: Since the target domain pseudo-label is not necessarily correct, constraining only in the label space is not enough to eliminate domain differences. Therefore, in the feature space without using the target domain pseudo-label as supervision information, the feature difference between the intermediate domain and the source domain / target domain is minimized to further achieve domain alignment. Cosine similarity is used as a measure of feature similarity; The similarity between the intermediate domain and the source / target domain features is expressed as: cos represents cosine similarity; in a mini-batch, represents the nth middle domain sample, Indicates the use to create The source domain sample is the nth source domain sample in the mini-batch. Represents the corresponding target domain sample, that is, the nth target domain sample in the mini-batch; definition With all The feature similarity of the source domain samples of the label is set to 1, otherwise it is set to 0, that is, represents the feature similarity label between the nth intermediate domain sample and the source domain sample in the mini-batch, represents the label of the nth source domain sample in the mini-batch, y s represents the labels of all source domain samples in the mini-batch; since the target domain has no reliable true labels, only the maximum With the target domain sample The feature similarity between them; the mixed loss in the feature space is expressed as: N represents the size of mini-batch, which is set to 8; 3) Hybrid Domain Confrontation: To force Generate domain invariant features, set the domain label of the source domain sample to 1, and the domain label of the target domain sample to 0; combined with the gradient reversal layer, a mixed domain adversarial loss is used on the intermediate domain, expressed as formula (19): Formula (19) indicates that the mixed domain adversarial loss is calculated by the binary cross entropy loss function and the degree of adversarial learning is controlled by the sample-level mixing ratio λ; Step 5: Data enhancement based on ChannelMix; In the source domain dataset And the target domain dataset Use ChannelMix to construct a data set represents the i-th sample And its corresponding label as well as represents the jth sample And its corresponding label n is and n it Respectively represent D is and D it The number of samples in; Using the given enhanced EEG data, train the classifier and domain discriminator as follows: