A Cross-Individual EEG Emotion Recognition Method Based on Contrastive Learning

Through the cross-individual EEG emotion recognition method based on contrast learning, combined with trial consistency strategy and timing multi-scale convolution-self-attention fusion model, the problem of insufficient generalization ability of EEG emotion recognition in cross-subject scenarios is solved, and higher robustness and recognition effect are achieved.

CN119848630BActive Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510315981.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-27
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

The existing EEG emotion recognition methods perform poorly in cross-subject scenarios, mainly due to the high individualization of EEG signals, the generalization ability of the model on new subjects is limited.

Method used

A cross-individual EEG emotion recognition method based on contrast learning is adopted, and a comparison learning sample is constructed through trial consistency strategies to strengthen emotion-related feature representations, and a time-sequence multi-scale convolution-self-attention fusion model is used to improve the generalization ability of the model.

Benefits of technology

It significantly improves the robustness of emotional recognition across subjects, can more effectively generalize between different subjects, and provides a scalable solution for the practical application of emotional brain-computer interfaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119848630B_ABST
    Figure CN119848630B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-subject EEG emotion recognition method based on contrastive learning. Through a fusion model of temporal multi-scale convolution and self-attention, combined with spatial feature encoding, multi-scale temporal convolution, a dimensional projector, and a spatio-temporal attention enhancement module, this method can efficiently extract and fuse spatio-temporal features across time scales, enhancing the ability to capture short-term and long-term features of EEG signals. To address the challenges in cross-subject emotion recognition, the present invention adopts a trial consistency contrastive learning pre-training strategy, inputs signal segments of multi-scale time windows, defines positive and negative sample pairs according to trial numbers, maps features to the contrast space through a trainable projection head, optimizes feature discriminability based on the contrast loss function, and finally fine-tunes through task adaptability, thereby improving the performance of cross-subject EEG emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of EEG emotion recognition, and specifically relates to a cross-subject EEG emotion recognition method based on contrastive learning. Background Art

[0002] Emotion recognition is one of the key technologies in the fields of human-computer interaction, intelligent health monitoring, etc.; electroencephalogram (EEG), as a physiological signal reflecting brain electrical activity, has become the research focus in the field of emotion computing in recent years due to its high temporal resolution and sensitivity to emotional states. Compared with emotion recognition methods based on non-physiological signals such as facial expressions, speech, or eye movements, EEG can provide more objective emotional information and effectively avoid the interference of deliberate disguise. However, due to the high individuality of EEG signals, cross-subject emotion recognition faces severe challenges in practical applications.

[0003] Currently, EEG emotion recognition methods are mainly divided into two categories: traditional machine learning and deep learning. Traditional machine learning methods rely on manual feature extraction and shallow classifiers (such as support vector machines, random forests, etc.), but it is difficult to fully capture the complex non-linear patterns in EEG signals. Deep learning methods can automatically extract emotion-related features from raw EEG data through end-to-end learning and have made significant progress in emotion classification tasks. For example, convolutional neural networks (CNNs) are good at extracting local temporal and spatial features, recurrent neural networks (RNNs) can model the temporal dependence of EEG signals, and the Transformer architecture is good at capturing long-range dependencies. However, most of these methods rely on within-subject training, that is, training and testing the model on the training data of the same subject, resulting in a significant decline in their performance in cross-subject scenarios.

[0004] The main challenge in cross-subject EEG emotion recognition lies in the individual differences of EEG signals, such as the changes in the distribution of EEG waveforms and the differences in neurophysiological characteristics among different subjects. These individual differences severely restrict the generalization ability of the model on new subjects. To address this issue, researchers have proposed some methods such as Domain Adaptation and Transfer Learning. For example, in the literature [Li Yang, Zheng Wenming, Zong Yuan, Cui Zhen, Zhang Tong, Zhou Xiaoyan. EEG Emotion Recognition Based on Dual-Hemisphere Domain Adversarial Neural Network [J]. IEEE Transactions on Affective Computing, 2018], a method based on adversarial domain adaptation was proposed. By designing a dual-hemisphere neural network model and using domain adversarial training to reduce the distribution differences of EEG signals among different subjects, this method improved the generalization ability of the model in cross-subject emotion recognition tasks. Another example is the literature [Li Jinpeng, Qiu Tianshuang, Shen Yingying, Liu Congli, He Huiguang. Application of Multi-Source Transfer Learning in Cross-Subject EEG Emotion Recognition [J]. IEEE Transactions on Cybernetics, 2020, 50(7): 3281-3293], which proposed a multi-source transfer learning framework. By integrating the EEG data of multiple source subjects to train shared feature representations and combining a small amount of data from the target subject to fine-tune the model, the recognition performance on new subjects was improved. However, these methods usually still rely on the data of the target subject for model adjustment, resulting in difficulties in data acquisition and high computational overhead in practical applications.

[0005] In addition, most existing EEG emotion recognition studies focus on instantaneous emotion classification within a short time window and rarely involve modeling long-term emotion changes. However, long-term emotion recognition is crucial for applications such as mental health monitoring and emotion trend analysis. Due to the time-varying dynamics of EEG signals and the influence of external environments and physiological states, traditional methods face great difficulties in capturing the long-term evolution process of emotions. Summary of the Invention

[0006] In view of the above, the present invention provides a cross-individual EEG emotion recognition method based on contrastive learning, which constructs contrastive learning samples using a trial consistency strategy and improves the generalization ability of the model among different subjects by strengthening emotion-related feature representations during the training process.

[0007] A cross-individual EEG emotion recognition method based on contrastive learning includes the following steps:

[0008] (1) Obtain an emotion EEG dataset, which stimulates the emotional state of the subjects through emotion induction stimuli and synchronously collects the EEG signals of the subjects and their corresponding emotion labels;

[0009] (2) Perform standardized preprocessing on the EEG signals in the dataset, and then divide the dataset into a training set and a test set;

[0010] (3) Construct a temporal multi-scale convolutional-self attention fusion model, which includes:

[0011] A feature extraction backbone network for extracting multi-resolution features in the spatio-temporal domain from the input EEG signals and mapping the extracted low-dimensional features to high-dimensional features;

[0012] A spatio-temporal attention enhancement module for constructing the context dependence of the spatio-temporal information in the high-dimensional features to generate context-aware enhanced features;

[0013] (4) Conduct trial-consistency contrastive learning pre-training on the above fusion model;

[0014] (5) Connect a classification module to the output end of the pre-trained fusion model and fine-tune the fusion model connected to the classification module using the training set to obtain a cross-subject EEG emotion recognition model;

[0015] (6) Input the EEG signals in the test set into the cross-subject EEG emotion recognition model, and the corresponding emotion categories can be recognized and output.

[0016] Further, the emotion EEG dataset obtained in step (1) comes from 2 different research institutions, namely the two publicly available datasets SEED and FACED. Each dataset contains EEG signals from multiple different subjects and their corresponding emotion labels. The EEG signals in the SEED dataset are 62 channels, and the EEG signals in the FACED dataset are 32 channels.

[0017] Further, the specific implementation method of the normalization preprocessing in step (2) is as follows: for the EEG signals in the SEED dataset, first downsample the signals to 200 Hz, then use a band-pass filter to limit the signal frequency range to 0.3 - 49 Hz, then remove the artifact interference including eye movement and muscle movement in the signals through the ICA (Independent Component Analysis) algorithm, and finally apply the CAR (Common Average Reference) method to complete the normalization preprocessing of the signals; for the EEG signals in the FACED dataset, complete the normalization preprocessing after downsampling the signals to 250 Hz and performing preliminary filtering and artifact removal.

[0018] Further, the specific implementation method for dividing the dataset in step (2) is as follows: For the SEED dataset, a leave-one-out cross-validation strategy is adopted: The EEG signals of 15 subjects are respectively divided into 15 independent groups. Each time, one group is selected as the test set, and the remaining 14 groups are used as the training set. For the FACED dataset, a 10-fold cross-validation strategy is adopted: The EEG signals of all subjects are evenly divided into 10 groups. Each time, one group is selected as the test set, and the remaining 9 groups are used as the training set, and it is ensured that the EEG signals of the same subject will not appear in the training set and the test set at the same time.

[0019] Further, the feature extraction backbone network includes:

[0020] A spatial feature encoding module, which is used to extract spatial domain features of the input EEG signal along the electrode dimension and capture the spatial dependence relationship between different electrodes;

[0021] A multi-scale temporal convolutional module, which is used to extract multi-resolution features of the output of the spatial feature encoding module along the time dimension;

[0022] A high-dimensional feature projection module, which is used to project the low-dimensional feature tensor output by the multi-scale temporal convolutional module into a high-dimensional hidden space to enhance the discriminative expression ability of the features.

[0023] Further, the spatial feature encoding module outputs the input EEG signal after passing through a two-dimensional convolutional neural network and a group normalization layer in sequence; the multi-scale temporal convolutional module inputs the output of the spatial feature encoding module into multiple parallel temporal convolutional neural networks respectively, and then splices the output results of each temporal convolutional neural network and processes them through an ELU (Exponential Linear Unit) activation function before outputting; the high-dimensional feature projection module first performs average pooling on the low-dimensional feature tensor output by the multi-scale temporal convolutional module, then constructs a non-linear mapping function using a three-dimensional convolutional kernel with a scale of 1 to map the pooled features into high-dimensional features, and finally processes them through an ELU activation function before outputting.

[0024] Further, the spatio-temporal attention enhancement module consists of layer normalization L1, a multi-head self-attention mechanism layer, layer normalization L2, and an MLP (Multilayer Perceptron) connected in sequence from input to output. Among them, the input of L1 is added to the output of the multi-head self-attention mechanism layer as the input of L2, and the input of L2 is added to the output of the MLP as the output of the spatio-temporal attention enhancement module.

[0025] Further, the specific implementation method of step (4) is as follows:

[0026] 4.1 For the EEG signals from different subjects and different emotion induction trials in the training set, a number of EEG segments are obtained by intercepting the EEG signals through a sliding time window at multiple different window scales. All combinations of EEG segments in the training set are traversed to generate a large number of sample pairs. Each sample pair contains two EEG segments at the same window scale. If these two EEG segments come from the same emotion induction trial, the sample pair is a positive sample pair; otherwise, it is a negative sample pair.

[0027] 4.2 Connect the output end of the temporal multi-scale convolutional-self attention fusion model to a non-linear projection network, which is composed of two connected MLPs.

[0028] 4.3 Input the sample pairs into the fusion model connected to the non-linear projection network in batches for pre-training, and use the InfoNCE (Information Noise Contrastive Estimation) loss function to iteratively optimize the model parameters by gradient descent.

[0029] Furthermore, the classification module connected in the step (5) is composed of a fully connected network and a Softmax activation function. During the process of fine-tuning the task adaptability of the fusion model, the cross-entropy loss function and the adaptive moment estimation optimizer are used to perform end-to-end training on the classification module at a fixed learning rate.

[0030] Based on the above technical solutions, the present invention innovatively solves two core challenges in cross-subject EEG emotion recognition through a trial consistency contrast learning framework and a multi-scale spatio-temporal feature fusion mechanism: First, through the trial consistency contrast pre-training strategy, the differences between subjects are effectively eliminated, and a unified feature space with cross-subject generalization ability is constructed; Second, by adopting a hybrid architecture of multi-scale temporal convolution and self-attention model, combined with the synergistic effect of local receptive fields and temporal global attention mechanisms, the continuous evolution modeling of emotional states across time scales is realized. The present invention not only significantly improves the cross-subject robustness of emotion recognition, but also provides an extensible solution for the actual application scenarios of emotional brain-computer interfaces through an end-to-end deep learning architecture. Description of the Drawings

[0031] Figure 1 It is a schematic flowchart of the cross-individual EEG emotion recognition method based on contrast learning of the present invention.

[0032] Figure 2 It is a schematic structural diagram of the temporal multi-scale convolutional-self attention fusion model of the present invention.

[0033] Figure 3 It is a schematic flowchart of the model pre-training, fine-tuning and testing based on trial consistency contrast learning of the present invention. Specific Embodiments

[0034] To describe the present invention more specifically, the technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0035] As Figure 1 shown, the cross - individual EEG emotion recognition method based on contrastive learning of the present invention includes the following steps:

[0036] (1) Obtain a standardized emotion EEG dataset. The dataset stimulates the emotional states of subjects through emotion - inducing stimuli (such as videos or audios), and synchronously collects multi - lead EEG signals and corresponding emotion labels.

[0037] In this embodiment, the datasets come from two different research institutions, namely SEED and FACED. Each dataset contains multi - lead EEG signals from different subjects and corresponding emotion label information. The SEED dataset was developed by the BCMI Laboratory of Shanghai Jiao Tong University. It uses 15 standardized movie clips as standardized stimulus materials, recruits 15 healthy subjects, and collects signals through a 62 - lead EEG system. The collected labeled data includes three emotion dimensions: positive, neutral, and negative. The FACED dataset was released by the Department of Psychology of Tsinghua University. It induces 9 types of emotions based on 28 standardized video stimuli, covering 32 - lead EEG records of 123 subjects. Its labels include 4 types of positive emotions (pleasure, inspiration, happiness, tenderness), 4 types of negative emotions (anger, nausea, fear, sadness), and neutral states.

[0038] (2) Perform standardized pre - processing on the original EEG signals in the dataset to generate normalized time - series data suitable for input to a deep neural network.

[0039] In this embodiment, for the SEED dataset, data containing 62 standard EEG recording channels is used. First, the signal is downsampled to 200 Hz, then the signal frequency range is restricted to 0.3 - 49 Hz using a band - pass filter. Subsequently, artifacts such as eye movements and muscle movements are removed through an independent component analysis algorithm. Finally, the standardized pre - processing is completed using the independent component value weighted average reference method. For the FACED dataset, data containing 32 standard EEG recording channels is used, which includes EEG signals that have been preliminarily filtered and artifact - removed at 250 Hz.

[0040] Next, divide the training set and the test set to ensure that the same training set is used in the pre-training and fine-tuning phases while keeping the test set invisible. For the SEED dataset, adopt the leave-one-subject cross-validation strategy, divide the 15 subjects into 14 independent groups, select 1 group as the test set for each experiment, and the remaining 14 groups as the training set. For the FACED dataset, adopt the 10-fold cross-validation strategy, evenly divide the subjects into 10 groups, select 1 group as the test set for each experiment, and the remaining 9 groups as the training set. The EEG data of the same subject will not appear in both the training set and the test set at the same time, and it is necessary to ensure the balance of the EEG data quantity of each group.

[0041] (3) Construct a temporal multi-scale convolutional-self attention fusion model, as Figure 2 shown, its architecture includes:

[0042] Spatial feature encoding module, using a two-dimensional convolutional neural network to extract spatial domain features along the electrode dimension, capturing the spatial dependence relationship between different electrodes, and the convolutional scale is the number of channels of the input signal. The specific expression is as follows:

[0043]

[0044] Among them: X represents the input signal, Conv S represents the spatial convolutional network, GN represents the group normalization layer;

[0045] Multi-scale temporal convolutional module, which consists of multiple parallel temporal convolutional branches. Each branch is configured with convolutional kernels of different time scales for multi-resolution feature extraction. The specific expression is as follows:

[0046]

[0047] Among them: represents the i th convolutional network of the multi-scale temporal module network, Concat represents parallel concatenation, E represents the ELU activation function.

[0048] High-dimensional feature projection module, using a three-dimensional convolutional kernel with a scale of 1 to construct a non-linear mapping function, projecting the low-dimensional feature tensor output by the previous module into a high-dimensional hidden space to enhance the discriminative expression ability of the features. The specific expression is as follows:

[0049]

[0050] Among them: Conv P represents the three-dimensional convolutional feature projector, AvgpoolIndicates average pooling.

[0051] The spatio-temporal attention enhancement module, based on the multi-head self-attention mechanism (Multi-Head Attention), realizes the dynamic modeling of the global spatio-temporal feature relationship by calculating the spatio-temporal joint attention weight matrix, and introduces a relative position encoding strategy to retain the temporal dependence features of bioelectrical signals. The specific expression is as follows:

[0052]

[0053] Where: Attention Represents the multi-head self-attention mechanism, Q , K and V Represent the query, key, and value matrices respectively, d K Is the dimension of the key.

[0054] (4)Design trial consistency contrastive learning pre-training. As Figure 3 shown, map the features to the contrast space through a trainable projection head, and optimize the feature discriminability based on the contrast loss function.

[0055] First, for each emotion induction trial, intercept several EEG segments through a sliding time window. The multi-scale time window spans , and the time windows corresponding to the three spans are Len1, Len2, and Len3 respectively; define sample pairs according to the trial numbers to which the samples belong ( X , X' ). Samples from the same trial are positive sample pairs, and samples from different trials are negative sample pairs.

[0056] Next, construct a trainable non-linear projection head network, which consists of two layers of MLP:

[0057]

[0058] Where: H Represents the feature vector obtained by the input signal passing through the fusion model, W 1 and W 2 are learnable parameter matrices, b 1 and b 2 are bias terms, σ Is the RELU activation function, and the output projection feature is Z .

[0059] Then, combine the multi-scale convolution-self-attention fusion model with the non-linear projection head to get:

[0060]

[0061] Where:M denotes the fusion model, g is the non - linear projection head.

[0062] Finally, the sample pairs are input into the combined model batch - by - batch to obtain the corresponding projected features , and the InfoNCE loss function is used to optimize the representation of the projected features output by the model in the feature space. The loss function takes the average value under inputs of different time - window lengths to obtain the joint loss L con :

[0063]

[0064]

[0065]

[0066] where: l represents time windows of different lengths, is the cosine similarity of the sample pairs, in which the numerator is the similarity of positive sample pairs in the batch, and the denominator is the sum of the similarities of all negative sample pairs in the batch, is the temperature hyper - parameter, N is the number of sample pairs in a batch.

[0067] (5) Fine - tune the pre - trained model for the sentiment classification task adaptability.

[0068] Based on the pre - trained fusion model, a learnable classification head is connected for fine - tuning of the sentiment classification task; in this embodiment, the classification head is implemented by connecting a fully - connected network and a Softmax activation function. The fully - connected network is composed of three - layer MLP and can be expressed by the following formula:

[0069]

[0070] where: h represents the feature vector obtained from the pre - trained fusion model, W 1, W 2, W 3 are learnable weight matrices, b 1, b 2, b 3 are bias terms, is the probability distribution of the sentiment categories output by the model, and finally the corresponding category is selected by taking the maximum value.

[0071] In the parameter optimization stage, this embodiment uses the cross - entropy loss function L clsAnd an adaptive moment estimation optimizer to perform end-to-end training on the classification layer at a fixed learning rate, and adopt a leave-one-subject cross-validation strategy to optimize the emotional category decision boundary through supervised learning.

[0072] (6) Deploy an emotion inference system across subjects.

[0073] Use the trained cross-subject emotion recognition model to predict the sleep stage on the test set data corresponding to each fold, calculate the prediction accuracy (Accuracy) and Macro-F1 score, and count the average value of the results of all folds and compare it with other existing methods. The comparison results are shown in Table 1. It can be seen from the table that the method of the present invention has a significant improvement in the recognition effect compared with other existing recognition methods.

[0074] Table 1

[0075]

[0076] The above description of the embodiments is to facilitate the understanding and application of the present invention by those of ordinary skill in the art. It is obvious that those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made by those skilled in the art to the present invention should be within the protection scope of the present invention according to the disclosure of the present invention.

Claims

1. A cross-individual EEG emotion recognition method based on contrastive learning, comprising the following steps: (1) Obtain an emotional EEG dataset, which stimulates the emotional state of the subject through emotion-inducing stimulation and synchronously collects the subject's EEG signal and its corresponding emotion label; (2) Standardize and preprocess the EEG signals in the dataset, and then divide the dataset into a training set and a test set; (3) Construct a temporal multi-scale convolution-self-attention fusion model, which includes: The feature extraction backbone network is used to extract multi-resolution features of the input EEG signal in the spatiotemporal domain and map the extracted low-dimensional features into high-dimensional features; A spatiotemporal attention enhancement module, used to construct contextual dependencies of spatiotemporal information in the high-dimensional features and generate context-aware enhanced features; The feature extraction backbone network comprises: The spatial feature encoding module is used to extract spatial domain features of the input EEG signal along the electrode dimension and capture the spatial dependencies between different electrodes; The multi-scale temporal convolution module is used to extract multi-resolution features along the time dimension from the output of the spatial feature encoding module; The high-dimensional feature projection module is used to project the low-dimensional feature tensor output by the multi-scale temporal convolution module into a high-dimensional latent space to enhance the discriminative expression ability of the feature; The spatial feature encoding module sequentially passes the input EEG signal through a two-dimensional convolutional neural network and a group normalization layer and then outputs it; the multi-scale temporal convolution module inputs the output of the spatial feature encoding module into multiple parallel temporal convolutional neural networks respectively, and then splices the output results of each temporal convolutional neural network and then processes it through the ELU activation function and then outputs it; the high-dimensional feature projection module first averages the low-dimensional feature tensor output by the multi-scale temporal convolution module, and then uses a three-dimensional convolution kernel with a scale of 1 to construct a nonlinear mapping function to map the pooled features to high-dimensional features, and finally processes it through the ELU activation function and then outputs it; The spatiotemporal attention enhancement module is composed of layer normalization L1, multi-head self-attention mechanism layer, layer normalization L2, and MLP connection from input to output, wherein the input of L1 is added to the output of the multi-head self-attention mechanism layer as the input of L2, and the input of L2 is added to the output of MLP as the output of the spatiotemporal attention enhancement module; (4) Pre-training the above fusion model with trial-to-trial consistency contrastive learning. The specific implementation method is as follows: 4.1 For the EEG signals from different emotion-induced trials of different subjects in the training set, the EEG signals are intercepted by sliding time windows at multiple different window scales to obtain several EEG segments, and all EEG segment combinations in the training set are traversed to generate a large number of sample pairs, wherein the sample pairs contain two EEG segments of the same window scale. If the two EEG segments come from the same emotion-induced trial, the sample pair is a positive sample pair, otherwise it is a negative sample pair; 4.2 Connect the output of the temporal multi-scale convolution-self-attention fusion model to a nonlinear projection network consisting of two MLP connections; 4.3 Input the sample pairs in batches into the fusion model connected to the nonlinear projection network for pre-training, and use the InfoNCE loss function to iteratively optimize the model parameters through the gradient descent method; (5) Connecting the classification module to the output of the pre-trained fusion model, and using the training set to fine-tune the task adaptability of the fusion model connected to the classification module, thereby obtaining a cross-individual EEG emotion recognition model; (6) The EEG signals in the test set are input into the cross-individual EEG emotion recognition model, which can identify the output and obtain the corresponding emotion category.

2. The method for cross-individual EEG emotion recognition based on contrastive learning according to claim 1, characterized in that: The emotional EEG datasets obtained in step (1) come from two different research institutions, namely, SEED and FACED, two public datasets. Each dataset contains EEG signals from multiple different subjects and their corresponding emotional labels. The EEG signal in the SEED dataset has 62 channels, and the EEG signal in the FACED dataset has 32 channels.

3. The method for cross-individual EEG emotion recognition based on contrastive learning according to claim 2, characterized in that: The specific implementation method of the standardized preprocessing in step (2) is as follows: for the EEG signal in the SEED data set, the signal is first downsampled to 200 Hz, and then a bandpass filter is used to limit the signal frequency range to 0.3-49 Hz, and then the ICA algorithm is used to remove artifact interference including eye movement and muscle movement in the signal, and finally the common mean reference method is used to complete the standardized preprocessing of the signal; for the EEG signal in the FACED data set, the signal is downsampled to 250 Hz and then subjected to preliminary filtering and artifact removal to complete the standardized preprocessing.

4. The method for cross-individual EEG emotion recognition based on contrastive learning according to claim 2, characterized in that: The specific implementation method of dividing the data set in step (2) is as follows: for the SEED data set, a leave-one-out cross-validation strategy is adopted: the EEG signals of the 15 subjects are divided into 15 independent groups, one group is selected as the test set each time, and the remaining 14 groups are used as training sets; for the FACED data set, a 10-fold cross-validation strategy is adopted: the EEG signals of all subjects are equally divided into 10 groups, one group is selected as the test set each time, and the remaining 9 groups are used as training sets, and it is ensured that the EEG signal of the same subject does not appear in both the training set and the test set.

5. The method for cross-individual EEG emotion recognition based on contrastive learning according to claim 1, characterized in that: The classification module connected in the step (5) is composed of a fully connected network and a Softmax activation function. In the process of fine-tuning the task adaptability of the fusion model, a cross entropy loss function and an adaptive moment estimation optimizer are used to perform end-to-end training on the classification module at a fixed learning rate.

Citation Information

Patent Citations

  • Dynamic nerve-muscle rehabilitation exercise method and system, electronic equipment and storage medium

    CN118680579A