Emotion recognition method and system based on generative self-supervised learning and EEG signals

Through generative self-supervised learning and multi-view mask auto-encoding model, the EEG signal is reconstructed and personalized training is solved, and the problem of the impact of EEG data noise in daily environments is achieved, and high-precision emotion recognition is achieved.

CN115590515BActive Publication Date: 2025-08-19SHANGHAI ZERO UNIQUE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211194404.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2025-08-19
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

EEG data collected in daily environments are susceptible to user and environmental noise, and it is difficult to train high-precision emotion recognition models.

Method used

Generative self-supervised learning method is adopted, and the differential entropy features are reconstructed in frequency domain, spatial and temporal dimensions using a multi-view masked self-coding model. A general feature extractor is obtained through pre-training, and emotional prediction is made based on the calibrated EEG signal of the target subjects.

Benefits of technology

Effective use of unlabeled and damaged EEG signals improves the accuracy and robustness of emotion recognition, and solves the problem of emotion decoding under a small amount of markings and damaged data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115590515B_ABST
    Figure CN115590515B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides an emotion recognition method and system based on generative self-supervised learning and EEG signals. The method includes: inputting the differential entropy features used to reflect the EEG signals of the subject into a multi-view masked autoencoder model, reconstructing the differential entropy features, pre-training the codec of the multi-view masked autoencoder model, and using the encoder as a universal feature extractor for EEG signals; performing personalized training on the universal feature extractor based on the calibrated EEG signals of the target subject and the baseline emotion labels to obtain an emotion predictor for self-supervised learning of the target subject; and performing personalized emotion prediction on the collected EEG data of the target subject based on the emotion predictor. The embodiment of the present invention uses the reconstruction of the masked EEG channel as a proxy task in the pre-training stage, mines the information of the unlabeled data and gives the model the ability to decode damaged EEG data, solving the problem of decoding emotions from a small amount of labeled and damaged EEG data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of emotional brain-computer interface technology, and in particular to an emotion recognition method and system based on generative self-supervised learning and electroencephalogram (EEG) signals. Background Art

[0002] Emotion recognition plays a crucial role in emotional brain-computer interfaces and mental health assessment. For example, many affective disorders are linked to emotions, and accurately assessing a patient's emotional state can aid in their treatment. Similarly, in interactions with intelligent assistants, accurately identifying a user's emotions can also enable the delivery of more personalized information and feedback, enhancing the user experience.

[0003] Although there are many ways to identify emotions, such as facial expressions, eye movements, skin conductance response, electrocardiogram and electroencephalogram, the use of electroencephalogram signals can reveal subtle changes in emotions with higher temporal resolution, making it more objective and accurate when analyzing emotional states.

[0004] With the rapid development of EEG emotion recognition technology, researchers can successfully decode labeled, high-quality EEG data collected in laboratory settings. Emotion recognition models are trained using this labeled, high-quality EEG data to accurately assess people's emotions.

[0005] In the process of implementing the present invention, the inventors discovered that there are at least the following problems in the related art:

[0006] EEG data annotation is time-consuming and labor-intensive, and collecting high-quality EEG data in large-scale lab environments is often difficult. This limits the collection of high-quality EEG data. While devices like portable dry-electrode EEG can be used to collect EEG data in everyday environments, these environments are susceptible to noise interference, and EEG signals are sensitive to noise. EEG data collected in everyday environments is easily corrupted by both the user and the environment, making it difficult to train high-precision emotion recognition models. Summary of the Invention

[0007] In order to at least solve the problem in the prior art that EEG data collected in daily environments are easily damaged by the user and the environment, making it difficult to train a high-precision emotion recognition model. In a first aspect, an embodiment of the present invention provides an emotion recognition method based on generative self-supervised learning and EEG signals, comprising:

[0008] Inputting differential entropy features used to reflect the subject's EEG signal into a multi-view masked autoencoder model, reconstructing the differential entropy features in the frequency domain and / or spatial and / or temporal dimensions to obtain multi-view reconstructed differential entropy features for simulating unlabeled and / or damaged EEG signals, pre-training an encoder and decoder of the multi-view masked autoencoder model based on the multi-view reconstructed differential entropy features, and using the obtained encoder as a universal feature extractor for the EEG signal;

[0009] Performing personalized training on the universal feature extractor based on the target subject's calibrated EEG signal and the baseline emotion label corresponding to the calibrated EEG signal to obtain an emotion predictor for the target subject through self-supervised learning;

[0010] Based on the emotion predictor, personalized emotion prediction is performed on the collected EEG data of the target subject, wherein the EEG data includes: unlabeled EEG signals and damaged EEG signals.

[0011] In a second aspect, an embodiment of the present invention provides an emotion recognition system, comprising:

[0012] A general feature program module is used to input differential entropy features used to reflect the subject's EEG signal into a multi-view masked autoencoder model, reconstruct the differential entropy features in the frequency domain and / or spatial and / or temporal dimensions to obtain multi-view reconstructed differential entropy features for simulating unlabeled and / or damaged EEG signals, pre-train an encoder and decoder of the multi-view masked autoencoder model based on the multi-view reconstructed differential entropy features, and use the obtained encoder as a general feature extractor for the EEG signal;

[0013] a personalized training program module for performing personalized training on the universal feature extractor based on the target subject's calibrated EEG signal and the reference emotion label corresponding to the calibrated EEG signal, to obtain an emotion predictor for the target subject through self-supervised learning;

[0014] The emotion recognition program module is used to perform personalized emotion prediction on the collected EEG data of the target subject based on the emotion predictor, wherein the EEG data includes: unlabeled EEG signals and damaged EEG signals.

[0015] According to a third aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the emotion recognition method based on generative self-supervised learning and electroencephalogram signals of any embodiment of the present invention.

[0016] In a fourth aspect, an embodiment of the present invention provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the emotion recognition method based on generative self-supervised learning and EEG signals of any embodiment of the present invention are implemented.

[0017] The beneficial effects of the embodiments of the present invention are: using the reconstruction of masked EEG channels as a proxy task in the pre-training stage, fully mining the information of unlabeled data and giving the model the ability to decode a small amount of labeled and damaged EEG data; the hybrid structure based on CNN-Transformer fully utilizes the information of the spectrum, time and space domains of EEG signals, and then solves the problem of decoding emotions from a small amount of labeled and damaged EEG data through generative self-supervised learning of reconstructing masked EEG channels. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 This is a flowchart of an emotion recognition method based on generative self-supervised learning and EEG signals provided by one embodiment of the present invention;

[0020] Figure 2 1. This is a schematic diagram of a self-supervised learning model architecture of a multi-view masked autoencoder for an emotion recognition method based on generative self-supervised learning and EEG signals provided by one embodiment of the present invention;

[0021] Figure 3 1 is a schematic diagram of an EEG channel of an emotion recognition method based on generative self-supervised learning and EEG signals provided by an embodiment of the present invention;

[0022] Figure 4 1 is a schematic diagram of average accuracy and standard deviation of an emotion recognition method based on generative self-supervised learning and EEG signals using all labeled training data provided by one embodiment of the present invention;

[0023] Figure 5 This is a schematic diagram of average accuracy and standard deviation of an emotion recognition method based on generative self-supervised learning and EEG signals provided by one embodiment of the present invention when using a small amount of labeled training data;

[0024] Figure 6 Schematic diagram of damaged labeled training data for an emotion recognition method based on generative self-supervised learning and EEG signals provided by one embodiment of the present invention;

[0025] Figure 7 Schematic diagram of ablation research on hierarchical performance of an emotion recognition method based on generative self-supervised learning and EEG signals provided by one embodiment of the present invention;

[0026] Figure 8 Schematic diagram of a confusion matrix for calibrating three or four emotional state distinctions using a small amount of or all labeled data for an emotion recognition method based on generative self-supervised learning and EEG signals, provided by one embodiment of the present invention;

[0027] Figure 9 This is a schematic diagram of the reconstruction visualization of the test data of an emotion recognition method based on generative self-supervised learning and EEG signals provided by one embodiment of the present invention after masking and damaging the EEG channel at different masking rates;

[0028] Figure 10 1 is a schematic diagram of the structure of an emotion recognition system based on generative self-supervised learning and EEG signals provided by one embodiment of the present invention;

[0029] Figure 11 A schematic structural diagram of an embodiment of an electronic device for emotion recognition based on generative self-supervised learning and EEG signals provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0031] like Figure 1 FIG2 is a flowchart of an emotion recognition method based on generative self-supervised learning and EEG signals provided by an embodiment of the present invention, comprising the following steps:

[0032] S11: Inputting the differential entropy features used to reflect the subject's EEG signals into a multi-view masked autoencoder model, reconstructing the differential entropy features in the frequency domain and / or spatial and / or temporal dimensions to obtain multi-view reconstructed differential entropy features for simulating unlabeled and / or damaged EEG signals, pre-training the encoder and decoder of the multi-view masked autoencoder model based on the multi-view reconstructed differential entropy features, and using the obtained encoder as a universal feature extractor for the EEG signals;

[0033] S12: performing personalized training on the universal feature extractor based on the calibrated EEG signal of the target subject and the reference emotion label corresponding to the calibrated EEG signal to obtain an emotion predictor for the target subject through self-supervised learning;

[0034] S13: Performing personalized emotion prediction on the collected EEG data of the target subject based on the emotion predictor, wherein the EEG data includes: unlabeled EEG signals and damaged EEG signals.

[0035] In this embodiment, the prepared emotion-inducing materials are provided for the subjects to watch. In daily scenarios, as long as the subjects have an EEG acquisition device, the EEG signal data of the subjects can be collected through the EEG acquisition device. Specifically, the EEG acquisition device can use an ESI NeuroScan wet electrode EEG cap to collect EEG signals. In this way, a large number of subjects' EEG signals can be obtained in daily scenarios as training data for emotion recognition model training, and then used for emotion recognition in various fields. For example, a daily scenario can be that the subject sits quietly in a room, and a display is placed in front of the subject to play an emotion-inducing video for the subject to watch. At this time, the EEG signal is collected based on the EEG acquisition device.

[0036] In step S11, after obtaining the EEG signals of the subjects, considering that these EEG signals are derived from daily scenes, in order to further improve the accuracy, the EEG signals of the subjects need to be pre-processed by filtering and noise reduction.

[0037] In this embodiment, preprocessing includes performing baseline correction, artifact removal, filtering, and other processing on the collected EEG data. Specifically, 50 Hz AC power noise is removed from the EEG signal; after denoising, a 1-75 Hz bandpass filter is used to remove low-frequency and high-frequency invalid signals from the EEG signal.

[0038] By determining the corresponding differential entropy features from the preprocessed EEG signal, a fixed-length Hanning window can be used to perform a fast Fourier transform on the EEG signal. The EEG signal's frequency spectrum in the frequency domain is then used to extract differential entropy features that reflect the energy of different frequency bands in the EEG signal. After obtaining the differential entropy features, a linear dynamic system smoothing process is performed. This results in a differential entropy feature that reflects the subject's EEG signal.

[0039] For the multi-view masked autoencoder model, it is based on a CNN (Convolutional Neural Network)-Transformer hybrid structure, which decodes emotion-related knowledge of EEG signals from spectral, spatial and temporal perspectives.

[0040] Spectral analysis can be performed in the frequency domain. According to Fourier's theorem, any continuously measured time series or signal can be represented as an infinite superposition of sinusoidal signals of varying frequencies. EEG signals can be viewed as a mixture of different sinusoidal signals. Using the Fourier transform, this mixture can be decomposed into sinusoidal waves of varying frequencies, thereby obtaining information in the frequency domain. Frequency domain analysis can be used not only to analyze task-based data but is also commonly used to analyze resting-state data. The time domain, from a temporal perspective, focuses on how EEG signal amplitude changes over time, enabling rapid analysis of amplitude changes caused by a specific event (stimulus). The spatial perspective extracts the dynamics of EEG signal channels and the dependencies between channels, objectively capturing the overall brain amplitude in response to a specific stimulus.

[0041] By reconstructing differential entropy features from multiple perspectives in the frequency, space, and time dimensions, we can obtain a large number of multi-perspective reconstructed differential entropy features for simulating unlabeled and / or damaged EEG signals. These multi-perspective reconstructed differential entropy features can be used to train a universal feature extractor, which can extract emotional features that may appear in all people. Since different people express emotions in different ways, their corresponding EEG signals are also different, and individual, detailed emotions may not be recognized.

[0042] Specifically, the multi-view masked autoencoder model consists of a spectrum embedding layer, a spatial position encoding layer, an EEG channel mask layer, a hybrid encoding block, and a hybrid decoding block symmetrical to the hybrid encoding block, wherein the spectrum embedding layer is used to extract spectrum information of differential entropy features;

[0043] The spatial position encoding layer is used to encode the spatial position of the EEG channel for destruction and reconstruction, wherein the encoding method of the spatial position of the EEG channel includes sine-cosine position encoding;

[0044] The EEG channel mask layer is used to divide the reconstructed differential entropy features into a visible subset and a mask subset according to the channel;

[0045] The hybrid coding block is used to capture the dependencies between EEG channels in the visible subset and determine the multi-view fusion features of EEG;

[0046] The hybrid decoding block is used to determine the original EEG features through a visible subset and a mask subset replaced by parameters, and pre-train the hybrid encoding block and the hybrid decoding block based on the reconstruction loss of the original EEG features and the reconstructed EEG features output by the decoder to obtain a universal feature extractor for the EEG signal.

[0047] In this embodiment, if Figure 2The figure shows the structure of the multi-view masked autoencoder model, which consists of a spectrum embedding layer, a spatial position encoding layer, an EEG channel mask, L CNN-Transformer hybrid encoding blocks, and L symmetric CNN-Transformer hybrid decoding blocks. The input consists of the frequency domain differential entropy features of the training EEG data from all subjects, which can be expressed as X = (x1, x2, ...x N ,)∈R N×F×V To obtain the EEG time series, the extracted frequency domain feature X is converted into

[0048] For the spectrum embedding layer, it projects the input EEG spectrum features (that is, differential entropy features) into the new D-dimensional spectrum space through linear transformation, embedding the spectrum information of the EEG signal. It is expressed as Where W and b are the weight matrix and bias of the linear transformation.

[0049] The spatial position encoding layer divides the EEG data into blocks based on their spatial channels, with each block representing a channel. To remember the location of each channel and facilitate reconstruction after damage, sine and cosine position encoding is added to the spatial dimensions of the EEG channels.

[0050] For EEG channel mask, EEG data is randomly divided into a visible subset by channel. and a mask subset Only the visible subset is used as the input of the hybrid encoder, which is expressed as:

[0051] For L CNN-Transformer hybrid encoding blocks, each CNN-Transformer hybrid encoding block includes a multi-scale temporal causal convolution layer, a multi-head spatial self-attention layer, a normalization layer, and a feedforward network layer.

[0052] Specifically, the temporal causal convolution layer includes a multi-scale temporal convolution kernel for extracting temporal information of EEG channels in the visible subset;

[0053] The multi-head spatial self-attention layer is used to capture the spatial dependencies between EEG channels of the visible subsets after division.

[0054] In this embodiment, the multi-scale temporal causal convolution layer includes causal convolution layer branches with long, medium and short scale convolution kernels. For each scale branch, the time dimension T of each EEG channel is convolved with a kernel size of K. l ×1,K m ×1,K s×1 causal convolution operation and batch normalization (BN), the embedded features are updated by the adjacent time of the same channel, and the updated features are expressed as:

[0055]

[0056]

[0057]

[0058] Among them, B in is the input of the multi-scale temporal causal convolutional layer, are the encoding feature outputs of the long, medium, and short scale causal convolutional layers, respectively. For the first layer of CNN-Transformer hybrid encoding block, B in Visible subset The input B of each subsequent CNN-Transformer hybrid encoding block in is the output of the encoding block in the previous layer.

[0059] After the temporal convolution layer, the multi-head spatial self-attention layer captures the dependencies between EEG channels of all visible subsets, and then fuses the EEG feature embeddings of the three scale branches through summation operation. The fused features Expressed as:

[0060]

[0061]

[0062]

[0063]

[0064] Thus, the multi-view fusion characteristics of EEG are determined.

[0065] For L symmetric CNN-Transformer hybrid decoding blocks, it consists of L CNN-Transformer hybrid decoding blocks and linear layers with the same structure as the encoding block. The input of the decoder is the visible subset of the encoding and mask subset The complete set of components, the mask subset The parameters are set to random initialization and concatenated with the encoded visible subset. The decoder outputs the reconstructed EEG features The reconstructed predicted value of each masked EEG channel is:

[0066] The masked EEG channel is predicted by MSE (Mean Square Error) calculation The value of the corresponding original EEG feature The reconstruction loss between is:

[0067]

[0068] Finally, by minimizing the reconstruction loss rec Until it reaches the preset reconstruction standard, for example, when the reconstruction loss loss rec When it is less than the set reconstruction threshold, the training is stopped and a pre-trained universal feature extractor E is obtained.

[0069] In step S12, considering that the trained general feature extractor may not be able to specifically extract the personalized emotions of different users, personalized training and self-supervised tuning are performed based on the pre-trained general feature extractor. Generally speaking, in daily scenarios, using a general feature extractor can meet the general needs of users. However, in order to further accurately identify the emotions of each user and further improve the accuracy of emotion recognition, detailed personalized training is performed.

[0070] As an embodiment, the personalized training of the universal feature extractor based on the calibrated EEG signal of the target subject and the baseline emotion label corresponding to the calibrated EEG signal includes: adding a linear layer for emotion classification to the universal feature extractor to obtain an initialized personalized emotion predictor; using the emotion predictor to determine the predicted emotion label of the calibrated EEG signal; training the emotion predictor based on the cross-entropy loss of the predicted emotion label and the baseline emotion label until the cross-entropy loss reaches a preset loss standard.

[0071] In this embodiment, for a specific subject s (that is, the user we want to train, for example, a patient with an affective disorder in the medical field, or a user using an intelligent voice assistant in the field of artificial intelligence), the calibration data of the subject is obtained at this time. and the corresponding baseline sentiment labels Obtain a personalized calibrated emotion predictor for subject s by fine-tuning the general feature extractor E

[0072] Add a linear layer for emotion classification to the pre-trained general feature extractor E, and use the parameters of the feature extractor E to initialize the personalized emotion predictor

[0073]

[0074] Then the calibration data Enter the sentiment predictor The calculated emotion categories of the predicted emotion labels and the baseline emotion labels Cross entropy loss loss cls , similarly, we can minimize the loss cls To fine-tune the sentiment predictor

[0075]

[0076] Then, an emotion predictor for self-supervised learning of the target subject is obtained.

[0077] For step S13, after the emotion predictor is trained, the EEG data of the target subject is collected. The EEG data input by this method can be a complete EEG signal, or an unlabeled EEG signal or a damaged EEG signal. Input to the sentiment predictor Predict the sentiment category:

[0078]

[0079] Finally, the emotional category of the target subject is obtained.

[0080] It can be seen from this implementation that reconstructing the masked EEG channel is used as a proxy task in the pre-training stage, which fully mines the information of unlabeled data and gives the model the ability to decode a small amount of labeled and damaged EEG data. The hybrid structure based on CNN-Transformer makes full use of the information of the spectrum, time and space domains of the EEG signal, and then solves the problem of decoding emotions from a small amount of labeled and damaged EEG data through generative self-supervised learning of reconstructing the masked EEG channel.

[0081] The method is described in detail in the experiment. The pre-training dataset is composed of the unlabeled training data of all subjects, represented as X = {X1, ..., X S}, where S represents the number of subjects. The concatenated EEG features are the extracted spectral features and can also be represented as a sequence Where N is the number of samples in the time series, C represents the number of EEG (electroencephalogram) channels, and F represents a set of frequency bands (δ: 1-4 Hz, θ: 4-8 Hz, α: 8-14 Hz, β: 14-31 Hz, γ: 31-50 Hz) transformed in the spectral domain by STFT (Short-time Fourier Transform). The pre-trained general feature extractor is denoted as E, and the subject-specific calibrated emotion predictor s is denoted as Where s represents the sth topic. and Represent the calibration data and labels respectively. The test data and labels of subject s are represented as and

[0082] This method designs a MV-SSSTMA (Multi-view Spectral-Spatial-Temporal Masked Autoencoder) based on multi-view CNN (Convolutional Neural Network)-Transformer, such as Figure 2 As shown in Figure 2. The entire model can be divided into three stages: pre-training stage, personalized calibration stage and individual testing stage. In the pre-training stage, the channels of the unlabeled EEG data X from all subjects are randomly masked and then reconstructed to learn the general information extracted by the feature extractor shared by all subjects. In the personalized calibration stage, only a few labeled data from a specific subject s are used. and Used to calibrate the personal emotion predictor from the pre-trained generalized feature extractor E During the testing phase, EEG data and damaged data Can be achieved through Decoding to identify emotional states. A general feature extractor E of generalized features is pre-trained, which learns knowledge of the unlabeled EEG data of all subjects, with the goal of better identifying the emotional state of a specific subject in the future. To address the problem of decoding emotions from less and damaged EEG data, generative learning of reconstructing masked EEG channels is chosen as a proxy task to learn a general representation of EEG data. Considering the characteristics of EEG signals, a pre-trained model based on a multi-view CNN-Transformer hybrid structure is designed, which consists of a spectral embedding layer, a spatial position encoding layer, L hybrid encoders, and L symmetric hybrid decoders. Each hybrid block includes a temporal multi-scale random convolution layer and a spatial multi-head self-attention layer.

[0083] Since DE (differential entropy) has been shown to have excellent performance in EEG-based emotion recognition tasks, the differential entropy features extracted from the spectral domain of the EEG signal are used as the input of the model. Convert to Sample Overlapping windows. For each sample i, In the spectral embedding layer, we first pass a linear layer to transform Projected into D-dimensional space to embed the spectral information of EEG signal. It is embedded in the shape of C×T×D and its expression is as follows:

[0084]

[0085] Among them, the weight vector and deviations

[0086] For the spatial position encoding layer, the EEG data is divided into multiple blocks C (spatial dimensions) according to the different dimensions of the EEG channels. One dimension represents one EEG channel. This is to remember the position of each EEG channel and reconstruct it in subsequent tasks.

[0087] For the masking step, randomly sample a visible subset and the mask subcollection Among them, C v UC m =C. Only Serves as input to the hybrid encoder.

[0088] In order to capture the temporal information of EEG signals, a multi-scale temporal causal convolutional layer is introduced to enable the model to learn dynamic temporal representations. Three random convolutional layer branches with long, medium and short kernel sizes are implemented, corresponding to Figure 2 The temporal convolution layers in

[15] compute the temporal brain summary for each EEG channel from the input spectral features.

[0089] The multi-scale temporal causal convolution layer uses temporal causal convolution with multiple convolution kernel lengths to capture different ranges of time steps. The short temporal kernel is designed to learn short-term representations, while the long temporal kernel is used to extract long-term representations. Through the multi-scale temporal kernel, the diverse representations of EEG data can be enriched and emotion-related information can be fully learned. Dynamic long-term and short-term temporal patterns are generated by applying multi-scale temporal kernels in parallel on the input EEG samples. The temporal convolution kernel size k for the temporal convolution layer - long, the temporal convolution layer - medium, and the temporal convolution layer - short t ×1 are set to k l ×1, k m ×1 and ks ×1.

[0090] Different from the temporal image of the video, the temporal sequence of the EEG signal is represented as a continuous sequence of each channel. For each channel c∈{1,...,C}, the embedding of c is updated by the adjacent frames of the same channel. In each EEG channel, the input The time dimension T is the kernel size K t ×1 convolution operation. Among them, K t To encode temporal information in the neighborhood.

[0091] Furthermore, causal convolution is used to enforce that information does not flow from the future to the past. Figure 3 As shown, the output at time t only depends on the input at time t and earlier. The channel temporal convolution implemented in the model does not change the shape of the vector, so a length of K is added t The zero padding of -1 is used to keep the shape unchanged. The temporal convolution of the three scale branches in the model can be expressed as:

[0092]

[0093]

[0094]

[0095] in, is the input spectral feature, which is set as and It is a batch normalization operation to maintain the stability of the model.

[0096] After the temporal convolutional layer, spatial multi-head self-attention is used to learn the dynamics and inter-channel dependencies of all visible EEG channels, as Figure 3 As shown. For the long-scale branches, Reshape to C v ×TD shape, then the EEG embedding can be expressed as The scaled dot product is used to explicitly capture the topological relationship between EEG channels, which is expressed as:

[0097]

[0098] where Q, K, and V represent the query vector, key vector, and value vector respectively, and TD is the dimension of the key vector used to scale the dot product.

[0099] The dot product similarity is evaluated between Q and K of the channel of interest. If Q and K are similar, meaning the attention weight is high, then the corresponding values are assumed to be correlated. Here, Q, K, and V vectors are the input brain embeddings Specifically, the spatial brain summary of the long-scale branches is expressed as Calculate the attention weights between EEG channels through multi-head attention:

[0100]

[0101]

[0102]

[0103]

[0104]

[0105] in and It is the weight matrix that connects the multi-head results and projects them back to the representation space. Spatial attention matrix Indicates the degree of attention one channel pays to another channel.

[0106] The other two branches are processed in the same way as the long-scale branch. The spatial brain embeddings in the three scale branches are fused through a summation operation:

[0107]

[0108] in and Represent the outputs of the short-scale branch, medium-scale branch, and long-scale branch in the spatial attention layer, respectively. The branches representing the holistic spatial brain summaries of all visible EEG channels at three scales. After the spatial attention, layer normalization and a feedforward network follow. A hybrid LCNN-Transformer encoder is stacked to update the embeddings and further extract EEG features.

[0109] The final embedding is represented as

[0110] After feature extraction, a symmetric decoder is used to reconstruct the masked EEG channel, which consists of L similar CNN-Transformer hybrid blocks and linear layers. The encoder-decoder adopts a symmetric structure to obtain a stronger decoder to reconstruct complex EEG data. The input to the decoder is the encoded visible channel and shielded channels A complete set of components. are set to randomly initialized parameters and concatenated with the encoded visible channels. The decoder outputs the reconstructed EEG features The reconstruction process predicts the value of each masked EEG channel. The loss is only in the reconstruction of the masked channel. The mean square error (MSE) is calculated between the corresponding original EEG features. Finally, the pre-trained universal feature extractor E is obtained by minimizing the construction loss.

[0111] For a personalized target subject s, the calibration data consists of a small number of labeled samples for each emotional state in the subject’s original training dataset, denoted as and Since EEG data are recorded in chronological order, it is reasonable to use the data at the beginning of the training dataset as calibration data. A personalized calibrated emotion predictor is obtained. By fine-tuning the generalized feature extractor E, we predict the emotion class through a linear layer. The classification loss is measured by cross entropy.

[0112] During the testing phase, the model accepts corrupted EEG data. A test set of subject s from the original test dataset is used, denoted as and To validate the personalized model To simulate corrupted data, channels are masked in the same way as in the pre-training stage.

[0113] The model of this method is evaluated on the emotional EEG datasets (SEED dataset and SEED-IV dataset), the stimulus materials of these datasets are all video clips. Among them, the SEED dataset contains EEG signals of 15 participants, which are divided into three emotional states: positive, neutral and negative. Each subject performed three 15 trials at different times. In each session, the first 9 trials are usually used as training data and the remaining 6 trials are used as test data. The SEED-IV dataset is collected for four emotional states: happiness, sadness, fear and neutral emotions. The 15 subjects participated in three trials on different days, with 24 trials each time. Generally speaking, the first 16 trials are training data and the remaining 8 trials are test data for each session.

[0114] To ensure comparability of our results, we employed the same common experimental setup as previously used for both datasets. Since the classes in the datasets are balanced, performance is evaluated using the average accuracy and standard deviation across sessions. For each experiment, our pre-training data X consists of the concatenation of the unlabeled raw training data for all subjects, including 9 trials for the SEED dataset and 16 trials for the SEED-IV dataset. Starting from the training dataset for the target subject, a small amount of labeled data, 10, 20, or 30 per emotional state, is used for calibration.

[0115] Pre-training data Transformed by overlapping windows of size T To maintain the same sample size of 10 samples as in the comparative experiments, C represents the number of EEG channels, equal to 62. The experiments used the PyTorch deep learning framework. For each experiment, the learning rate of the proposed model ranged from 0.001 to 0.00001. Furthermore, the spectral embedding size D was set to 16, the number of hybrid blocks L was equal to 6, and the multi-head dimension H was set to 6.

[0116] The baseline models involved include:

[0117] STRNN: Spatiotemporal recurrent neural network is based on a unified spatiotemporal dependency model and learns information from both spatiotemporal and temporal aspects.

[0118] DGCNN: Dynamic Graph Convolutional Neural Network dynamically learns the representation of EEG signals through graph convolution for EEG-based emotion recognition.

[0119] BiDANN: A dual-domain adversarial neural network focuses on the discriminative features of EEG signals from the left and right hemispheres of the brain for EEG-based emotion recognition.

[0120] BiHDM: The bihemispheric difference model studies the asymmetric differences between the left and right hemispheres of the brain.

[0121] R2G-STNN: A regional-to-global spatiotemporal neural network model learns global and regional EEG representations of EEG signals in space and time.

[0122] RGNN: Regularized Graph Neural Networks for Exploring the Topology of EEG Channels via Graph Convolution.

[0123] MD-AGCN: Multi-domain adaptive graph convolutional network, fully utilizing features from different domains.

[0124] MAE: Masked Autoencoder as a scalable self-supervised learner by reconstructing missing patches in computer vision images.

[0125] Results for different amounts of calibration data. Figure 4 and Figure 5 In Figure 2, we present comparisons between our proposed model and a baseline model using both all labeled training data and a limited amount of labeled training data (10, 20, and 30 per emotion state, respectively) from the SEED and SEED-IV datasets. For each emotion, the 10, 20, and 30 labeled data points were collected from the beginning of the same epoch of trials. It is important to note that our results are only compared to models that followed the same general experimental setup.

[0126] like Figure 4As shown, compared with supervised methods, our model achieves state-of-the-art results on the SEED and SEED-IV datasets, demonstrating that the pre-training process can improve the model's generalization and efficiency, especially on problems with a high number of emotion classes. Specifically, our model achieves a recognition accuracy of 95.32% on the SEED dataset with a standard deviation of 3.05%. On the SEED-IV dataset, our model achieves significant improvements, with the highest accuracy of 92.82% and the lowest standard deviation of 5.03%. Furthermore, the MAE method also outperforms the baseline method on SEED-IV and outperforms some supervised models on SEED. This may be due to the fact that these supervised models take into account the temporal information of the EEG.

[0127] When only a small amount of labeled data is available for calibration, the self-supervised method MAE and the supervised method MD-AGCN are used to evaluate the MV-SSMA of our method. Figure 5 As shown, the # labeled data column indicates the amount of labeled training data for each emotional state for our model. All models improve in accuracy when using more labeled data. The increase can be small because different amounts of labeled data are adjacent and come from the same epoch, meaning there's a lack of diversity. Furthermore, our model outperforms both MAE and MD-ADCN in every case.

[0128] We also tested the performance of MVSTMA and MAE for all subjects in all the above cases, as well as the performance of MV-SSSTMA and MD-AGCN. In all cases, the significance level was well below 1%, indicating that there were significant differences between them.

[0129] (1) The pre-training phase captures a generalized representation of the EEG signal, while the calibration process transfers the training of the model to the specified target subjects.

[0130] (2) This method model fully utilizes EEG signals in the spectral, temporal and spatial domains.

[0131] like Figure 6 As shown in Figure 1, 10 labeled calibration data on the SEED-IV dataset demonstrate the results of different proportions of channel damage rates in the test data. Each column represents the percentage of damaged channels in the test data. The reason for using the SEED-iv dataset is that the four emotion categories in SEED-iv include all three emotional states in SEED. Figure 6 As can be seen, when 30% of the channels in the test data are damaged, MV-SSTMA can well identify emotional states, achieving 73.68% and a standard deviation of 7.58% with only 10 labeled calibration data. In addition, even when more EEG signal channels are damaged, the proposed model can still distinguish emotional states well.

[0132] To demonstrate the effect of the channel-wise incidental convolutional layers in the hybrid encoder block, ablation studies were performed by replacing them with temporal embeddings (i.e., NoHybrid). In the NoHybrid model, temporal information is still considered by adding temporal embeddings to the original spectral embedding layer, but it cannot be viewed interchangeably with the spatial information in the L encoder block. Ablation studies were also performed by reducing the multi-scale temporal branches of MV-SSTMA to use only a single scale branch in the model, called SingleScale. The contribution of causal convolutions in the single-scale model was evaluated by replacing them with ordinary convolution operations.

[0133] like Figure 7 Figure 3 shows the performance of our MV-SSTMA, the NoHybrid model, and the SingleScale model on the SEED and SEED-IV datasets, with varying amounts of labeled data for each emotional state. Our model consistently outperforms the NoHybrid and SingleScale models, a fact that highlights the importance of the channel-wise random convolutional layers and the multi-scale branches with random convolutions in the hybrid encoder block. Furthermore, since the NoHybrid and SingleScale models also consider temporal information, they continue to outperform MAE.

[0134] like Figure 8 The confusion matrix of our MV-SSTMA, calibrated on the SEED and SEED-IV datasets using both 10 and all labeled training data, illustrates its ability to distinguish each emotional state. For the SEED dataset, our model was best able to identify the positive emotional state and had the most difficulty identifying the neutral emotional state across both the 10 and all labeled training data. For the SEED-IV dataset, the 10 labeled calibration data were the most difficult emotional state to identify, while the neutral state was the easiest to identify. Furthermore, when calibrated using all labeled training data, our model still decoded the neutral state better than all three other emotional states, with the fear state being the most difficult to distinguish.

[0135] We further studied the ability of this method model to reconstruct damaged EEG channels from test data. Figure 9 Figure 1 shows reconstructed test data that was manually corrupted by randomly masking the EEG channels using different masking rates. As can be seen, EEG features can be reconstructed well at masking rates of 30% and 50%. At a masking rate of 70%, features can generally be reconstructed, but some details may be lost. However, at a masking rate of 90%, EEG features are more difficult to recover.

[0136] In summary, this method's self-supervised learning of a multi-view spectral-spatial-temporal masked autoencoding model addresses the problem of decoding emotion from scantly labeled and corrupted EEG data. This model exploits the spectral, spatial, and temporal characteristics of EEG data through a multi-view CNN-Transformer hybrid architecture, fully leveraging the EEG signal. The three-stage pre-training, calibration, and testing ensure the generalization, customization, and efficiency of the overall framework.

[0137] Extensive experiments on the SEED and SEED-IV datasets demonstrate the superior performance of our model compared to various advanced baseline models. Results on both minimally labeled and impaired EEG data demonstrate that the MV-SSSMA model can learn EEG representations from a large amount of unlabeled data and effectively decode emotional states from minimally labeled and even impaired EEG data. Visualization of reconstructed impaired EEG channels on test data demonstrates the effectiveness and ability of our model to restore missing channels in emotional EEG data, boosting the performance of EEG-based emotion recognition in a self-supervised manner.

[0138] like Figure 10 The figure shows a structural diagram of an emotion recognition system based on generative self-supervised learning and EEG signals provided by one embodiment of the present invention. The system can execute the emotion recognition method based on generative self-supervised learning and EEG signals described in any of the above embodiments and be configured in a terminal.

[0139] This embodiment provides an emotion recognition system 10 based on generative self-supervised learning and EEG signals, including: a general feature program module 11 , a personalized training program module 12 and an emotion recognition program module 13 .

[0140] Among them, the general feature program module 11 is used to input the differential entropy feature used to reflect the subject's EEG signal into the multi-view masked autoencoder model, reconstruct the differential entropy feature in the frequency domain and / or space and / or time dimension, and obtain the multi-view reconstructed differential entropy feature for simulating unlabeled and / or damaged EEG signals, pre-train the encoder and decoder of the multi-view masked autoencoder model based on the multi-view reconstructed differential entropy feature, and use the obtained encoder as the general feature extractor of the EEG signal; the personalized training program module 12 is used to perform personalized training on the general feature extractor based on the calibrated EEG signal of the target subject and the baseline emotion label corresponding to the calibrated EEG signal, and obtain an emotion predictor for self-supervised learning of the target subject; the emotion recognition program module 13 is used to perform personalized emotion prediction on the collected EEG data of the target subject based on the emotion predictor, wherein the EEG data includes: unlabeled EEG signals and damaged EEG signals.

[0141] An embodiment of the present invention further provides a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions can execute the emotion recognition method based on generative self-supervised learning and EEG signals in any of the above method embodiments;

[0142] As an embodiment, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are configured as follows:

[0143] Inputting differential entropy features used to reflect the subject's EEG signal into a multi-view masked autoencoder model, reconstructing the differential entropy features in the frequency domain and / or spatial and / or temporal dimensions to obtain multi-view reconstructed differential entropy features for simulating unlabeled and / or damaged EEG signals, pre-training an encoder and decoder of the multi-view masked autoencoder model based on the multi-view reconstructed differential entropy features, and using the obtained encoder as a universal feature extractor for the EEG signal;

[0144] Performing personalized training on the universal feature extractor based on the target subject's calibrated EEG signal and the baseline emotion label corresponding to the calibrated EEG signal to obtain an emotion predictor for the target subject through self-supervised learning;

[0145] Based on the emotion predictor, personalized emotion prediction is performed on the collected EEG data of the target subject, wherein the EEG data includes: unlabeled EEG signals and damaged EEG signals.

[0146] A non-volatile computer-readable storage medium can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods described in the embodiments of the present invention. One or more program instructions stored in the non-volatile computer-readable storage medium, when executed by a processor, perform the emotion recognition method based on generative self-supervised learning and EEG signals described in any of the aforementioned method embodiments.

[0147] Figure 11 This is a hardware structure diagram of an electronic device for an emotion recognition method based on generative self-supervised learning and EEG signals provided in another embodiment of the present application, such as Figure 11 As shown, the device includes:

[0148] One or more processors 1110 and memory 1120, Figure 11 A processor 1110 is used as an example. The device for emotion recognition method based on generative self-supervised learning and EEG signals may further include: an input device 1130 and an output device 1140 .

[0149] The processor 1110, the memory 1120, the input device 1130 and the output device 1140 may be connected via a bus or other means. Figure 11 The bus connection is taken as an example.

[0150] Memory 1120, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the emotion recognition method based on generative self-supervised learning and EEG signals in the embodiments of the present application. Processor 1110 executes the non-volatile software programs, instructions, and modules stored in memory 1120 to execute various server functional applications and data processing, thereby implementing the emotion recognition method based on generative self-supervised learning and EEG signals in the above-mentioned method embodiment.

[0151] The memory 1120 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data, etc. In addition, the memory 1120 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 1120 may optionally include a memory remotely located relative to the processor 1110, and these remote memories may be connected to the mobile device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0152] The input device 1130 can receive input digital or character information. The output device 1140 can include a display device such as a display screen.

[0153] The one or more modules are stored in the memory 1120 and, when executed by the one or more processors 1110 , perform the emotion recognition method based on generative self-supervised learning and EEG signals in any of the above method embodiments.

[0154] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.

[0155] The non-volatile computer-readable storage medium may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0156] An embodiment of the present invention also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the emotion recognition method based on generative self-supervised learning and electroencephalogram signals of any embodiment of the present invention.

[0157] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to:

[0158] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and are primarily designed to provide voice and data communications. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.

[0159] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, have computing and processing capabilities, and generally also have mobile Internet access. These terminals include PDAs, MIDs, and UMPC devices, such as tablet computers.

[0160] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.

[0161] (4) Other electronic devices with data processing functions.

[0162] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include" and "comprise" include not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or device. In the absence of further limitations, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device that includes the elements.

[0163] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0164] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An emotion recognition method based on generative self-supervised learning and EEG signals, comprising: The differential entropy features used to reflect the subject's EEG signal are input into a multi-view masked autoencoder model, and the differential entropy features are reconstructed in the frequency domain and / or spatial and / or temporal dimensions to obtain multi-view reconstructed differential entropy features for simulating unlabeled and / or damaged EEG signals. The encoding and decoding modules of the multi-view masked autoencoder model are pre-trained based on the multi-view reconstructed differential entropy features, and the obtained encoding module is used as a universal feature extractor for the EEG signal, wherein the multi-view masked autoencoder model consists of a spectrum embedding layer, a spatial position encoding layer, an EEG channel mask layer, a hybrid encoding block, and a hybrid decoding block symmetrical to the hybrid encoding block. The spectrum embedding layer is used to extract the spectrum information of the differential entropy feature. The spatial position coding layer is used to encode the spatial position of the EEG channel for destruction and reconstruction, wherein the spatial position coding method of the EEG channel includes sine-cosine position coding; The EEG channel mask layer is used to divide the reconstructed differential entropy features into visible subsets and mask subsets according to channels. The hybrid coding block is used to capture the dependencies between EEG channels in the visible subset and determine the multi-view fusion features of EEG. The hybrid decoding block is used to reconstruct EEG features through the visible subset and the mask subset outputs, pre-train the hybrid coding block and the hybrid decoding block based on the reconstruction loss of the original EEG features and the reconstructed EEG features output by the decoding module to obtain a universal feature extractor for the EEG signal, and add a linear layer for emotion classification to the universal feature extractor to obtain an initialized personality emotion predictor; Performing personalized training on the emotion predictor based on the target subject's calibrated EEG signal and a reference emotion label corresponding to the calibrated EEG signal to obtain a self-supervised learning emotion predictor for the target subject; Based on the emotion predictor, personalized emotion prediction is performed on the collected EEG data of the target subject, wherein the EEG data includes: unlabeled EEG signals and damaged EEG signals.

2. The method according to claim 1, wherein The hybrid coding block includes: a temporal causal convolution layer and a multi-head spatial self-attention layer, wherein: The temporal causal convolution layer includes a multi-scale temporal convolution kernel for extracting temporal information of EEG channels in the visible subset; The multi-head spatial self-attention layer is used to capture the spatial dependencies between EEG channels of the visible subsets after division.

3. The method according to claim 1, wherein The pre-training of the hybrid coding block and the hybrid decoding block based on the reconstruction loss of the original EEG features and the reconstructed EEG features output by the decoding module comprises: The determined mean square error between the reconstructed EEG features output by the decoding module and the original EEG features is used as the reconstruction loss, and the hybrid coding block and the hybrid decoding block are pre-trained based on the reconstruction loss until the reconstruction loss reaches a preset reconstruction standard.

4. The method according to claim 1, wherein The personalized training of the emotion predictor based on the target subject's calibrated EEG signal and the reference emotion label corresponding to the calibrated EEG signal includes: determining a predicted emotion label of the calibrated EEG signal using the emotion predictor; The emotion predictor is trained based on the cross entropy loss between the predicted emotion label and the baseline emotion label until the cross entropy loss reaches a preset loss standard.

5. The method according to claim 1, wherein The differential entropy feature for reflecting the subject's EEG signal is determined by the frequency spectrum of the subject's EEG signal in the frequency domain.

6. The method according to claim 5, wherein: Before determining the frequency spectrum of the subject's EEG signal in the frequency domain, the method further includes pre-processing the subject's EEG signal by filtering and noise reduction.

7. An emotion recognition system based on generative self-supervised learning and EEG signals, comprising: A general feature program module is used to input the differential entropy feature used to reflect the subject's EEG signal into a multi-view masked autoencoder model, reconstruct the differential entropy feature in the frequency domain and / or spatial and / or temporal dimensions to obtain a multi-view reconstructed differential entropy feature for simulating unlabeled and / or damaged EEG signals, pre-train the encoding and decoding modules of the multi-view masked autoencoder model based on the multi-view reconstructed differential entropy feature, and use the obtained encoding module as a general feature extractor for the EEG signal, wherein the multi-view masked autoencoder model consists of a spectral embedding layer, a spatial position encoding layer, an EEG channel mask layer, a hybrid encoding block, and a hybrid decoding block symmetrical to the hybrid encoding block. The spectrum embedding layer is used to extract the spectrum information of the differential entropy feature. The spatial position coding layer is used to encode the spatial position of the EEG channel for destruction and reconstruction, wherein the spatial position coding method of the EEG channel includes sine-cosine position coding; The EEG channel mask layer is used to divide the reconstructed differential entropy features into visible subsets and mask subsets according to channels. The hybrid coding block is used to capture the dependencies between EEG channels in the visible subset and determine the multi-view fusion features of EEG. The hybrid decoding block is used to reconstruct EEG features through the visible subset and the mask subset outputs, pre-train the hybrid coding block and the hybrid decoding block based on the reconstruction loss of the original EEG features and the reconstructed EEG features output by the decoding module to obtain a universal feature extractor for the EEG signal, and add a linear layer for emotion classification to the universal feature extractor to obtain an initialized personality emotion predictor; a personalized training program module for performing personalized training on the emotion predictor based on the target subject's calibrated EEG signal and a reference emotion label corresponding to the calibrated EEG signal, to obtain an emotion predictor for the target subject through self-supervised learning; The emotion recognition program module is used to perform personalized emotion prediction on the collected EEG data of the target subject based on the emotion predictor, wherein the EEG data includes: unlabeled EEG signals and damaged EEG signals.

8. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1 to 6.

9. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Intelligent emotional state recognition and adjustment method based on electroencephalogram signals

    CN114384998A

  • Brain tumor self-supervision pre-training method and device based on attention symmetry self-coding

    CN115035093A