EEG emotion recognition system and method based on multi-scale space-time diagram convolution and comparative learning

By using multi-scale spatiotemporal graph convolution and contrast learning technology in the EEG emotion recognition system, the problem of existing algorithms ignoring spatial and frequency information and individual differences is solved, and higher recognition accuracy and generalization performance are achieved.

CN119970033AActive Publication Date: 2025-05-13SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202411961308.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-13
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The existing EEG emotion recognition algorithm ignores the spatial and frequency information in the EEG signal, lacks the ability to model complex relationships between different brain regions, and individual differences lead to insufficient recognition accuracy and generalization performance.

Method used

The EEG emotion recognition system based on multi-scale spatiotemporal graph convolution and contrast learning is adopted to extract hidden spatial representations of EEG signals through a stack autoencoder, construct an undirected graph structure of brain source signals, and use a multi-scale spatiotemporal graph convolution network to extract emotions-related features, and combine it with contrast learning to eliminate individual differences.

Benefits of technology

Effectively capture the spatial characteristics in the EEG signal, improve the accuracy and robustness of emotion recognition, enhance the generalization performance of the model, and significantly improve the effect of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119970033A_ABST
    Figure CN119970033A_ABST
Patent Text Reader

Abstract

The invention relates to the field of deep learning and emotion recognition, in particular to an EEG emotion recognition system and method based on multi-scale space-time diagram convolution and comparative learning, and the method comprises the steps: carrying out the preprocessing of an EEG signal; constructing positive and negative sample pairs based on the emotion categories and the identities of the subjects; extracting hidden space representation of the electroencephalogram signal by adopting a stack auto-encoder, and dividing the electroencephalogram signal into a plurality of functional brain regions to obtain a brain source signal after dimension reduction; constructing an undirected graph structure of the brain source signal; constructing a multi-scale space-time diagram convolutional network as a feature extraction network, inputting an undirected graph structure, dynamically learning and updating an adjacent matrix, expressing features of brain source signals in the undirected graph structure, and extracting emotion related features; based on a multi-scale space-time diagram convolution and comparative learning framework, an EEG emotion recognition model is built and trained to eliminate individual differences. According to the method, the multi-dimensional information of the EEG signals is fully utilized, and the accuracy, robustness and generalization performance of emotion recognition are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning and emotion recognition, and in particular to an EEG emotion recognition system and method based on multi-scale spatiotemporal graph convolution and contrastive learning. Background Art

[0002] Emotions not only profoundly affect people's cognitive and behavioral activities, but are also an important factor affecting mental health. In recent years, with the development of deep learning technology, emotion recognition algorithms based on physiological signals have made significant progress, especially electroencephalogram (EEG) signals, which have received widespread attention due to their close relationship with emotions.

[0003] Currently, many EEG emotion recognition algorithms mainly use one-dimensional convolutional neural networks (CNNs) to process data of a single dimension. These methods are effective to a certain extent, but they usually ignore the spatial and frequency information in the EEG signal. In addition, these algorithms lack the ability to model the complex relationships between different brain regions, which limits the accuracy and comprehensiveness of emotion recognition. In addition, the problem of individual differences that are prevalent in physiological signals will lead to poor results and low accuracy of general discrimination methods when applied to new subjects. How to eliminate individual differences has also become a key issue.

[0004] In view of the above challenges, it is of great significance to develop an EEG emotion recognition system and method based on multi-scale spatiotemporal graph convolution and contrastive learning. Summary of the invention

[0005] In order to overcome the shortcomings and deficiencies of the prior art, the present invention provides an EEG emotion recognition system and method based on multi-scale spatiotemporal graph convolution and contrastive learning.

[0006] The emotion recognition method of the present invention adopts the following technical solution: an EEG emotion recognition method based on multi-scale spatiotemporal graph convolution and contrastive learning, comprising the following steps:

[0007] S1. Collect the EEG signals of the subjects and pre-process them to obtain sample data that truly reflects the emotional processing process of the subjects; construct positive and negative sample pairs based on the emotion category and the subject identity for comparative learning; construct a cross-subject dataset with one subject left out;

[0008] S2, using stacked autoencoders to extract the latent space representation of EEG signals, dividing EEG signals into multiple functional brain regions based on medical prior information, and obtaining brain source signals after dimensionality reduction;

[0009] S3, constructing an undirected graph structure of brain source signals;

[0010] S4. Construct a multi-scale spatiotemporal graph convolutional network as a feature extraction network, input the undirected graph structure into the feature extraction network, dynamically learn and update the adjacency matrix, express the characteristics of brain source signals in the undirected graph structure, and extract emotion-related features;

[0011] S5. Based on the multi-scale spatiotemporal graph convolution and contrastive learning framework, an EEG emotion recognition model is built and trained to eliminate individual differences.

[0012] The emotion recognition system of the present invention adopts the following technical solution: an EEG emotion recognition system based on multi-scale spatiotemporal graph convolution and contrastive learning, comprising the following modules:

[0013] The data set construction module is used to collect the EEG signals of the subjects and pre-process the EEG signals to obtain sample data that truly reflects the emotional processing process of the subjects; and to construct a cross-subject data set with one subject left out;

[0014] The brain source signal acquisition module uses a stacked autoencoder to extract the latent space representation of the EEG signal, divides the EEG signal into multiple functional brain areas based on medical prior information, and obtains the brain source signal after dimensionality reduction;

[0015] Undirected graph construction module, used to construct the undirected graph structure of brain source signals;

[0016] Convolutional network construction module, used to construct a multi-scale spatiotemporal graph convolutional network as a feature extraction network, input the undirected graph structure into the feature extraction network, dynamically learn and update the adjacency matrix, express the characteristics of brain source signals in the undirected graph structure, and extract emotion-related features;

[0017] The recognition model building and training module is used to build and train the EEG emotion recognition model based on the multi-scale spatiotemporal graph convolution and contrastive learning framework to eliminate individual differences.

[0018] Compared with the prior art, the technical effects achieved by the present invention include:

[0019] The graph convolutional network can effectively capture the spatial features in the EEG signal, while the autoencoder can perform feature compression and dimensionality reduction processing, and add contrastive learning to eliminate the differences between different individuals and improve the generalization performance of the model; it makes full use of the multi-dimensional information of the EEG signal and significantly improves the accuracy, robustness and generalization performance of emotion recognition, which has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a flowchart of multi-scale spatiotemporal graph convolution in an embodiment of the present invention;

[0021] Figure 2 Schematic diagram of the algorithm model in the embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the present invention more obvious, the exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described here; the exemplary embodiments and their descriptions are only used to explain the present invention, and are not intended to limit the present invention.

[0023] Furthermore, in the present specification and the drawings, steps and elements having substantially the same or similar features are denoted by the same or similar reference numerals, and repeated description of these steps and elements will be omitted.

[0024] Example

[0025] This embodiment proposes an EEG emotion recognition method based on multi-scale spatiotemporal graph convolution and contrastive learning. The implementation process is based on a deep learning framework of stacked autoencoders and spatiotemporal graph convolution, which is used to extract high-quality brain source signals and hidden features from EEG data. First, the original electrode channel is reduced in dimension by a stacked autoencoder to extract key brain source information from redundant signals. Then, in the feature extraction stage, modeling is performed based on the spatial structure and temporal structure of EEG data: in the spatial dimension, an undirected graph is constructed and a graph convolutional network (GCN) based on an attention mechanism is introduced to capture the spatial correlation between electrodes; in the temporal dimension, a multi-scale convolutional network (MSCN) is used to extract dynamic features at different time resolutions. In the overall framework, contrastive learning is introduced to reduce the impact of individual differences and enhance the generalization ability of emotion classification. In the contrastive learning strategy, pairs of the same emotion data from different subjects are selected as positive pairs, and pairs of different emotion data from the same subject are selected as negative pairs. This contrastive target is used to eliminate individual differences, thereby modeling emotion-related features more accurately.

[0026] Specifically, if Figure 1 , Figure 2 As shown, the EEG emotion recognition method based on multi-scale spatiotemporal graph convolution and contrastive learning in this embodiment includes the following steps:

[0027] S1. Collect the EEG signals of the subjects and pre-process them to obtain sample data that truly reflects the subjects' emotion processing process; construct positive and negative sample pairs based on emotion categories and subject identities for comparative learning; and construct a cross-subject dataset with one subject left out.

[0028] This embodiment is based on the public EEG emotion dataset "SJTU Emotion EEG Dataset". This dataset induces emotional reactions through video media and collects EEG signals while the subjects watch different emotion-triggering videos. Specifically, video media are divided into three categories according to the induced emotion categories: positive emotions, neutral emotions, and negative emotions, and the collected EEG signals are labeled with the corresponding emotion categories. This dataset contains EEG signals of 15 subjects in total.

[0029] In this embodiment, the data is divided into 1 second time windows without overlap. Since the sampling rate of the data is 200 Hz, the time series length of each data segment is 200, and each data segment contains 62 channels, corresponding to 62 EEG acquisition electrodes. The experiment is based on the Leave-One-Subject-Out (LOSO) method across subjects, that is, each time the sample of one subject is used as the test set, and the samples of all other subjects are used as the training set, and the data of each subject is used as the test set in turn to construct a cross-subject data set.

[0030] Preprocessing includes filtering, denoising, etc. For subsequent comparative learning, in the preprocessing process, the collected EEG signals are divided into positive pairs and negative pairs. The positive pairs here are the same emotions of different subjects, and the negative pairs are different emotions of the same subject. Specifically: for the SEED data set, there are 15 subjects. For each emotion (positive, neutral, negative), data pairs of different subjects are selected, such as: the positive emotions of subject 1 and the positive emotions of subject 2 constitute a positive pair; the positive emotions of subject 1 and the positive emotions of subject 3 constitute a positive pair, etc. Similarly, the positive emotions of subject 1 and the neutral emotions of subject 1 constitute a negative pair; the neutral emotions of subject 2 and the negative emotions of subject 2 constitute a negative pair. Since the number of such permutations and combinations is relatively large, in order to reduce the computational complexity and ensure the diversity of training data and the generalization ability of the model, this embodiment adopts a random sampling strategy to randomly extract some samples from all positive pairs and negative pairs for training. Through this sampling method, the amount of calculation can be effectively reduced and the training efficiency can be improved on the basis of retaining sufficient data diversity.

[0031] S2. Use stacked autoencoders to extract the latent space representation of EEG signals, divide the EEG signals into multiple functional brain areas based on medical prior information, and obtain brain source signals after dimensionality reduction.

[0032] In this step, in order to extract the latent spatial feature representation of the EEG signal, a stacked autoencoder (SAE) is used, and the EEG signal is divided into 12 functional brain regions based on medical prior information to generate brain source signal features after dimensionality reduction. Specifically, the division of these 12 brain regions is based on physiological and medical prior information, by dividing the brain into left and right hemispheres (hemispheres), and combining the horizontal plane (anterior, lateral, posterior) and vertical plane (inferior, superior), and is formed based on the electrode recording position of the international 10-20 system. In addition, this embodiment also studies the influence of different numbers of source signal channels (such as 6 channels and 7 channels), but the experimental results show that the brain region division using 12 channels is the best, which is also one of the important reasons why this embodiment selects 12 brain regions as the basis for functional division.

[0033] The stacked autoencoder SAE is used to reduce the dimensionality of multi-channel brain source signals to reconstruct brain source signals and extract key features. SAE maps high-dimensional input data to a low-dimensional hidden space through unsupervised learning while retaining the core information of brain source signals. Specifically, the encoder part of the stacked autoencoder compresses high-dimensional brain source signals into low-dimensional hidden features, and the decoder part reconstructs brain source signals. By minimizing the error between the input signal and the reconstructed brain source signal, it is ensured that the low-dimensional hidden features can truly and comprehensively represent the brain source signal. These features provide a high-quality foundation for subsequent feature extraction and analysis.

[0034] Before the stacked autoencoder SAE performs dimensionality reduction on multi-channel brain source signals, the SAE is trained first. During the training process of SAE, a layer-by-layer unsupervised pre-training strategy is adopted to learn deep feature representations. Taking a network with an n→m→k structure as an example, the training process is divided into two steps: first, train an n→m→n autoencoder to learn the feature mapping from n to m; then, train an m→k→m autoencoder to further extract feature representations from m to k. Finally, these features are gradually stacked to construct a complete deep network structure n→m→k. This layer-by-layer optimization process is similar to building a building, gradually building a complete model through a solid foundation.

[0035] The training process of stacked autoencoder SAE is as follows:

[0036] First layer autoencoder training: First, the original brain source signal x(k) is passed into the first sparse autoencoder to train the network to learn the first-order feature representation h of the input data. (1) (k). This step is equivalent to learning the mapping from the original brain source signal to the low-dimensional feature space.

[0037] Second layer autoencoder training: Next, the first-order feature representation h (1) (k) is used as a new input and passed into the second sparse autoencoder to train it to learn the second-order feature representation h (2) (k). This step further extracts higher-level features of the data and gradually reduces the data dimension.

[0038] Classifier training: After obtaining the second-order feature representation h (2) (k), these features can be used as input to train a Softmax classifier to map high-level features to specific emotion classification labels. In this way, the network can eventually perform emotion recognition based on the features after dimensionality reduction. The whole process ensures that each layer can effectively extract different levels of features in the brain source signal through layer-by-layer training.

[0039] To ensure the performance of the model, each sparse autoencoder is also optimized by reconstruction loss. Specifically, the reconstruction loss measures the difference between the output and input of the sparse autoencoder and is defined as:

[0040]

[0041] in, It is the brain source signal reconstructed by the autoencoder. By minimizing the reconstruction loss, it is ensured that the feature representation after dimensionality reduction still retains the key information of the original brain source signal. Ultimately, this layer-by-layer training forms a stacked autoencoder network with two hidden layers (reducing the 62 channels to 12 channels respectively).

[0042] S3. Construct an undirected graph structure of brain source signals.

[0043] Brain-source signals are naturally suitable for corresponding to undirected graph structures. Each brain-source signal channel can be regarded as a node in the undirected graph, and the relationship between channel signals is mapped to the edge features of the undirected graph. Therefore, the brain-source signals obtained from the stacked autoencoder are converted into data representations of the undirected graph structure. Specifically, the edge features of the undirected graph are defined by constructing an adjacency matrix, and the weight of the edge is measured by the phase-locked value of the signal between the nodes, that is, by calculating the instantaneous phase relationship between the two signals, the dependency and correlation between the channels are characterized.

[0044] S31. For the 12-dimensional EEG channel brain source signal, convert it into a data representation of an undirected graph structure.

[0045] Specifically, the brain source signal of each channel is first normalized to reduce its range to [-1,1]. The normalization formula is:

[0046]

[0047] Where x is the original brain source signal, μ is the mean of the brain source signal, and σ is the standard deviation of the brain source signal. The standardized signal value will be used as the feature of each node in the undirected graph. Each channel corresponds to a node in the undirected graph, and there are 12 nodes in total. Specifically, these standardized brain source signals are used as input node features in the graph convolution operation to construct the initial representation of the undirected graph. In the subsequent graph convolution process, these node features are optimized through the propagation and update mechanism in the graph convolution network. The graph convolution network will update the feature value of each node based on the adjacent relationship between nodes (that is, the edges of the undirected graph), thereby gradually extracting more complex and high-level features in space.

[0048] S32. In order to define the connection relationship between channels, the phase locking value (PLV) is used to measure the phase synchronization between the brain source signals of two EEG channels.

[0049] In this embodiment, the calculation formula of the phase lock value PLV is:

[0050]

[0051] where φ i (t) is the instantaneous phase of the ith channel at time t, φ j (t) is the instantaneous phase of the jth channel at time t, and T is the length of the time series. The closer the phase-locked value PLV is to 1, the stronger the synchronization of the two channels in phase and the closer the connection relationship. When constructing an undirected graph structure, the PLV value will be used as the weight of the edge to represent the mutual correlation between channels.

[0052] In the undirected graph modeling of EEG multi-channel signals, PLV is used as the weight of the edge in the undirected graph. This method can effectively capture the phase synchronization between brain source signal channels and reveal the pattern of collaborative work of different brain regions. EEG is a multi-channel time series signal, each channel records the electrical activity from different brain regions. Due to the complex functional interactions between different brain regions, studying the phase synchronization between channels helps to describe the functional connections of brain regions. Specifically, the value of PLV is between 0 and 1, 1 means that the phases of the two signals are completely synchronized, and 0 means that the phases of the signals are completely unrelated. Compared with directly using the signal amplitude, PLV focuses on the relative phase changes between signals, which enables it to more stably describe the relationship between brain regions under different subjects and experimental conditions. In addition, PLV is more robust to noise, is not affected by amplitude changes, and is suitable for processing frequency band information in EEG signals. Therefore, using PLV as an edge feature, the phase synchronization between each pair of EEG channels can be represented as a weight, and an undirected graph structure reflecting the functional connection of the brain can be constructed.

[0053] S33. Introduce a preset threshold τ, perform sparse processing on the matrix A of the phase-locked values, and construct a sparse adjacency matrix of the phase-locked values.

[0054] In order to construct a sparse adjacency matrix and avoid excessive noise or weakly correlated connections in the undirected graph structure interfering with the modeling effect, this embodiment introduces a thresholding strategy. Specifically, after the PLV value is calculated, the PLV value less than the set threshold is directly set to 0, and only the strongly correlated connections are retained. This sparse process not only reduces the computational complexity, but also highlights the more significant functional connections between brain regions. The specific PLV matrix is ​​A∈R N×N , where N represents the number of EEG channels, and the element A in the matrix ij Represents the PLV value between the i-th and j-th channels. In order to construct a sparse adjacency matrix, a preset threshold τ is introduced to perform sparse processing on the PLV matrix A, and the following formula is defined:

[0055]

[0056] in is the adjacency matrix element after sparseness. Through this thresholding operation, weakly correlated edges with PLV less than τ are removed, and only strongly correlated connections are retained.

[0057] Correspondingly, the sparse adjacency matrix can be expressed as:

[0058] A sas =A⊙M,

[0059] Where M is a mask matrix, defined as:

[0060]

[0061] By constructing the undirected graph structure of EEG through the adjacency sparse matrix of PLV, it is possible to capture the static functional connectivity characteristics between brain regions and explore the dynamic connectivity changes under different emotional states. This undirected graph representation provides rich and stable structured information for subsequent emotion classification tasks based on spatiotemporal features.

[0062] S4. Construct a multi-scale spatiotemporal graph convolutional network as a feature extraction network, input the undirected graph structure into the feature extraction network, dynamically learn and update the adjacency matrix, express the characteristics of brain source signals in the undirected graph structure, and extract emotion-related features.

[0063] The multi-scale spatiotemporal graph convolutional network consists of two main modules: the spatial feature extraction module and the temporal feature extraction module. The spatial feature extraction module uses multiple convolution kernels of different scales to extract information from nodes in the undirected graph and unifies and merges the extracted features to form a complete spatial representation. The temporal feature extraction module uses the self-attention mechanism to dynamically focus on important time points, thereby extracting key features in the time series. Through the collaborative work of these two modules, a comprehensive feature expression is formed to provide support for subsequent emotion recognition tasks.

[0064] For the spatial feature extraction module, this embodiment proposes a global graph convolutional attention module (GGCAM), which integrates the multi-head self-attention mechanism and the graph convolutional network (GCN) to effectively aggregate local and global node feature information. GGCAM first calculates the relationship matrix between nodes through the multi-head self-attention mechanism. It contains 4 self-attention heads, each of which independently captures the node dependencies on different channel feature dimensions. The self-attention mechanism not only enhances the information flow between nodes, but also avoids the influence of a single feature dimension on the global relationship. Then, a learnable weight vector Integrate the multi-head outputs into the global attention relationship matrix A through weighted fusion global =wG. Since a dense attention matrix easily causes the problem of over-smoothing, this embodiment ensures the sparsity of the matrix by retaining the first 20% of connections in the adjacency matrix, thereby maintaining the sparsity of the undirected graph structure while aggregating global information.

[0065] Get the global attention matrix A global After that, this embodiment further calculates the corresponding Laplace matrix And introduce it into the graph convolutional network (GCN) for feature propagation. Through the graph convolution operation, the global node features are aggregated layer by layer. The specific formula is:

[0066]

[0067] Among them, O (l) Represents the node features of the lth layer input, O (l+1) is the node feature outputted by the l+1th layer, W (l) is a learnable weight matrix, and σ is a nonlinear activation function. Initial node feature O (0) is the original node feature X of the input meso This embodiment achieves effective fusion of local information and global information by introducing the GGCAM module, and enhances the emotion recognition ability of the model on the basis of global feature learning.

[0068] For the time feature extraction module, this embodiment adopts a multi-scale convolution time feature extraction mechanism to capture key signal features at different time scales. In order to adapt to time signals with different frequency components, this embodiment designs three different convolution kernel scales, corresponding to 1 times, 0.5 times and 0.25 times the signal sampling rate. This multi-scale design can ensure that features of different time lengths are extracted, thereby achieving a balance between short-term and long-term dependencies, and effectively improving the richness and robustness of feature representation.

[0069] Specifically, let the input signal be X time ∈R N×C×T , where N is the number of samples, C is the number of channels, and T is the time step. This embodiment designs three convolution operations, corresponding to different convolution kernel lengths K 1 , K 2 and K 3 , whose sizes are 1, 0.5 and 0.25 times of the original EEG signal sampling rate respectively. The convolution operation formula is:

[0070] X conv1 =Conv1D(X time ,W 1 ,K 1 , stride=1)

[0071] X conv2 =Conv1D(X time ,W 2 ,K 2 , stride=1)

[0072] X conv3 =Conv1D(X time ,W 3 ,K 3 , stride=1)

[0073] Among them, W 1 , W 2 , W 3 are the learnable parameters of the corresponding convolution kernels, K 1 , K 2 , K 3 are the sizes of the convolution kernels respectively. Through these convolution operations of different scales, it is ensured that the signal features in different time windows can be effectively captured. Subsequently, this embodiment concatenates the output features of the multi-scale convolution to form a comprehensive time feature representation:

[0074] X multi =concat(X conv1 ,X conv2 ,X conv3 )

[0075] The spliced ​​temporal feature X multi It includes feature representations at different time lengths, thereby improving the model's ability to perceive signals at different time scales. Through this multi-scale convolutional time feature extraction mechanism, this embodiment can effectively extract multi-level temporal features from complex time series data, providing a strong feature foundation for subsequent classification and recognition tasks.

[0076] S5. Based on the multi-scale spatiotemporal graph convolution and contrastive learning framework, an EEG emotion recognition model is built and trained to eliminate individual differences.

[0077] In view of the problems of low spatial resolution and obvious individual differences of EEG signals, this embodiment has built a contrastive learning module. In this framework, the feature extractor adopts a shared weights strategy. Whether the input is a positive sample pair (the same emotional data of different subjects) or a negative sample pair (different emotional data of the same subject), the input EEG signal passes through the same feature extractor. This weight sharing mechanism ensures that the features extracted between different subjects are consistent, thereby enhancing the alignment effect in contrastive learning. The shared feature extractor consists of two modules: a temporal feature extraction module and a spatial feature extraction module based on graph convolution. The temporal feature extraction module captures temporal information of different temporal resolutions through multi-scale convolution, and the spatial feature extraction module extracts the spatial features of EEG signals through a graph convolution network. The feature extractor combining these two modules can extract local and global temporal and spatial features from different scales, thereby comprehensively representing the emotional features in the EEG signal.

[0078] The EEG emotion recognition model training process includes the following steps:

[0079] S51, firstly, the original EEG data is divided according to the subjects and the emotion categories. For each emotion category, the same emotion data of different subjects are selected as positive sample pairs, and the different emotion data of the same subject are selected as negative sample pairs.

[0080] S52, then, the positive and negative sample pairs are respectively input into the feature extraction network with shared weights to extract high-dimensional feature representation. In order to ensure the consistency of feature extraction, all samples are processed by the same network structure and parameters. The similarity between the extracted high-dimensional feature vectors is measured by cosine similarity, and the calculation formula is as follows:

[0081]

[0082] Among them, z i and z jare the two vectors in the sample pair respectively.

[0083] S53. Next, use the contrast loss function to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs. The core goal of the loss function is to make samples of the same emotion cluster together and samples of different emotions stay away from each other by adjusting the feature space. The specific contrast loss function is:

[0084]

[0085] Among them, τ is the temperature parameter, is an indicator function used to ensure that k≠i. This loss function gradually optimizes the feature space by maximizing the similarity of positive samples and minimizing the similarity of negative samples to achieve generalization of emotion classification.

[0086] Specifically, for each input pair (positive or negative), the two input samples are passed through the same feature extractor to generate feature representations and Due to the weight sharing of the feature extractor, the extracted features have good comparability, which is conducive to optimizing the distance between similar samples in contrastive learning. This embodiment introduces a projection head to further optimize the contrast of features. The projection head maps the features to a low-dimensional space, which facilitates the fine optimization of the contrast loss on the feature distance between different samples. The projection head usually adopts a multi-layer perceptron (MLP) structure, which is specifically in the form of:

[0087] h(z)=W 2 σ(W 1 z),

[0088] Where W 1 and W 2 is a learnable weight matrix, and σ is a nonlinear activation function. The role of the projection head is to project the original feature z into a subspace that is more suitable for calculating the contrast loss to improve the effect of contrastive learning. In contrastive learning, the representations after the feature extractor and the projection head are denoted as and These projected features will be used to calculate the contrastive loss. The specific loss function is NT-Xent (Normalized Temperature-scaled CrossEntropy Loss), which is a commonly used loss function in contrastive learning to maximize the similarity between positive pairs and minimize the similarity between negative pairs.

[0089] The formula is as follows:

[0090]

[0091] in and is the representation after the feature extractor and projection head. represents the cosine similarity between two features. τ is the temperature parameter that controls the smoothness of the distribution. is an indicator function, ensuring that the denominator does not contain its own samples. N represents the batch size, that is, the number of sample pairs in each batch. Through this contrast loss, this embodiment can significantly enhance the alignment effect of similar emotions between different subjects, eliminate the interference caused by individual differences, and thus improve the generalization ability of sentiment classification.

[0092] Based on the same inventive concept, this embodiment also provides an EEG emotion recognition system based on multi-scale spatiotemporal graph convolution and contrastive learning, including the following modules:

[0093] The dataset construction module is used to collect the EEG signals of the subjects and pre-process them to obtain sample data that truly reflects the emotional processing process of the subjects; construct positive and negative sample pairs based on the emotion category and the subject identity for comparative learning; and construct a cross-subject dataset with one subject left out;

[0094] The brain source signal acquisition module uses a stacked autoencoder to extract the latent space representation of the EEG signal, divides the EEG signal into multiple functional brain areas based on medical prior information, and obtains the brain source signal after dimensionality reduction;

[0095] Undirected graph construction module, used to construct the undirected graph structure of brain source signals;

[0096] Convolutional network construction module, used to construct a multi-scale spatiotemporal graph convolutional network as a feature extraction network, input the undirected graph structure into the feature extraction network, dynamically learn and update the adjacency matrix, express the characteristics of brain source signals in the undirected graph structure, and extract emotion-related features;

[0097] The recognition model building and training module is used to build and train the EEG emotion recognition model based on the multi-scale spatiotemporal graph convolution and contrastive learning framework to eliminate individual differences.

[0098] It should be noted that the above modules are respectively used to implement the corresponding steps of the EEG emotion recognition method in this embodiment. For example, the brain source signal acquisition module uses a stacked autoencoder SAE to reduce the dimension of multi-channel brain source signals to reconstruct brain source signals and extract key features; the stacked autoencoder maps high-dimensional input data to a low-dimensional hidden space through unsupervised learning, while retaining the core information of the brain source signal.

[0099] For example, the undirected graph construction module regards each brain source signal channel as a node in the undirected graph, and the relationship between channel signals is mapped to the edge features of the undirected graph, converting the brain source signal obtained by dimensionality reduction from the stacked autoencoder into a data representation of the undirected graph structure.

[0100] For more detailed implementation processes, please refer to the description of steps S1-S5.

[0101] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. An EEG emotion recognition method based on multi-scale spatiotemporal graph convolution and contrastive learning, characterized in that: The following steps are involved: S1. Collect the EEG signals of the subjects and pre-process them to obtain sample data that truly reflects the emotional processing process of the subjects; construct positive and negative sample pairs based on the emotion category and the subject identity for comparative learning; Construct a cross-subject dataset with one subject left out; S2, using stacked autoencoders to extract the latent space representation of EEG signals, dividing EEG signals into multiple functional brain regions based on medical prior information, and obtaining brain source signals after dimensionality reduction; S3, constructing an undirected graph structure of brain source signals; S4. Construct a multi-scale spatiotemporal graph convolutional network as a feature extraction network, input the undirected graph structure into the feature extraction network, dynamically learn and update the adjacency matrix, express the characteristics of brain source signals in the undirected graph structure, and extract emotion-related features; S5. Based on the multi-scale spatiotemporal graph convolution and contrastive learning framework, an EEG emotion recognition model is built and trained to eliminate individual differences.

2. The EEG emotion recognition method according to claim 1, characterized in that: In step S2, a stacked autoencoder (SAE) is used to reduce the dimensionality of multi-channel brain source signals to reconstruct the brain source signals and extract key features. The stacked autoencoder maps high-dimensional input data to a low-dimensional latent space through unsupervised learning while retaining the core information of the brain source signals.

3. The EEG emotion recognition method according to claim 2, characterized in that: Step S2: Before the stacked autoencoder SAE performs dimensionality reduction on the multi-channel brain source signal, the SAE is first trained, including the following steps: The first layer of autoencoder training, the original brain source signal x(k) is passed into the first sparse autoencoder, and the network is trained to learn the first-order feature representation h of the input data (1) (k); The second layer of autoencoder training is used to represent the trained first-order features h (1) (k) is used as a new input and passed into the second sparse autoencoder to train it to learn the second-order feature representation h (2) (k); Classifier training, the second-order feature representation h (2) (k) As input, train a Softmax classifier to map high-level features to specific emotion classification labels; Each sparse autoencoder is also optimized by the reconstruction loss, which is used to measure the difference between the output of the sparse autoencoder and the input, and is defined as: in, It is the brain source signal reconstructed by the autoencoder; by minimizing the reconstruction loss, it ensures that the feature representation after dimensionality reduction still retains the key information of the original brain source signal; Finally, a stacked autoencoder network with two hidden layers is formed through layer-by-layer training.

4. The EEG emotion recognition method according to claim 1, characterized in that: In step S3, each brain source signal channel is regarded as a node in an undirected graph, and the relationship between the signals between the channels is mapped to the edge features of the undirected graph, converting the brain source signal obtained by dimensionality reduction from the stacked autoencoder into a data representation of the undirected graph structure.

5. The EEG emotion recognition method according to claim 4, characterized in that: Step S3 defines the edge features of the undirected graph by constructing an adjacency matrix. The edge weights are measured by the phase-locked values ​​of the signals between nodes, including: S31. For the multi-dimensional EEG channel brain source signal, convert it into a data representation of an undirected graph structure; First, the brain source signal of each channel is standardized. The standardized brain source signal is used as the input node feature in the graph convolution operation to construct the initial representation of the undirected graph. During the graph convolution process, the node features are optimized through the propagation and update mechanism in the graph convolution network. The graph convolution network will update the feature value of each node in combination with the edges of the undirected graph, thereby gradually extracting more complex and high-level features in space. S32. Use the phase-locking value PLV to measure the phase synchronization between the brain source signals of two EEG channels; when constructing an undirected graph structure, the phase-locking value is used as the weight of the edge to represent the mutual correlation between the channels; S33. A preset threshold τ is introduced to perform sparse processing on the matrix A of the phase-locked values, and a sparse adjacency matrix of the phase-locked values ​​is constructed; the undirected graph structure of EEG is constructed through the sparse adjacency matrix of the phase-locked values ​​to capture the static functional connectivity characteristics between brain regions and explore the dynamic connectivity changes under different emotional states.

6. The EEG emotion recognition method according to claim 1, characterized in that: The multi-scale spatiotemporal graph convolutional network constructed in step S4 includes a spatial feature extraction module and a temporal feature extraction module; The spatial feature extraction module uses multiple convolution kernels of different scales to extract information from nodes in the undirected graph and merges the extracted features to form a complete spatial representation. The temporal feature extraction module uses the self-attention mechanism to dynamically focus on important time points, thereby extracting key features in the time series.

7. The EEG emotion recognition method according to claim 1, characterized in that: Step S5 is a process of training the EEG emotion recognition model, comprising the following steps: First, the original EEG data are divided according to subjects and emotion categories. For each emotion category, the same emotion data of different subjects are selected as positive sample pairs, and the different emotion data of the same subject are selected as negative sample pairs. Then, the positive and negative sample pairs are input into the feature extraction network with shared weights to extract high-dimensional feature representations; the similarity between the extracted high-dimensional feature vectors is measured by cosine similarity; Next, the contrast loss function is used to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs. The loss function is used to adjust the feature space so that samples with the same emotion are clustered together, while samples with different emotions are far away from each other.

8. An EEG emotion recognition system based on multi-scale spatiotemporal graph convolution and contrastive learning, characterized in that: Includes the following modules: The data set construction module is used to collect the EEG signals of the subjects and pre-process them to obtain sample data that truly reflects the emotional processing process of the subjects; construct positive and negative sample pairs based on the emotion category and the subject identity for comparative learning; Construct a cross-subject dataset with one subject left out; The brain source signal acquisition module uses a stacked autoencoder to extract the latent space representation of the EEG signal, divides the EEG signal into multiple functional brain areas based on medical prior information, and obtains the brain source signal after dimensionality reduction; Undirected graph construction module, used to construct the undirected graph structure of brain source signals; Convolutional network construction module, used to construct a multi-scale spatiotemporal graph convolutional network as a feature extraction network, input the undirected graph structure into the feature extraction network, dynamically learn and update the adjacency matrix, express the characteristics of brain source signals in the undirected graph structure, and extract emotion-related features; The recognition model building and training module is used to build and train the EEG emotion recognition model based on the multi-scale spatiotemporal graph convolution and contrastive learning framework to eliminate individual differences.

9. The EEG emotion recognition system according to claim 8, characterized in that: The brain source signal acquisition module uses a stacked autoencoder (SAE) to reduce the dimensionality of multi-channel brain source signals in order to reconstruct brain source signals and extract key features. The stacked autoencoder maps high-dimensional input data to a low-dimensional latent space through unsupervised learning while retaining the core information of brain source signals.

10. The EEG emotion recognition system according to claim 8, characterized in that: The undirected graph construction module regards each brain source signal channel as a node in the undirected graph, and the relationship between channel signals is mapped to the edge features of the undirected graph, converting the brain source signal obtained by dimensionality reduction from the stacked autoencoder into a data representation of the undirected graph structure.

Citation Information

Patent Citations

  • Emotional dialogue generation method and device, and emotional dialogue model training method and device

    CN111966800A

  • Electroencephalogram emotion classification method based on multi-scale connectivity features and meta transfer learning

    CN113057657A

  • Facial expression recognition method and system

    CN113688715A

  • Multi-source conditional domain adaptive electroencephalogram emotion recognition method based on dynamic graph

    CN117668664A

  • Electroencephalogram emotion recognition method based on comparative learning in combination with double-layer high-speed noise and meiosis

    CN118114136A

Cited By

  • Electroencephalogram signal fatigue detection method and device and computer readable storage medium

    CN119700138A

  • Electroencephalogram signal fatigue detection method, device and computer readable storage medium

    CN119700138B

  • Cross-individual EEG emotion recognition method based on course learning and multi-source domain adaptation

    CN120470543A

  • Electroencephalogram emotion recognition method based on deep learning

    CN120514387A

  • Electroencephalogram signal source space function connection estimation system and method based on deep learning

    CN120763621A