Electroencephalogram emotion recognition method and system based on multi-task self-supervision and dynamic graph fusion network
By constructing a dynamic graph structure and multi-task self-supervised pre-training tasks, combined with cross-modal contrastive learning and attention mechanisms, the problems of static graph structure and modality simplification in EEG emotion recognition are solved, achieving efficient emotion recognition results.
Patent Information
- Application Number
- CN202511691921.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-10
AI Technical Summary
Existing EEG emotion recognition methods suffer from problems such as static graph structure, single modality, and reliance on large amounts of labeled data, making it difficult to effectively capture the dynamic characteristics of brain functional connectivity and the complementarity of multimodal information, thus limiting recognition performance.
We employ a method based on multi-task self-supervised and dynamic graph fusion networks. By constructing a dynamic graph structure, we fuse spatial distance and functional connectivity, combine multi-task self-supervised pre-training tasks and cross-modal contrastive learning, and utilize attention mechanisms to adaptively fuse multimodal features to achieve EEG emotion recognition.
It significantly improves the accuracy and generalization ability of emotion recognition, dynamically captures brain functional connectivity, makes full use of the complementarity of multimodal signals, and reduces the dependence on labeled data.
Smart Images

Figure CN121489503A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and physiological signal processing technology, specifically relating to a brainwave emotion recognition method and system based on multi-task self-supervised and dynamic graph fusion network. Background Technology
[0002] Electroencephalogram (EEG) signals are spontaneous potentials originating from the bioelectrical activity of neuronal groups in the brain. Their different frequency rhythms are closely related to the physiological and psychological states of the human body: delta waves (1-3 Hz) are commonly seen in deep sleep, theta waves (4-7 Hz) are associated with the early stages of sleep or drowsiness, alpha waves (8-13 Hz) reflect a relaxed state when quiet and with eyes closed, beta waves (14-30 Hz) characterize mental tension and concentration, while gamma waves (31-50 Hz) are associated with higher-level emotions and cognitive activities. These rhythms together constitute an electrophysiological spectrum reflecting the physiological and psychological states of the human body.
[0003] Based on these characteristics, electroencephalogram (EEG) signals acquired using non-invasive electrodes have become one of the key technologies for brain-computer interfaces (BCI). Early BCI technology primarily served as an adjunct rehabilitation tool for patients with neurological dysfunctions; today, its applications have expanded to multiple fields, including educational attention monitoring, sleep quality intervention, VR / AR interaction, and status monitoring of personnel in special operations. Against this backdrop, EEG-based emotion recognition, as a crucial link in enhancing the environmental perception and user understanding capabilities of BCI systems, has received widespread attention.
[0004] Currently, EEG emotion recognition methods mainly follow two major technical routes: traditional machine learning and deep learning. Traditional methods often rely on manually extracted time-frequency domain features, which have limitations in generalization ability and insufficient representation of the brain's dynamic responses. Although graph convolutional neural networks (GNNs) have been introduced in recent years to model brain region topological relationships, driving the development of this field, existing methods still have significant bottlenecks: First, most GNN models rely on fixed adjacency matrices (such as those based on spatial distance), making it difficult to capture the dynamic and time-varying characteristics of brain functional connections during emotion induction, thus limiting the model's ability to characterize complex neural patterns.
[0005] Secondly, existing methods mostly focus on a single EEG modality and fail to effectively integrate peripheral physiological signals such as eye movement, skin conductance, and electrocardiogram. They lack in-depth mining and semantic alignment mechanisms for the complementarity of multimodal information, which restricts further improvement in recognition performance.
[0006] Furthermore, models often neglect the joint modeling of functional connectivity and frequency band energy distribution in spatially adjacent brain regions within EEG signals, failing to fully extract discriminative spatiotemporal-frequency domain features. Simultaneously, the high cost and subjectivity of acquiring high-quality sentiment-annotated data also pose challenges to model training due to small sample sizes and insufficient generalization.
[0007] Therefore, developing a recognition method that can dynamically model brain region connectivity, effectively integrate multimodal information, and utilize self-supervised learning to mine discriminative features from unlabeled data has become the key to promoting the practical application of this technology. Summary of the Invention
[0008] To overcome the technical bottlenecks of static graph structures, single modality, and reliance on large amounts of labeled data in existing EEG emotion recognition technologies, this invention provides an EEG emotion recognition method and system based on multi-task self-supervised and dynamic graph fusion networks. This method can dynamically capture brain functional connectivity, make full use of the complementarity of multimodal signals, and learn robust features from unlabeled data, thereby significantly improving the accuracy and generalization ability of emotion recognition.
[0009] To achieve the above objectives, the present invention provides the following solution: A brainwave emotion recognition method based on multi-task self-supervised and dynamic graph fusion networks, the method comprising: Acquire and preprocess EEG and other physiological signals to extract multi-band energy features; A dynamic graph structure is constructed, with electrodes as nodes and frequency band energy as features, and spatial distance and functional connectivity are integrated to generate a dynamic adjacency matrix; Construct multi-task self-supervised pre-training tasks to learn general representations; The preprocessed raw data and the data generated by the self-supervised task are input in parallel into the dynamic graph convolutional neural network to achieve adaptive fusion of multimodal features; Based on the fused multimodal features, independent classification heads are set for each supervised task. Through optimization by joint loss function, the final output of EEG emotion recognition results is achieved.
[0010] The preferred method for generating a dynamic adjacency matrix, which uses electrodes as nodes and frequency band energy as characteristics, and integrates spatial distance and functional connectivity, is as follows: Constructing dynamic images of electroencephalograms , where the set of nodes Corresponding EEG channels, node feature matrix Energy values for each channel across multiple frequency bands Initial adjacency matrix Calculated through electrode spatial distance or functional connectivity; Utilizing attention mechanisms By integrating spatial and functional information, a dynamic adjacency matrix that changes over time is obtained. This allows for the construction of dynamic graphs that reflect the time-varying interactive relationships between EEG nodes.
[0011] Preferred multi-task self-supervised pre-training tasks include spatial jigsaw puzzle, frequency jigsaw puzzle, and cross-modal contrastive learning tasks; The spatial jigsaw puzzle task involves dividing EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffling their order, then training a model to predict the original arrangement. The frequency jigsaw puzzle task is used to divide EEG signals into multiple blocks according to different frequency bands and randomly shuffle their order, then train the model to restore the original frequency band order. A cross-modal contrastive learning task is used to construct positive sample pairs based on EEG signals of the same emotional sample and other modal physiological signals, and to construct negative sample pairs based on signals of different emotional samples, and to optimize feature space alignment through contrastive loss. One method involves dividing the EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffling their order, then training the model to predict the original arrangement. Total N C =62 EEG channels, and the features of each channel node are represented by multi-band energy and other modal feature vectors as follows: ; in, Indicates the capability of the corresponding frequency band. Indicates the number of brainwave channels. Indicates other modal characteristics synchronized with this channel; After dividing the 62 channels into 10 brain regions according to international standards, the features of each brain region were obtained by averaging or weighted summing the channel features within the region: ; in, C b Indicates the first b The set of channels contained in each brain region w i Channel weights are dynamically adjusted based on the importance of brain regions or node degree. The spatial jigsaw puzzle task is described as combining the features of 10 blocks. Randomly shuffle the order The original permutation index [1,2,…,10] is used as a pseudo-label, and the shuffled block features are... Input spatial classification head for prediction; One method involves dividing EEG signals into multiple blocks according to different frequency bands and randomly shuffling their order, then training a model to recover the original frequency band order. Preset EEG signals total Each channel has a multi-band energy characteristic as follows: ; Other modal features associated with each channel are denoted as Then, the EEG signals were divided into frequency bands. F =5 frequency bands ( Each frequency band block aggregates the energy characteristics and other modal characteristics of all channels in that frequency band: ; The frequency mosaicking task can be described as combining the features of 5 frequency band blocks. Randomly shuffle the order The original permutation index [1,2,…,5] is used as a pseudo-label, and the shuffled block features are... Input a frequency classification head for prediction; Among them, the method of constructing positive sample pairs based on EEG signals of the same emotional sample and other modal physiological signals, and constructing negative sample pairs based on signals of different emotional samples, and optimizing feature space alignment through contrastive loss includes: For the same emotional sample, EEG features reconstructed through brain segmentation and frequency band segmentation are paired with corresponding physiological features from other modalities to form positive sample pairs; corresponding features from different emotional samples are paired to form negative sample pairs. Specifically, let the same emotional sample... The features are obtained by reconstructing brain region features by segmenting and frequency bands. The corresponding other modal physiological characteristics are Positive sample pairs are represented as Different emotional samples The following forms negative sample pairs ; Features are mapped to a shared feature space using a projection head. A contrastive loss function based on cosine similarity is then used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. Specifically, the loss function for the contrastive learning task is expressed as follows: ; in, and yes and The normalized eigenvectors obtained after projection and normalization >0 represents the temperature parameter.
[0012] Preferably, the dynamic graph convolutional neural network includes a shared feature extraction module, which employs graph convolution operations based on Chebyshev multinomials and embeds a cross-modal attention module to adaptively fuse multimodal features; The operational expression for graph convolution based on Chebyshev polynomials is as follows: ; ; Where σ(·) is the activation function, X is the input node feature data, and β k These are the parameters learned during network training, Tk (·) is a Chebyshev polynomial of order K, where K is the order of the polynomial and takes the value 2, and λ max It is the largest eigenvalue of the Laplace matrix L. This is the scaled Laplace matrix; The cross-modal attention module is used to adaptively calculate the attention weights between EEG features and peripheral physiological signal features, and to perform weighted fusion of features to enhance emotion-related cross-modal collaborative patterns.
[0013] The present invention also provides an EEG emotion recognition system based on multi-task self-supervision and dynamic graph fusion network. The system is used to implement the aforementioned method and includes: a data acquisition and preprocessing module, a dynamic graph construction module, a multi-task self-supervision and management module, a dynamic graph fusion network module, and a model training and optimization module. The data acquisition and preprocessing module is used to acquire and preprocess EEG and other physiological signals, and extract multi-frequency energy features; The dynamic graph construction module is used to construct a dynamic graph structure, using electrodes as nodes and frequency band energy as features, and integrating spatial distance and functional connectivity to generate a dynamic adjacency matrix. The multi-task self-supervised supervision module is used to construct multi-task self-supervised pre-training tasks to learn general representations; The dynamic graph fusion network module is used to input the preprocessed raw data and the data generated by the self-supervised task into the dynamic graph convolutional neural network in parallel to achieve adaptive fusion of multimodal features; The model training and optimization module is used to set independent classification heads for each supervised task based on fused multimodal features, optimize through a joint loss function, and finally output the EEG emotion recognition result.
[0014] Preferably, the process of generating a dynamic adjacency matrix by using electrodes as nodes and frequency band energy as characteristics, and integrating spatial distance and functional connectivity, is as follows: Constructing dynamic images of electroencephalograms , where the set of nodes Corresponding EEG channels, node feature matrix Energy values for each channel across multiple frequency bands Initial adjacency matrix Calculated through electrode spatial distance or functional connectivity; Utilizing attention mechanisms By integrating spatial and functional information, a dynamic adjacency matrix that changes over time is obtained. This allows for the construction of dynamic graphs that reflect the time-varying interactive relationships between EEG nodes.
[0015] Preferred multi-task self-supervised pre-training tasks include spatial jigsaw puzzle, frequency jigsaw puzzle, and cross-modal contrastive learning tasks; The spatial jigsaw puzzle task involves dividing EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffling their order, then training a model to predict the original arrangement. The frequency jigsaw puzzle task is used to divide EEG signals into multiple blocks according to different frequency bands and randomly shuffle their order, then train the model to restore the original frequency band order. A cross-modal contrastive learning task is used to construct positive sample pairs based on EEG signals of the same emotional sample and other modal physiological signals, and to construct negative sample pairs based on signals of different emotional samples, and to optimize feature space alignment through contrastive loss. One method involves dividing the EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffling their order, then training the model to predict the original arrangement. Total N C =62 EEG channels, and the features of each channel node are represented by multi-band energy and other modal feature vectors as follows: ; in, Indicates the capability of the corresponding frequency band. Indicates the number of brainwave channels. Indicates other modal characteristics synchronized with this channel; After dividing the 62 channels into 10 brain regions according to international standards, the features of each brain region were obtained by averaging or weighted summing the channel features within the region: ; in, C b Indicates the first b The set of channels contained in each brain region w i Channel weights are dynamically adjusted based on the importance of brain regions or node degree. The spatial jigsaw puzzle task is described as combining the features of 10 blocks. Randomly shuffle the order The original permutation index [1,2,…,10] is used as a pseudo-label, and the shuffled block features are... Input spatial classification head for prediction; One method involves dividing EEG signals into multiple blocks according to different frequency bands and randomly shuffling their order, then training a model to recover the original frequency band order. Preset EEG signals total Each channel has a multi-band energy characteristic as follows: ; Other modal features associated with each channel are denoted as Then, the EEG signals were divided into frequency bands. F =5 frequency bands ( Each frequency band block aggregates the energy characteristics and other modal characteristics of all channels in that frequency band: ; The frequency mosaicking task can be described as combining the features of 5 frequency band blocks. Randomly shuffle the order The original permutation index [1,2,…,5] is used as a pseudo-label, and the shuffled block features are... Input a frequency classification head for prediction; Among them, the method of constructing positive sample pairs based on EEG signals of the same emotional sample and other modal physiological signals, and constructing negative sample pairs based on signals of different emotional samples, and optimizing feature space alignment through contrastive loss includes: For the same emotional sample, EEG features reconstructed through brain segmentation and frequency band segmentation are paired with corresponding physiological features from other modalities to form positive sample pairs; corresponding features from different emotional samples are paired to form negative sample pairs. Specifically, let the same emotional sample... The features are obtained by reconstructing brain region features by segmenting and frequency bands. The corresponding other modal physiological characteristics are Positive sample pairs are represented as Different emotional samples The following forms negative sample pairs ; Features are mapped to a shared feature space using a projection head. A contrastive loss function based on cosine similarity is then used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. Specifically, the loss function for the contrastive learning task is expressed as follows: ; in, and yes and The normalized eigenvectors obtained after projection and normalization >0 represents the temperature parameter.
[0016] Preferably, the dynamic graph convolutional neural network includes a shared feature extraction module, which employs graph convolution operations based on Chebyshev multinomials and embeds a cross-modal attention module to adaptively fuse multimodal features; The operational expression for graph convolution based on Chebyshev polynomials is as follows: ; ; Where σ(·) is the activation function, X is the input node feature data, and β kThese are the parameters learned during network training, T k (·) is a Chebyshev polynomial of order K, where K is the order of the polynomial and takes the value 2, and λ max It is the largest eigenvalue of the Laplace matrix L. This is the scaled Laplace matrix; The cross-modal attention module is used to adaptively calculate the attention weights between EEG features and peripheral physiological signal features, and to perform weighted fusion of features to enhance emotion-related cross-modal collaborative patterns.
[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the aforementioned method.
[0018] The present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements the aforementioned method.
[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: Dynamism: By integrating spatial and functional connections and introducing attention mechanisms, the graph structure can dynamically evolve with emotional states, more accurately depicting the spatiotemporal dependencies of the brain.
[0020] Complementarity: Through cross-modal contrastive learning and attention fusion modules, the complementary advantages of EEG and peripheral physiological signals in emotional representation are deeply explored, enhancing the discriminative power of features.
[0021] Efficiency: By adopting a multi-task self-supervised paradigm, robust feature representations are learned by making full use of unlabeled data, which effectively alleviates the dependence on large-scale labeled data and improves the generalization ability of the model.
[0022] Integration: The method is encapsulated into a complete system, and the collaborative relationship between each functional module is clearly defined, providing clear architectural support for the actual deployment and application of the technology. Attached Figure Description
[0023] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A schematic diagram illustrating the steps of an EEG emotion recognition method based on multi-task self-supervision and dynamic graph fusion network provided by the present invention; Figure 2This is a framework diagram of an EEG emotion recognition method based on multi-task self-supervised and dynamic graph fusion network provided by the present invention; Figure 3 This is a schematic diagram of a self-supervised pre-training task for an EEG emotion recognition method based on multi-task self-supervision and dynamic graph fusion network provided by the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] Example 1 This invention provides a brainwave emotion recognition method based on multi-task self-supervised and dynamic graph fusion networks, applicable to scenarios such as mental health assessment, brain-computer interfaces, and emotion recognition. The method includes the following steps: S1. Acquire the user's original EEG signal and at least one other modality of physiological signal, preprocess the original EEG signal and extract energy features of multiple frequency bands, and simultaneously extract features from the other modality of physiological signal to form a multimodal feature matrix aligned with the EEG signal time sequence. The original EEG signal With at least one other modality of physiological signal (Such as heart rate, skin conductance, respiration, etc.) are first collected synchronously. This is to improve the effectiveness and robustness of subsequent feature extraction; Preprocessing of EEG signals mainly involves... Apply multiple bandpass filters, each corresponding to δ (0.5–4Hz) θ (4–8 Hz) α (8–13 Hz) β (13–30 Hz) γ In the 30–50 Hz frequency band, to remove power frequency interference and low-frequency drift, unlike existing fixed bandpass filtering methods, this invention introduces an adaptive artifact suppression module. Dynamic artifact suppression is achieved by dynamically adjusting filter parameters based on the inter-channel covariance matrix. ; in, Indicates the corresponding frequency band type; The center frequency and bandwidth are represented by the frequency band. The determined bandpass filtering operation; This is an adaptive artifact suppression function that automatically adjusts the suppression weights based on the time-varying statistical characteristics of the filtered EEG signal (such as power spectral density and covariance matrix). Indicates at time The EEG signal after filtering and artifact suppression.
[0028] For the filtered EEG signal, energy features can be calculated using short-time Fourier transform or sliding window power, and are often generated using node features from subsequent self-supervised tasks. ; in, Indicates time Time corresponding frequency band The average energy; The length of the sliding window; It represents the time offset within the window, used to smooth the energy within a continuous time segment and reflect the local time-varying energy distribution characteristics of the signal.
[0029] By integrating the energy of each frequency band, the overall EEG energy feature vector can be obtained: in, Indicates at time Multi-band EEG energy characteristics are used to characterize the comprehensive time-frequency dynamic characteristics of EEG signals.
[0030] Other modal physiological signals The processing of signals such as heart rate, skin conductance, and respiration is similar to that of electroencephalogram (EEG) signals. Feature sequences can also be obtained through bandpass filtering and sliding window energy calculation. The processed brainwave energy characteristics Other modal features aligned with synchronization The features are concatenated to form a multimodal temporal feature matrix: in, N E For EEG frequency band features, N P For other modal feature dimensions, T This represents the length of the time series. This matrix is used as input for subsequent graph neural network or deep learning models.
[0031] S2. Constructing dynamic images of electroencephalograms (EEGs). , where the set of nodes Corresponding EEG channels, node feature matrix Energy values for each channel across multiple frequency bands Initial adjacency matrix It is calculated through electrode spatial distance or functional connectivity. Subsequently, an attention mechanism is utilized. By integrating spatial and functional information, a dynamic adjacency matrix that changes over time is obtained. This allows for the construction of dynamic graphs that reflect the time-varying interactive relationships between EEG nodes.
[0032] S3. Construct multi-task self-supervised pre-training tasks, including spatial jigsaw puzzle tasks, frequency jigsaw puzzle tasks, and cross-modal contrastive learning tasks; The spatial jigsaw puzzle task divides the EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffles their order, then trains the model to predict the original arrangement. The frequency jigsaw puzzle task divides the EEG signal into multiple blocks according to different frequency bands and randomly shuffles their order, then trains the model to restore the original frequency band order. The cross-modal contrastive learning task constructs positive sample pairs based on EEG signals of the same emotional sample and physiological signals of other modalities, and constructs negative sample pairs based on signals of different emotional samples, and optimizes feature space alignment through contrastive loss. S4. Input the preprocessed raw data and the data generated by the self-supervised task into the dynamic graph convolutional neural network in parallel. The network includes a shared feature extraction module, which adopts graph convolution operation based on Chebyshev multinomials and embeds a cross-modal attention module to adaptively fuse multimodal features. S5. In the classification module, independent classification heads are set for the main task of emotion recognition and their respective supervision tasks. The network is trained end-to-end through a multi-task joint loss function, and finally the emotion recognition result is output. In this embodiment, the construction of the initial adjacency matrix in step S2 is specifically as follows: The Euclidean distance is calculated based on the three-dimensional spatial coordinates of the electrodes on the scalp surface and converted into a spatial distance weight matrix using a Gaussian kernel function; the phase lock value between each pair of electrodes is calculated based on the EEG signal to generate a functional connectivity matrix. Specifically, the formula for calculating the spatial distance weight matrix is as follows: ; ; in Indicates the first i The three-dimensional coordinates of each electrode. Indicates the first i The first electrode and the first j The Euclidean distance between the electrodes Indicates the first i The first electrode and the firstj The spatial distance weights between the electrodes are mapped to weights using a Gaussian kernel function, where... The Gaussian kernel bandwidth can be adaptively adjusted according to the electrode distribution; On the other hand, the phase lock value (PLV) is used to measure the degree of phase synchronization between two EEG signals within a certain frequency band. Let the EEG signals... , The instantaneous phase is obtained by filtering and Hilbert transforming a certain frequency band. and The functional connectivity matrix is calculated as follows: ; The spatial distance weight matrix and the functional connectivity matrix are dynamically weighted and fused using an attention mechanism to generate a dynamic adjacency matrix. And based on this, construct the Laplace matrix. Used for graph convolution operations, where D It is a diagonal matrix.
[0033] In this embodiment, the spatial jigsaw puzzle task in step S3 is as follows: based on the international 10-20 system, 62 EEG electrodes are divided into 10 brain blocks. The frequency band energy features of the corresponding channel are aggregated in each block and other modal features are spliced together. After randomly shuffling the block order, the original permutation index is used as a pseudo-label and predicted by the spatial classification head.
[0034] Specifically, there are a total of N C =62 EEG channels, and the features of each channel node, consisting of multi-band energy and other modal feature vectors, can be represented as follows: ; in, Indicates the capability of the corresponding frequency band. Indicates the number of brainwave channels. This indicates other modal characteristics that are synchronized with this channel.
[0035] After dividing the 62 channels into 10 brain regions according to international standards, the features of each brain region were obtained by averaging or weighted summing the channel features within the region: ; in, C b Indicates the first b The set of channels contained in each brain region w i The channel weights can be dynamically adjusted based on the importance of brain regions or the degree of node involvement.
[0036] The spatial jigsaw puzzle task can be described as combining the features of 10 blocks. Randomly shuffle the order The original permutation index [1,2,…,10] is used as a pseudo-label, and the shuffled block features are... Input spatial classification head for prediction.
[0037] In this embodiment, the frequency jigsaw puzzle task in step S3 specifically involves dividing the five frequency bands of the EEG signal (δ, θ, α, β, γ) into five frequency band blocks. Each block contains the energy characteristics of all channels in the corresponding frequency band and other associated modal signals. After shuffling the order of the frequency band blocks, the original frequency band order is used as a pseudo-label for prediction by the frequency classification head.
[0038] The specific process and spatial jigsaw puzzle task types first assume that there are a total of EEG signals. Each channel has a multi-band energy characteristic as follows: ; Other modal features associated with each channel are denoted as Then, the EEG signals were divided into frequency bands. F =5 frequency bands ( Each frequency band block aggregates the energy characteristics and other modal characteristics of all channels in that frequency band: ; The frequency mosaicking task can be described as combining the features of 5 frequency band blocks. Randomly shuffle the order The original permutation index [1,2,…,5] is used as a pseudo-label, and the shuffled block features are... Input a frequency classification head for prediction.
[0039] In this embodiment, the cross-modal contrastive learning task in step S3 specifically includes: Under the same emotional sample, the EEG features reconstructed by brain segmentation and frequency band segmentation are combined with the corresponding other modal physiological features to form positive sample pairs; corresponding features under different emotional samples are combined to form negative sample pairs. Specifically, let's assume the same emotional sample The features are obtained by reconstructing brain region features by segmenting and frequency bands. The corresponding other modal physiological characteristics are Positive sample pairs can be represented as Different emotional samples The following forms negative sample pairs ; Features are mapped to a shared feature space using a projection head, and a contrastive loss function based on cosine similarity is used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs.
[0040] Specifically, the loss function for the contrastive learning task can be expressed as follows: ; in, and yes and The normalized eigenvectors obtained after projection and normalization >0 represents the temperature parameter (which controls the smoothness of the similarity distribution). In this embodiment, the feature extraction module in step S4 employs a K-order graph convolution based on Chebyshev polynomials, the operation of which is expressed as: ; ; Where σ(·) is the activation function, X is the input node feature data, and β k These are the parameters learned during network training, T k (·) is a Chebyshev polynomial of order K, where K is the order of the polynomial and takes the value 2, and λ max It is the largest eigenvalue of the Laplace matrix L. This is the scaled Laplace matrix. I It is an identity matrix.
[0041] In this embodiment, the cross-modal attention module mentioned in step S4 is used to adaptively calculate the attention weight between EEG features and peripheral physiological signal features, and to perform weighted fusion of features to enhance the cross-modal collaborative mode related to emotion.
[0042] Specifically, firstly, the time The following EEG feature matrix Peripheral physiological signal feature matrix Mapping these to the query, key, and value spaces respectively yields: ; in Indicates the number of brainwave channels (e.g., 64 channels). This represents the feature dimension of each channel (such as the feature length after convolution or time-frequency transformation). This indicates the number of channels for peripheral signals (such as the number of sensors for heart rate, electrooculography, electromyography, etc.). Indicates the corresponding feature dimension; A set of query vectors representing EEG modalities. , Let each represent a set of key and value vectors representing a peripheral physiological modality. , , The linear projection weight matrix is learnable. This represents the dimension of the key vector, used to control the feature size of the attention space.
[0043] Then, adaptively scaling cross-modal attention weights were designed, as follows: ; In the above formula Indicates time t At that time, from multimodal signals P to EEG signals E The cross-modal attention weight matrix reflects the multimodal signal. P Features of EEG signals E The degree of attention allocation can be used for subsequent weighted fusion of cross-modal features. Among them, For normalization function, This is an adaptive temperature factor used to adjust the smoothness of the attention distribution.
[0044] This method differs from traditional methods that use fixed scaling or fixed temperature. τ (t) is driven by EEG multi-frequency energy (τ becomes smaller when the energy is high, making attention sharper), thereby automatically enhancing the weight of significant cross-modal coupling signals at the moment of emotional activation.
[0045] Secondly, attention weight matrices are used to weight peripheral physiological features to obtain cross-modal feature mappings. : ; To further enhance the flexibility of local fusion, the study also introduces a node-level gating mechanism, using gating vectors... Controlling the fusion ratio of EEG and peripheral modal information: in, The Sigmoid activation function is used. This indicates a channel pooling operation. Indicates feature splicing, This represents element-wise multiplication. For learnable parameter matrix and bias terms, This is the output of the fused cross-modal features.
[0046] In this embodiment, the multi-task joint loss function in step S5 is a weighted sum of the cross-entropy loss of the emotion recognition main task and the loss of each of the respective supervised tasks. The self-supervised task loss includes the cross-entropy loss of the spatial jigsaw puzzle task and the frequency jigsaw puzzle task, as well as the normalized cross-entropy loss of the cross-modal contrastive learning task.
[0047] In this embodiment, the training process in step S5 adopts a dynamic monitoring and adaptive adjustment mechanism. When the difference between the accuracy of the training set and the test set exceeds a preset threshold, it is determined to be overfitting, and the learning rate reduction and data resampling strategy are automatically started.
[0048] Example 2 like Figures 1-3 As shown, this invention provides a brainwave emotion recognition method based on multi-task self-supervised and dynamic graph fusion networks, comprising the following steps: I. In the stage of acquiring raw EEG data and other physiological signals and performing preprocessing, subjects wore 62-channel electrode caps conforming to the international 10-20 lead system, sat quietly in an electromagnetically shielded room, and watched psychologically validated emotion-evoking videos in sequence, continuously inducing various emotional states such as happiness, sadness, fear, disgust, anger, fun, joy, and neutrality. Simultaneously, EEG signals were recorded at a sampling rate of 1000 Hz, and continuous waveforms were equidistantly segmented into 1-second non-overlapping segments. At the same time, eye movement tracking, high-definition facial images, speech signals, skin conductance responses, respiratory waveforms, and electrocardiogram signals were acquired through a preamplifier and synchronization interface, forming a multimodal data stream strictly aligned with the timing of the EEG segments. Subsequently, baseline removal, 50 Hz notch filtering, and bandpass filtering of 0.5-50 Hz were performed on the EEG segments. Then, differential entropy or short-time Fourier transform was calculated to obtain the energy characteristics of five subbands: δ (1-3 Hz), θ (4-7 Hz), α (8-13 Hz), β (14-30 Hz), and γ (31-50 Hz).
[0049] Simultaneously, facial action unit features are extracted from facial images, Mel frequency cepstral coefficient features are extracted from speech, and temporal statistical features are extracted from electrodermatology, respiration, and electrocardiogram. Finally, these features are spliced together to form a multimodal feature matrix that corresponds one-to-one with the EEG segments. The training set and test set are divided according to the subject-dependent 9:6, 16:8, 21:7 or leave-one-out cross-validation strategy.
[0050] II. In the dynamic graph construction phase, EEG channels are defined as graph nodes, and node features are represented by energy values of each frequency band. The initial adjacency matrix construction considers both electrode spatial distance and functional connectivity, where functional connectivity is quantified using phase-locked values. To fully utilize the spatiotemporal topological characteristics of multimodal emotional data and compensate for the spatial information loss problem in traditional EEG analysis, this embodiment adopts a cross-modal fusion framework based on dynamic graph neural networks. The differential entropy features or short-time Fourier transform features of multi-channel EEG signals are divided into five frequency bands: delta, theta, alpha, beta, and gamma bands, and mapped to a graph node feature matrix. Each node corresponds to an electrode in a specific brain region. Simultaneously, facial action unit features and Mel frequency cepstral coefficient features are introduced as cross-modal constraints, and a multimodal feature tensor is constructed through temporal alignment.
[0051] Dynamic adjacency matrix The construction integrates two correlations: first, a spatial distance matrix based on the three-dimensional coordinates of the electrodes, which retains prior information about the brain's anatomical structure after Gaussian kernel function transformation; second, a functional connectivity matrix based on phase-locked values, used to capture the dynamic synchronization characteristics of neural activity. The weights of spatial distance and phase-locked values are adaptively integrated through an attention mechanism to generate a time-varying adjacency matrix, which is then combined with the degree matrix to construct a Laplacian operator. : ; Where D is a diagonal matrix, A(t) is a time-varying adjacency matrix generated by fusing spatial distance and phase-locking values through an attention mechanism, and t represents the time index. This dynamic graph structure enables graph neural networks to simultaneously learn the spatial topological evolution of EEG signals and the temporal dependencies of cross-modal features, effectively extracting neurobehavioral coordination patterns highly correlated with emotional states.
[0052] Third, in the self-supervised pre-training task construction stage, to improve the accuracy and generalization of EEG emotion recognition, three self-supervised pre-training tasks—spatial jigsaw puzzle, frequency jigsaw puzzle, and contrastive learning—were constructed based on multimodal data. Specifically, in the spatial jigsaw puzzle task, multimodal data was divided into blocks by brain regions and shuffled to transform it into a permutation recognition task to learn EEG spatial coordination patterns; in the frequency jigsaw puzzle task, the five frequency bands δ, θ, α, β, and γ were treated as independent blocks, and pseudo-labels were assigned after shuffling the frequency band order to form a frequency band permutation prediction task; in the cross-modal contrastive learning task, positive and negative sample pairs were constructed by reorganizing brain regions and frequency band blocks to maximize the feature similarity of samples with the same emotion to standardize the cross-modal feature space. The outputs of the three tasks and the original data were subjected to max-min normalization to ensure that the feature value range fell within the 0-1 interval, eliminating dimensional differences and accelerating gradient convergence.
[0053] Based on the embodiments of this invention, the spatial jigsaw puzzle task involves dividing multimodal data into blocks according to different brain regions and shuffling them to obtain a randomly arranged EEG emotion data, and assigning a label (pseudo-label) to this arrangement. To divide the multimodal data into blocks based on the physical location of brain regions, firstly, based on the international 10-20 system, the 62-channel EEG is divided into 10 brain regions. Each region aggregates the frequency band energy of the corresponding channel and splices cross-modal features such as facial AU and speech MFCC. Then, the block order is randomly shuffled, and the original arrangement index is used as a pseudo-label, transforming the task into an arrangement recognition task. This process forces the graph convolutional neural network to learn the spatial semantic associations between brain regions, such as the collaborative patterns between the prefrontal cortex and the limbic system, and captures emotion-related spatial topological features by reconstructing the correct arrangement order.
[0054] The frequency jigsaw puzzle task involves dividing EEG emotion data into blocks according to different frequency ranges and shuffling them to obtain a randomly arranged EEG emotion data set. A pseudo-label is then assigned to this arrangement. The multimodal data is divided into five frequency bands according to the EEG frequency bands δ, θ, α, β, and γ. Each band contains the energy characteristics of all channels in the corresponding frequency band and associated cross-modal signals. After shuffling the frequency bands and assigning pseudo-labels, the task is transformed into a frequency band arrangement prediction task. The model learns the correct frequency band order and automatically selects frequency components key to emotion classification, such as the increased θ wave energy during sadness, thus enhancing the frequency discriminative power of the features.
[0055] The cross-modal contrastive learning task involves randomly recombining EEG data of the same emotional sample, segmenting it by brain region and frequency band, to generate positive sample pairs (combinations of different segments of the same emotion). Combinations of segments from different emotional samples are then used as negative sample pairs. A projection head maps multimodal features to a common space, maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs. This forces the spatial frequency features of EEG to align with facial and speech features in terms of emotional semantics; for example, the alpha wave brain region features of a pleasant emotion are adjacent to the alpha wave features of a smiling emotion in the feature space. This further normalizes the model's feature space, learns the intrinsic representation of EEG emotional signals, and improves the accuracy of EEG emotion recognition.
[0056] The raw data and extracted features from the three tasks were subjected to maximum and minimum value normalization, scaling the numerical range of each feature to the [0,1] interval. This eliminated the influence of dimensions between different modalities, improved the convergence speed of gradient descent, and enhanced the model's learning efficiency for spatial distribution patterns, frequency domain characteristics, and cross-modal correlation information. Through the collaborative constraint mechanism among the three tasks, the model can simultaneously grasp the spatial topological structure, frequency-sensitive characteristics, and multimodal semantic alignment relationships of EEG signals, thereby significantly improving the discriminative power of emotion-related features and the model's generalization ability.
[0057] Fourth, in the feature extraction and fusion stage using a dynamic graph convolutional neural network (GCNN), the normalized raw data and three self-supervised data sets are input into the GCNN in parallel. This network includes a graph input module, a feature extraction module, and a classification module. The graph input module is responsible for receiving and processing input data from multiple tasks in parallel. The shared feature extraction module, based on a graph convolutional neural network constructed using Chebyshev multinomials, performs unified feature extraction on the data from all tasks. The classification module consists of multiple independent graph convolutional classifiers, each receiving its corresponding features and performing a specific classification task.
[0058] Specifically, the graph input module, based on a multi-task learning framework, inputs the raw EEG data corresponding to the main emotion recognition task, along with self-supervised data from three auxiliary tasks—spatial mosaic, frequency mosaic, and contrastive learning—in parallel into the graph convolutional neural network. The spatial mosaic task data contains the spatial topological relationships of EEG signals, the frequency mosaic task data characterizes the correlation of its frequency band features, and the contrastive learning task strengthens cross-modal semantic consistency by constructing positive and negative sample pairs. Through a multi-task collaborative mechanism, the three auxiliary tasks jointly guide the main task to effectively integrate complementary information from spatially adjacent brain structures and the intrinsic correlation of frequency features in emotion recognition, thereby improving recognition accuracy. Simultaneously, during the multi-task iterative learning process, knowledge complementarity is formed among the tasks, further enhancing the model's ability to discriminate EEG emotion signals.
[0059] The feature extraction module, as a core component of multi-task sharing, employs a dynamic graphical convolutional network based on Chebyshev polynomials. Building upon the effective fusion of spatial and frequency information, it further introduces a cross-modal attention module to construct a complementary enhancement mechanism between EEG and peripheral physiological signals such as eye movements and electrodermal activity. This module dynamically strengthens the synergistic association between EEG signals and peripheral physiological signals by adaptively learning the importance weights between features of different modalities. When EEG prefrontal alpha wave energy and eye movement pupil dilation signals occur simultaneously, the attention mechanism significantly increases their weights, thereby highlighting the neuro-behavioral synergistic pattern under pleasurable emotions. This cross-modal information selective enhancement strategy enables the module to accurately capture complementary emotional cues in multimodal data and, combined with the spatial topology and frequency features extracted through multi-task self-supervised learning, jointly enhances the representational ability of EEG emotional signals and the overall model discrimination accuracy. Feature extraction is performed using a dynamic graphical convolutional neural network based on Chebyshev polynomials, the expression of which is: ; ; Where σ(·) is the activation function, X is the input data, and β is the input data. k These are the parameters learned during network training, T k (·) is a Chebyshev polynomial of order K, where K = 2 and λmax It is the largest eigenvalue of the Laplacian matrix L. For each task, the input data has been manually extracted features, therefore the input feature dimension is . The output feature dimension is In this way, the output feature dimension is higher than the input feature dimension, with richer high-dimensional feature information. High-dimensional feature information has a high degree of abstraction and contains richer semantic information.
[0060] The classification module utilizes a fully connected layer to design a multi-task-specific classification framework, achieving fine-grained mapping of the feature space: the spatial jigsaw puzzle task is equipped with a spatial classification head with an output dimension of 128, which predicts 128 spatial category labels, such as the connection patterns between the prefrontal and parietal lobes, thereby performing fine-grained modeling of the spatial topological information of EEG signals; the frequency jigsaw puzzle task corresponds to a frequency classification head with an output dimension of 120, which predicts 120 frequency category labels, such as the energy distribution of the δ, θ, α, β, and γ frequency bands, thereby enhancing the feature extraction capability of emotion-sensitive frequency bands; the contrastive learning task... The algorithm utilizes a projection head and employs cosine similarity to measure the similarity between cross-modal sample pairs, such as EEG and eye-tracking feature pairs or EEG and electrodermal feature pairs. By maximizing the similarity of positive cross-modal sample pairs under the same emotion, such as the feature combination of prefrontal β wave and electrodermal response in the anxiety state, it promotes the alignment of EEG features and peripheral physiological signals in the emotional semantic space, thereby significantly improving the discriminative performance of cross-modal features. The main emotion recognition task sets the output dimension based on the number of emotion categories in the SEED, SEED-IV, and MPED datasets to generate the final emotion prediction label.
[0061] Leveraging the knowledge-sharing mechanism of the feature extraction module, the emotion classification head integrates three key pieces of information when generating predicted labels: first, the functional connectivity patterns of brain regions captured by the spatial classification head; second, the frequency band energy variation patterns extracted by the frequency classification head; and third, the cross-modal semantic associations established through the contrastive learning task. This cross-task knowledge transfer strategy allows predicted emotion labels to naturally contain triple information: brain spatial topology, frequency characteristics, and cross-modal synergy. This not only significantly improves the accuracy of EEG emotion recognition but also enhances the model's generalization ability to multimodal data.
[0062] The main emotion recognition task classifies raw EEG emotion data from the SEED, SEED-IV, and MPED datasets using an emotion classification head. Its output dimension strictly matches the number of emotion categories in the dataset. The emotion classification head is responsible for outputting the final predicted emotion label, while the spatial classification head and frequency classification head generate predicted spatial and frequency category labels, respectively. The projection head is only used to perform cross-modal feature mapping and does not generate labels. Specifically, the contrastive learning task constructs cross-modal positive and negative sample pairs. For example, EEG and electrodermal feature pairs for the same emotion are used as positive samples, and EEG and eye-tracking feature pairs for different emotions are used as negative samples. A cosine similarity metric is used to drive the clustering of cross-modal features in the emotion space. The cross-modal semantic associations formed in this process are directly integrated into the prediction logic of the emotion classification head through knowledge sharing from the feature extraction module. For example, when the model learns the strong correlation between weakened prefrontal alpha waves and pupillary dilation in pleasurable emotions, it assigns a higher weight to this cross-modal feature combination during emotion classification. This triple feature representation, which integrates spatial topology, frequency characteristics, and cross-modal collaboration, ultimately leads to a significant improvement in the accuracy of EEG emotion recognition using the method of this invention.
[0063] V. In the system implementation and training phase of the graphical convolutional neural network, the system of this invention can be deployed on an electronic device containing a processor and a memory. The memory stores a computer program, which, when executed by the processor, implements the functions of all the above modules.
[0064] A dynamic graph convolutional neural network (GCNN) is jointly trained using multimodal data from the training set and a pre-built self-supervised pre-training task. During training, the network performance is validated using a test set. If overfitting is detected, the GCNN is retrained by adjusting the learning rate, ultimately resulting in a pre-trained dynamic graph convolutional neural network. The number of training iterations is set to 100 to 300, with 100 to 300 samples input per task. The cross-entropy loss function is used as the optimization objective for model training. The cross-entropy loss function is: ; Where L is the cross-entropy loss, N is the number of samples, and y i For one-hot encoding of labels, H is the classification head or projection head, F is the graph convolutional neural network, and X is the image. iThe input data is used as input data. The final loss function consists of a weighted sum of losses from multiple tasks, with the specific weights manually adjusted based on actual experimental results to ensure that the features learned by the model are more focused on solving the main task of emotion recognition. The graph convolutional neural network uses adaptive time estimation as the optimizer, with an initial learning rate set to 0.001 and a random deactivation mechanism, with a node retention ratio set to 0.5. During training, 100 samples are extracted from the training set each time to generate corresponding training data for each self-supervised task, and then the data is input into the graph input module of the constructed graph convolutional neural network. After obtaining the feature representation of the EEG emotion signal through the feature extraction module, the feature is sequentially subjected to nonlinear transformation, downsampling, logarithmic operation, and random deactivation. Finally, the processed feature is sent to the corresponding classification module to complete the classification task.
[0065] During the model training phase, a multi-task joint optimization strategy is employed to improve learning efficiency: the spatial jigsaw puzzle and frequency jigsaw puzzle tasks utilize cross-entropy loss functions to measure the difference between the predicted labels output by the classification head and the preset pseudo-labels, thereby guiding the model to effectively learn the spatial topological structure and frequency distribution characteristics of EEG signals. The multimodal contrastive learning task uses normalized cross-entropy loss to increase the similarity between cross-modal positive sample pairs and strengthen the semantic alignment of EEG with peripheral signals such as eye movement and electrodermal activity at the feature level. The main emotion recognition task calculates cross-entropy loss based on real emotion labels, directly optimizing the model's accuracy in emotion classification. The loss functions of the above four tasks are weighted and fused, and then backpropagated through adaptive optimizers such as AdamW to iteratively update the parameters of the dynamic graph convolutional neural network, aiming to gradually reduce the overall loss and continuously improve the model's ability to represent the features of EEG emotion signals. A dynamic monitoring and adaptive adjustment mechanism is introduced during training: after each complete training cycle, the system evaluates the current model performance using the test set every 10 iterations, continuously comparing the trends in accuracy between the training set and the test set. If the difference between the two exceeds 20%, the model is considered overfitted. In this case, a dual optimization strategy is automatically initiated: firstly, mechanisms such as cosine annealing are used to gradually reduce the learning rate, weakening the model's overfitting to the training data; secondly, 100 samples are randomly selected from the training set to construct a new batch, improving the model's generalization performance through data distribution perturbation. If the accuracy difference remains within a reasonable range, a graph convolutional neural network model with high generalization ability and high recognition accuracy can be obtained after 1000 iterations of training. The preprocessed multimodal fusion feature data of the test set is directly input into the trained dynamic graph convolutional neural network for classification. Statistical analysis of the output results yields the network's final recognition accuracy on the test set. If the test accuracy does not meet the preset accuracy requirements, the learning rate needs to be adjusted and the training process restarted to iteratively optimize the graph convolutional neural network until it meets the predetermined performance standards.
[0066] Example 3 This invention provides an EEG emotion recognition system based on multi-task self-supervision and dynamic graph fusion network. The system is used to implement the method described in Embodiment 1. The system includes: a data acquisition and preprocessing module, a dynamic graph construction module, a multi-task self-supervision and management module, a dynamic graph fusion network module, and a model training and optimization module. The data acquisition and preprocessing module is used to acquire and preprocess EEG and other physiological signals, and extract multi-band energy features; The dynamic graph construction module is used to construct dynamic graph structures, using electrodes as nodes and frequency band energy as features, and integrating spatial distance and functional connectivity to generate a dynamic adjacency matrix. The multi-task self-supervised supervision module is used to construct multi-task self-supervised pre-training tasks to learn general representations; The dynamic graph fusion network module is used to input preprocessed raw data and data generated by self-supervised tasks into the dynamic graph convolutional neural network in parallel to achieve adaptive fusion of multimodal features; The model training and optimization module is used to set independent classification heads for each supervised task based on fused multimodal features, optimize through a joint loss function, and finally output the EEG emotion recognition results.
[0067] In this embodiment, the process of generating a dynamic adjacency matrix by fusing spatial distance and functional connectivity, using electrodes as nodes and frequency band energy as features, is as follows: Constructing dynamic images of electroencephalograms , where the set of nodes Corresponding EEG channels, node feature matrix Energy values for each channel across multiple frequency bands Initial adjacency matrix Calculated through electrode spatial distance or functional connectivity; Utilizing attention mechanisms By integrating spatial and functional information, a dynamic adjacency matrix that changes over time is obtained. This allows for the construction of dynamic graphs that reflect the time-varying interactive relationships between EEG nodes.
[0068] In this embodiment, the multi-task self-supervised pre-training tasks include spatial jigsaw puzzle, frequency jigsaw puzzle, and cross-modal contrastive learning tasks; The spatial jigsaw puzzle task involves dividing EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffling their order, then training a model to predict the original arrangement. The frequency jigsaw puzzle task is used to divide EEG signals into multiple blocks according to different frequency bands and randomly shuffle their order, then train the model to restore the original frequency band order. A cross-modal contrastive learning task is used to construct positive sample pairs based on EEG signals of the same emotional sample and other modal physiological signals, and to construct negative sample pairs based on signals of different emotional samples, and to optimize feature space alignment through contrastive loss. One method involves dividing the EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffling their order, then training the model to predict the original arrangement. Total N C =62 EEG channels, and the features of each channel node are represented by multi-band energy and other modal feature vectors as follows: ; in, Indicates the capability of the corresponding frequency band. Indicates the number of brainwave channels. Indicates other modal characteristics synchronized with this channel; After dividing the 62 channels into 10 brain regions according to international standards, the features of each brain region were obtained by averaging or weighted summing the channel features within the region: ; in, C b Indicates the first b The set of channels contained in each brain region w i Channel weights are dynamically adjusted based on the importance of brain regions or node degree. The spatial jigsaw puzzle task is described as combining the features of 10 blocks. Randomly shuffle the order The original permutation index [1,2,…,10] is used as a pseudo-label, and the shuffled block features are... Input spatial classification head for prediction; One method involves dividing EEG signals into multiple blocks according to different frequency bands and randomly shuffling their order, then training a model to recover the original frequency band order. Preset EEG signals total Each channel has a multi-band energy characteristic as follows: ; Other modal features associated with each channel are denoted as Then, the EEG signals were divided into frequency bands. F =5 frequency bands ( Each frequency band block aggregates the energy characteristics and other modal characteristics of all channels in that frequency band: ; The frequency mosaicking task can be described as combining the features of 5 frequency band blocks. Randomly shuffle the order The original permutation index [1,2,…,5] is used as a pseudo-label, and the shuffled block features are... Input a frequency classification head for prediction; Among them, the method of constructing positive sample pairs based on EEG signals of the same emotional sample and other modal physiological signals, and constructing negative sample pairs based on signals of different emotional samples, and optimizing feature space alignment through contrastive loss includes: For the same emotional sample, EEG features reconstructed through brain segmentation and frequency band segmentation are paired with corresponding physiological features from other modalities to form positive sample pairs; corresponding features from different emotional samples are paired to form negative sample pairs. Specifically, let the same emotional sample... The features are obtained by reconstructing brain region features by segmenting and frequency bands. The corresponding other modal physiological characteristics are Positive sample pairs are represented as... Different emotional samples The following forms negative sample pairs ; Features are mapped to a shared feature space using a projection head. A contrastive loss function based on cosine similarity is used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. Specifically, the loss function for the contrastive learning task can be expressed as follows: ; in, and yes and The normalized eigenvectors obtained after projection and normalization >0 represents the temperature parameter.
[0069] In this embodiment, the dynamic graph convolutional neural network includes a shared feature extraction module, which employs graph convolution operations based on Chebyshev multinomials and embeds a cross-modal attention module to adaptively fuse multimodal features. The operational expression for graph convolution based on Chebyshev polynomials is as follows: ; ; Where σ(·) is the activation function, X is the input node feature data, and β k These are the parameters learned during network training, T k (·) is a Chebyshev polynomial of order K, where K is the order of the polynomial and takes the value 2, and λ max It is the largest eigenvalue of the Laplace matrix L. This is the scaled Laplace matrix. I It is the identity matrix; The cross-modal attention module is used to adaptively calculate the attention weights between EEG features and peripheral physiological signal features, and to perform weighted fusion of features to enhance emotion-related cross-modal collaborative patterns.
[0070] This invention captures the spatiotemporal dynamic characteristics of the brain through dynamic graph structures, learns robust features from unlabeled data using multi-task self-supervision, and fully exploits the complementarity of multi-source signals by combining cross-modal attention, which significantly improves the accuracy and generalization ability of EEG emotion recognition.
[0071] Example 4 The present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in Embodiment 1.
[0072] Example 5 The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in Embodiment 1.
[0073] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A brainwave emotion recognition method based on multi-task self-supervised and dynamic graph fusion networks, characterized in that, The method includes: Acquire and preprocess EEG and other physiological signals to extract multi-band energy features; A dynamic graph structure is constructed, with electrodes as nodes and frequency band energy as features, and spatial distance and functional connectivity are integrated to generate a dynamic adjacency matrix; Construct multi-task self-supervised pre-training tasks to learn general representations; The preprocessed raw data and the data generated by the self-supervised task are input in parallel into the dynamic graph convolutional neural network to achieve adaptive fusion of multimodal features; Based on the fused multimodal features, independent classification heads are set for each supervised task. Through optimization by joint loss function, the final output of EEG emotion recognition results is achieved.
2. The method according to claim 1, characterized in that, The method for constructing a dynamic graph structure, using electrodes as nodes and frequency band energy as features, and integrating spatial distance and functional connectivity to generate a dynamic adjacency matrix is as follows: Constructing dynamic images of electroencephalograms , where the set of nodes Corresponding EEG channels, node feature matrix Energy values for each channel across multiple frequency bands Initial adjacency matrix Calculated through electrode spatial distance or functional connectivity; Utilizing attention mechanisms By integrating spatial and functional information, a dynamic adjacency matrix that changes over time is obtained. This allows for the construction of dynamic graphs that reflect the time-varying interactive relationships between EEG nodes.
3. The method according to claim 1, characterized in that, Multi-task self-supervised pre-training tasks include spatial jigsaw puzzle, frequency jigsaw puzzle, and cross-modal contrastive learning tasks; The spatial jigsaw puzzle task involves dividing EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffling their order, then training a model to predict the original arrangement. The frequency jigsaw puzzle task is used to divide EEG signals into multiple blocks according to different frequency bands and randomly shuffle their order, then train the model to restore the original frequency band order. A cross-modal contrastive learning task is used to construct positive sample pairs based on EEG signals of the same emotional sample and other modal physiological signals, and to construct negative sample pairs based on signals of different emotional samples, and to optimize feature space alignment through contrastive loss. One method involves dividing the EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffling their order, then training the model to predict the original arrangement. Total N C =62 EEG channels, and the features of each channel node are represented by multi-band energy and other modal feature vectors as follows: ; in, Indicates the capability of the corresponding frequency band. Indicates the number of brainwave channels. Indicates other modal characteristics synchronized with this channel; After dividing the 62 channels into 10 brain regions according to international standards, the features of each brain region were obtained by averaging or weighted summing the channel features within the region: ; in, C b Indicates the first b The set of channels contained in each brain region w i Channel weights are dynamically adjusted based on the importance of brain regions or node degree. The spatial jigsaw puzzle task is described as combining the features of 10 blocks. Randomly shuffle the order The original permutation index [1,2,…,10] is used as a pseudo-label, and the shuffled block features are... Input spatial classification head for prediction; One method involves dividing EEG signals into multiple blocks according to different frequency bands and randomly shuffling their order, then training a model to recover the original frequency band order. Preset EEG signals total Each channel has a multi-band energy characteristic as follows: ; Other modal features associated with each channel are denoted as Then, the EEG signals were divided into frequency bands. F =5 frequency bands ( Each frequency band block aggregates the energy characteristics and other modal characteristics of all channels in that frequency band: ; The frequency mosaicking task can be described as combining the features of 5 frequency band blocks. Randomly shuffle the order The original permutation index [1,2,…,5] is used as a pseudo-label, and the shuffled block features are... Input a frequency classification head for prediction; Among them, the method of constructing positive sample pairs based on EEG signals of the same emotional sample and other modal physiological signals, and constructing negative sample pairs based on signals of different emotional samples, and optimizing feature space alignment through contrastive loss includes: For the same emotional sample, EEG features reconstructed through brain segmentation and frequency band segmentation are paired with corresponding physiological features from other modalities to form positive sample pairs; corresponding features from different emotional samples are paired to form negative sample pairs. Specifically, let the same emotional sample... The features are obtained by reconstructing brain region features by segmenting and frequency bands. The corresponding other modal physiological characteristics are Positive sample pairs are represented as Different emotional samples The following forms negative sample pairs ; Features are mapped to a shared feature space using a projection head. A contrastive loss function based on cosine similarity is then used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. Specifically, the loss function for the contrastive learning task is expressed as follows: ; in, and yes and The normalized eigenvectors obtained after projection and normalization >0 represents the temperature parameter.
4. The method according to claim 1, characterized in that, The dynamic graph convolutional neural network includes a shared feature extraction module that employs graph convolution operations based on Chebyshev multinomials and embeds a cross-modal attention module to adaptively fuse multimodal features. The operational expression for graph convolution based on Chebyshev polynomials is as follows: ; ; Where σ(·) is the activation function, X is the input node feature data, and β k These are the parameters learned during network training, T k (·) is a Chebyshev polynomial of order K, where K is the order of the polynomial and takes the value 2, and λ max It is the largest eigenvalue of the Laplace matrix L. This is the scaled Laplace matrix; The cross-modal attention module is used to adaptively calculate the attention weights between EEG features and peripheral physiological signal features, and to perform weighted fusion of features to enhance emotion-related cross-modal collaborative patterns.
5. A brainwave emotion recognition system based on multi-task self-supervised and dynamic graph fusion network, the system being used to implement the method described in any one of claims 1-4, characterized in that, The system includes: a data acquisition and preprocessing module, a dynamic graph construction module, a multi-task self-supervision and management module, a dynamic graph fusion network module, and a model training and optimization module; The data acquisition and preprocessing module is used to acquire and preprocess EEG and other physiological signals, and extract multi-frequency energy features; The dynamic graph construction module is used to construct a dynamic graph structure, using electrodes as nodes and frequency band energy as features, and integrating spatial distance and functional connectivity to generate a dynamic adjacency matrix. The multi-task self-supervised supervision module is used to construct multi-task self-supervised pre-training tasks to learn general representations; The dynamic graph fusion network module is used to input the preprocessed raw data and the data generated by the self-supervised task into the dynamic graph convolutional neural network in parallel to achieve adaptive fusion of multimodal features; The model training and optimization module is used to set independent classification heads for each supervised task based on fused multimodal features, optimize through a joint loss function, and finally output the EEG emotion recognition result.
6. The system according to claim 5, characterized in that, The method for constructing a dynamic graph structure, using electrodes as nodes and frequency band energy as features, and integrating spatial distance and functional connectivity to generate a dynamic adjacency matrix is as follows: Constructing dynamic images of electroencephalograms , where the set of nodes Corresponding EEG channels, node feature matrix Energy values for each channel across multiple frequency bands Initial adjacency matrix Calculated through electrode spatial distance or functional connectivity; Utilizing attention mechanisms By integrating spatial and functional information, a dynamic adjacency matrix that changes over time is obtained. This allows for the construction of dynamic graphs that reflect the time-varying interactive relationships between EEG nodes.
7. The system according to claim 5, characterized in that, Multi-task self-supervised pre-training tasks include spatial jigsaw puzzle, frequency jigsaw puzzle, and cross-modal contrastive learning tasks; The spatial jigsaw puzzle task involves dividing EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffling their order, then training a model to predict the original arrangement. The frequency jigsaw puzzle task is used to divide EEG signals into multiple blocks according to different frequency bands and randomly shuffle their order, then train the model to restore the original frequency band order. A cross-modal contrastive learning task is used to construct positive sample pairs based on EEG signals of the same emotional sample and other modal physiological signals, and to construct negative sample pairs based on signals of different emotional samples, and to optimize feature space alignment through contrastive loss. One method involves dividing the EEG electrodes into multiple blocks according to anatomical brain regions and randomly shuffling their order, then training the model to predict the original arrangement. Total N C =62 EEG channels, and the features of each channel node are represented by multi-band energy and other modal feature vectors as follows: ; in, Indicates the capability of the corresponding frequency band. Indicates the number of brainwave channels. Indicates other modal characteristics synchronized with this channel; After dividing the 62 channels into 10 brain regions according to international standards, the features of each brain region were obtained by averaging or weighted summing the channel features within the region: ; in, C b Indicates the first b The set of channels contained in each brain region w i Channel weights are dynamically adjusted based on the importance of brain regions or node degree. The spatial jigsaw puzzle task is described as combining the features of 10 blocks. Randomly shuffle the order The original permutation index [1,2,…,10] is used as a pseudo-label, and the shuffled block features are... Input spatial classification head for prediction; One method involves dividing EEG signals into multiple blocks according to different frequency bands and randomly shuffling their order, then training a model to recover the original frequency band order. Preset EEG signals total Each channel has a multi-band energy characteristic as follows: ; Other modal features associated with each channel are denoted as Then, the EEG signals were divided into frequency bands. F =5 frequency bands ( Each frequency band block aggregates the energy characteristics and other modal characteristics of all channels in that frequency band: ; The frequency mosaicking task can be described as combining the features of 5 frequency band blocks. Randomly shuffle the order The original permutation index [1,2,…,5] is used as a pseudo-label, and the shuffled block features are... Input a frequency classification head for prediction; Among them, the method of constructing positive sample pairs based on EEG signals of the same emotional sample and other modal physiological signals, and constructing negative sample pairs based on signals of different emotional samples, and optimizing feature space alignment through contrastive loss includes: For the same emotional sample, EEG features reconstructed through brain segmentation and frequency band segmentation are paired with corresponding physiological features from other modalities to form positive sample pairs; corresponding features from different emotional samples are paired to form negative sample pairs. Specifically, let the same emotional sample... The features are obtained by reconstructing brain region features by segmenting and frequency bands. The corresponding other modal physiological characteristics are Positive sample pairs are represented as Different emotional samples The following forms negative sample pairs ; Features are mapped to a shared feature space using a projection head. A contrastive loss function based on cosine similarity is then used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. Specifically, the loss function for the contrastive learning task is expressed as follows: ; in, and yes and The normalized eigenvectors obtained after projection and normalization >0 represents the temperature parameter.
8. The system according to claim 5, characterized in that, The dynamic graph convolutional neural network includes a shared feature extraction module that employs graph convolution operations based on Chebyshev multinomials and embeds a cross-modal attention module to adaptively fuse multimodal features. The operational expression for graph convolution based on Chebyshev polynomials is as follows: ; ; Where σ(·) is the activation function, X is the input node feature data, and β k These are the parameters learned during network training, T k (·) is a Chebyshev polynomial of order K, where K is the order of the polynomial and takes the value 2, and λ max It is the largest eigenvalue of the Laplace matrix L. This is the scaled Laplace matrix; The cross-modal attention module is used to adaptively calculate the attention weights between EEG features and peripheral physiological signal features, and to perform weighted fusion of features to enhance emotion-related cross-modal collaborative patterns.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 4.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 4.
Citation Information
Cited By
State emotion classification method and system and model training method
CN121765648A
Emotion recognition method based on multi-band adaptive graph convolution
CN121997275A