A CNN-Transformer-based EEG emotion recognition method, device, and system

By segmenting and feature mapping EEG signals using a CNN-Transformer network, and combining positional encoding and multi-scale Transformer modules, the problem of capturing spatiotemporal features and global temporal dependencies in EEG emotion recognition is solved, thereby improving recognition efficiency and accuracy.

CN120724310BActive Publication Date: 2025-10-28CHANGCHUN UNIV OF SCI & TECH

Patent Information

Application Number
CN202511134081.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-10-28
Estimated Expiration
2045-08-14

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively capture the spatiotemporal characteristics and global time dependencies of EEG signals, resulting in low efficiency in EEG emotion recognition and high computational resource requirements.

Method used

The CNN-Transformer network is used to extract multiple frequency band features by segmenting the EEG signal and mapping them into a two-dimensional EEG image. Combined with position encoding, hemispherical attention and multi-scale Transformer modules, comprehensive spatiotemporal features are extracted.

Benefits of technology

It significantly improves the efficiency and accuracy of EEG emotion recognition, and solves the problems of insufficient feature coupling modeling and inadequate utilization of neuroscience priors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724310B_ABST
    Figure CN120724310B_ABST
Patent Text Reader

Abstract

This invention discloses a CNN-Transformer-based EEG emotion recognition method, device, and system in the fields of emotion computing and brain-computer interface technology. The method includes: acquiring raw EEG signals; segmenting the raw EEG signals and extracting features from multiple frequency bands within each segment; extracting differential entropy features and power spectral density features from the multi-frequency band features within a set time window; extracting comprehensive spatiotemporal frequency features based on the hybrid features of differential entropy and power spectral density using a CNN-Transformer model; and inputting the spatiotemporal frequency features into a multilayer perceptron to obtain the emotion recognition result. This technical solution effectively captures the spatiotemporal frequency interaction characteristics of EEG signals using a CNN-Transformer network, solving the problems of insufficient feature coupling modeling and inadequate utilization of neuroscience priors in traditional methods, thereby significantly improving the efficiency and accuracy of EEG emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of emotion computing and brain-computer interface technology, specifically to an EEG emotion recognition method, device, and system based on CNN-Transformer. Background Technology

[0002] To achieve accurate emotion recognition, researchers have employed various methods, including emotion recognition technologies based on facial expressions, voice, body movements, and physiological signals. However, behavioral signals such as facial expressions and voice are easily influenced by individual subjective consciousness; individuals can intentionally feign their true emotions. Furthermore, different cultural backgrounds and environmental factors can also interfere with the accuracy of emotion recognition. In contrast, physiological signals directly reflect internal physiological changes and are less susceptible to subjective manipulation. Among these, electroencephalography (EEG), due to its high temporal resolution and direct origin from the cerebral cortex, can capture subtle differences in emotional changes, making it an important research subject in the field of emotion recognition.

[0003] In the field of EEG signal processing, Convolutional Neural Networks (CNNs) are widely used in emotion recognition due to their powerful spatial feature extraction capabilities. However, CNNs often struggle to capture long-range temporal dependencies when processing time-series data. To address this issue, the Transformer model was introduced into EEG emotion recognition. Transformers, through their self-attention mechanism, can capture long-range dependencies in the data, thus better learning global contextual information. However, while Transformers can handle global temporal dynamics, they often lack the ability to extract local features like CNNs when processing spatially structured data. Furthermore, Transformers have high computational resource requirements, especially when processing long-term sequences, which can lead to inefficient training.

[0004] Based on the aforementioned problems, combining the advantages of CNNs and Transformers becomes an ideal solution. By combining the local feature extraction capabilities of CNNs with the advantages of Transformers in modeling global temporal dependencies, the spatial and temporal features of EEG signals can be processed more effectively. Furthermore, neuroscience research shows that the left and right hemispheres of the human brain exhibit a natural asymmetric response to emotional stimuli. This hemispheric lateralization in neuroscience provides a theoretical basis for emotion recognition algorithms that incorporate physiological priors. Therefore, how to fully utilize neuroscience prior knowledge to design a deep learning model that can both capture the spatiotemporal features of EEG signals and efficiently model global temporal dependencies has become a key problem urgently needing to be solved in the field of EEG emotion recognition. Summary of the Invention

[0005] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0006] Therefore, the purpose of this invention is to provide an EEG emotion recognition method, device, and system based on CNN-Transformer. By utilizing the CNN-Transformer network to effectively capture the spatiotemporal frequency interaction characteristics in EEG signals, this invention solves the problems of insufficient feature coupling modeling and inadequate utilization of neuroscience priors in traditional methods, thereby significantly improving the efficiency and accuracy of EEG emotion recognition.

[0007] To address the aforementioned technical problems, according to one aspect of the present invention, the present invention provides the following technical solution:

[0008] A CNN-Transformer-based EEG emotion recognition method, characterized by the following steps:

[0009] S1. Segment the original EEG signal and extract multiple frequency band features within each segment;

[0010] S2. Map the multiple frequency band features into a two-dimensional brain electrode map to simultaneously retain time domain, frequency domain, and spatial domain information;

[0011] S3. The two-dimensional EEG is processed using the CNN-Transformer model to extract comprehensive spatiotemporal features;

[0012] S4. Generate emotion recognition results based on the spatiotemporal features.

[0013] As a preferred embodiment of the CNN-Transformer-based EEG emotion recognition method described in this invention, in step S1, the original EEG signal is segmented using a sliding window technique, wherein the sliding window duration is 0.5 seconds, the sampling frequency is 128Hz, and the differential entropy features and power spectral density features of the four frequency bands θ, α, β, and γ are calculated simultaneously.

[0014] As a preferred embodiment of the CNN-Transformer-based EEG emotion recognition method described in this invention, the differential entropy is used to quantify the complexity and randomness of the EEG signal, as shown in the following formula:

[0015]

[0016] in, Indicates EEG signal The differential entropy, yes The probability density function, when Follows Gaussian distribution hour, The above formula can be simplified to:

[0017] in The mean, For variance, It is the natural logarithm.

[0018] As a preferred embodiment of the CNN-Transformer-based EEG emotion recognition method described in this invention, the power spectral density is used to describe the energy distribution of the EEG signal in each frequency band, and is estimated using the Welch method. Divided into There are overlapping sub-segments, each sub-segment having a length of [number]. And apply a window function to each segment. To reduce spectrum leakage;

[0019] Window function The energy normalization factor is defined as:

[0020] ;

[0021] For the Sub-segment Its periodic diagram The calculation formula is as follows:

[0022] ;

[0023] in Sampling frequency, As a frequency variable, subsequently, by applying all The signal is obtained by averaging the periodicity of each segment. In frequency Power spectral density estimation on: .

[0024] As a preferred embodiment of the CNN-Transformer-based EEG emotion recognition method of the present invention, step S2, which maps the multiple frequency band features into a two-dimensional EEG to simultaneously retain time domain, frequency domain, and spatial domain information, specifically includes: mapping differential entropy features and power spectral density features to a two-dimensional EEG after cross-stacking along the channel dimension, wherein each data point on the two-dimensional EEG corresponds to a feature value at a specific electrode location.

[0025] As a preferred embodiment of the EEG emotion recognition method based on CNN-Transformer described in this invention, in step S3, the CNN-Transformer model includes: a position encoding module, a hemispherical attention module, and a multi-scale Transformer module;

[0026] The location encoding module inputs a two-dimensional EEG image. Global normalization is performed, and the calculation formula is as follows:

[0027]

[0028] in, , The normalized data matrix, Representative Channel Features across the entire 2D space This represents the mean value across all spatial locations within the channel. This represents the standard deviation of all spatial locations within the channel;

[0029] Normalized features It is composed of all time points The normalized characteristic matrix X t The three-dimensional tensor formed by ′, with respect to the normalized features Embedded learnable positional encoding pos, denoted as , , Indicates in Based on this, location coding information is embedded, which can greatly enhance the perception of spatial structure;

[0030] The hemispherical attention module uses dual-scale deep convolution to extract spatial and frequency domain features from the left and right hemispheres, and combines the attention mechanism to model the frequency domain dependency. First, different receptive fields are used... of Convolution kernel pair A dual-scale convolutional transformation is performed to capture local high-frequency and global low-frequency information. Then, a 1D convolution and a depthwise convolution are used to extract the query, key, and value vectors. The calculation formula is as follows:

[0031] Among them, the dimensions of Q, K, and V are... Consistent, This indicates batch normalization, SELU indicates activation function, Reshape indicates data reshaping, and Split indicates splitting.

[0032] The obtained Q, K, V are then adjusted for dimensionality using Reshape(·) and normalized using Norm(·), and then subjected to matrix multiplication. and learnable temperature coefficient After calculating the attention weights, the softmax(·) activation function is applied to normalize them into a probability distribution. Finally, the attention weights are compared with... Multiplication yields the output:

[0033]

[0034] Output After adjusting the dimensions using Reshape(·), and then processing by the Dropout layer and linear projection layer, the enhanced feature representation is obtained through residual connections and BatchNorm(·) normalization. :

[0035]

[0036] Since the feature extraction process for the left and right hemispheres is exactly the same, the above operations can be performed separately for each hemisphere to extract their respective enhanced features. and This process involves concatenating and fusing enhanced features to obtain a final representation with stronger expressive power. ;

[0037] The multi-scale Transformer module achieves dynamic capture of short-term, medium-term, and long-term temporal features by fusing multi-scale convolution and self-attention mechanisms. First, in order to capture global temporal dependencies, the input signal... Perform a dimensional transformation using permute(·) to obtain To adapt it to the multi-head self-attention mechanism MHA:

[0038]

[0039] Subsequently, the time-domain weighted feature representation is computed using the multi-head attention mechanism (MHA):

[0040]

[0041] in, This is the weighted self-attention output; to further compress the time dimension and extract key features, adaptive pooling is used to obtain preliminary time-dependent features:

[0042]

[0043] After obtaining preliminary time-dependency features, to further extract information at different time scales, multi-scale 1D convolution is employed to simultaneously capture short-term, medium-term, and long-term time-dependency features. The convolution kernel size is set to... ,in These correspond to short-term, medium-term, and long-term, respectively. The feature extraction process is as follows:

[0044]

[0045] in, and These represent short-term, medium-term, and long-term time-dependent characteristics, respectively. Convolution operations representing different scales ensure that features at different time scales have consistent dimensions.

[0046] As a preferred embodiment of the EEG emotion recognition method based on CNN-Transformer described in this invention, a lightweight weight generation module is proposed to adaptively adjust the importance of features at different time scales. The calculation method of the lightweight weight generation module is as follows:

[0047]

[0048] in, The weight values ​​are calculated by a lightweight weight generation module and are used to dynamically adjust the importance of features at different time scales. Perform feature transformation. Enhance nonlinear expressive power, is a learnable weight matrix, and b is a learnable bias term. Normalized feature distribution, This is applied to weight normalization to ensure that its value range is within a certain range. Between, for At each time scale, all weight tensors are stacked using the stack(·) operation and normalized to a global weight tensor:

[0049] Based on weight tensor The features at different scales are weighted and summed to obtain the final global fusion features. :

[0050]

[0051]

[0052] in, This indicates the removal of redundant dimensions. This represents element-wise multiplication.

[0053] To further enhance the ability to represent time-domain features, the features will be... With self-attention output Perform element-wise multiplication and further fuse the results through convolution:

[0054] ;

[0055] After feature fusion, the input is processed further through a feedforward network (FFN). First, the input... The model is first transformed into hidden layer dimensions through the first linear transformation layer (linear1), then non-linearly transformed using the ReLU activation function, and then subjected to Dropout to enhance its generalization ability. Next, the output passes through the second linear transformation layer (linear2) to compress the dimensions, followed by another Dropout operation. Finally, it undergoes LayerNorm normalization to obtain the final output containing rich contextual information. The formula is as follows:

[0056] .

[0057] An apparatus for implementing an EEG emotion recognition method based on CNN-Transformer, the apparatus comprising:

[0058] The data acquisition unit is used to acquire raw EEG signals in real time. This unit uses an EEG cap and signal acquisition circuit, and has an adaptive impedance matching function to adapt to different users' scalp contact conditions.

[0059] The preprocessing unit, connected to the data acquisition unit, performs sliding window segmentation processing on the raw EEG signal and extracts the differential entropy features and power spectral density features of the four frequency bands θ, α, β, and γ.

[0060] The feature mapping unit is used to cross-stack the frequency band features output by the preprocessing unit and map them into a two-dimensional brain electrode map, where each data point corresponds to the feature value of the corresponding electrode position, thereby preserving time domain, frequency domain and spatial domain information.

[0061] The model training unit is used to input the two-dimensional EEG generated by the feature mapping unit into the CNN-Transformer-based deep learning model for training. This unit integrates position encoding, hemispherical attention and multi-scale Transformer functions, and is equipped with a fully connected layer and loss function calculation module to achieve efficient extraction and representation of emotion-related features.

[0062] The emotion recognition unit is used to perform emotion recognition on real-time acquired EEG signals based on the trained CNN-Transformer model, and to post-process the model output to generate multi-dimensional emotion classification results.

[0063] The communication unit is used to transmit emotion recognition results to external terminals or cloud platforms to achieve real-time data interaction.

[0064] An EEG emotion recognition system based on CNN-Transformer includes the aforementioned device, a user interface, a data storage unit, and a cloud collaboration module.

[0065] The user interface is used to display EEG signal waveforms, emotion recognition results, and signal quality feedback in real time.

[0066] The data storage unit is used to store the raw EEG signal, processed feature data, and historical identification records.

[0067] The cloud-based collaboration module is used to upload emotion recognition results to a cloud platform for data analysis and supports remote model updates.

[0068] Compared with the prior art, the beneficial effects of the present invention are: the present invention utilizes the CNN-Transformer network to effectively capture the spatiotemporal frequency interaction characteristics in EEG signals, solving the problems of insufficient feature coupling modeling and insufficient utilization of neuroscience priors in traditional methods, thereby significantly improving the efficiency and accuracy of EEG emotion recognition. Attached Figure Description

[0069] To more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and detailed embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0070] Figure 1 This is a schematic diagram of the EEG emotion recognition method based on CNN-Transformer provided in an embodiment of the present invention;

[0071] Figure 2 This is a schematic diagram of data preprocessing provided in an embodiment of the present invention;

[0072] Figure 3 This is a schematic diagram of feature mapping provided in an embodiment of the present invention;

[0073] Figure 4 This is a diagram of the CNN-Transformer model framework provided in this embodiment of the invention;

[0074] Figure 5 This is a schematic diagram of the hemispherical attention module provided in an embodiment of the present invention;

[0075] Figure 6 This is a schematic diagram of the multi-scale Transformer module provided in an embodiment of the present invention;

[0076] Figure 7 This is a diagram illustrating the accuracy of the method provided in this embodiment of the invention on the DREAMER dataset;

[0077] Figure 8 This is a schematic diagram of the F1 score of the method provided in this embodiment of the invention on the DREAMER dataset;

[0078] Figure 9 This is a schematic diagram illustrating the accuracy of the method provided in this embodiment of the invention on the DEAP dataset;

[0079] Figure 10 This is a schematic diagram of the F1 score of the method provided in this embodiment of the invention on the DEAP dataset;

[0080] Figure 11 This is a schematic diagram of the structure of an EEG emotion recognition device based on CNN-Transformer provided in an embodiment of the present invention;

[0081] Figure 12 This is a schematic diagram of an EEG emotion recognition system based on CNN-Transformer provided in an embodiment of the present invention. Detailed Implementation

[0082] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0083] This invention provides an EEG emotion recognition method, device, and system based on CNN-Transformer. It effectively captures the spatiotemporal frequency interaction characteristics of EEG signals by utilizing the CNN-Transformer network, which solves the problems of insufficient feature coupling modeling and insufficient utilization of neuroscience priors in traditional methods, thereby significantly improving the efficiency and accuracy of EEG emotion recognition.

[0084] Figure 1 This is a flowchart illustrating the EEG emotion recognition method based on CNN-Transformer provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes the following:

[0085] S101, EEG signal acquisition: EEG signals are acquired through the brain electrode cap and amplified and suppressed by the signal acquisition circuit to ensure the stability and reliability of signal quality.

[0086] Furthermore, the EEG signal covers multiple channels, each corresponding to the electrode location in a different brain region, and records time-series data at a set sampling frequency to reflect the dynamic changes in brain activity.

[0087] S102. Data Preprocessing: Data preprocessing includes segmenting and filtering the EEG signal, and manually extracting frequency domain features to remove noise and retain key information. Further... Figure 2 This is a schematic diagram of the data preprocessing process provided in an embodiment of the present invention, such as... Figure 2 As shown, data preprocessing includes the following steps:

[0088] Data filtering employs a sliding window technique to segment the raw EEG signal and extract the frequency band characteristics of each segment. Specifically, the sliding window duration is set to 0.5 seconds, the sampling frequency is 128Hz, and a Butterworth bandpass filter is used to filter the segmented EEG signal to extract features from four frequency bands: θ (4-8Hz), α (8-13Hz), β (13-30Hz), and γ (30-50Hz), thereby preserving key information from different frequency bands.

[0089] Feature extraction involves further calculating the differential entropy (DE) and power spectral density (PSD) features within each frequency band to extract the time-frequency information of the EEG signal. Specifically, the differential entropy (DE) is used to quantify the complexity and randomness of the EEG signal, as shown in the following formula:

[0090]

[0091] Where H( ) indicates EEG signal The differential entropy, yes The probability density function, when Follows Gaussian distribution hour, The above formula can then be simplified to:

[0092] in The mean, For variance, It is the natural logarithm;

[0093] Specifically, power spectral density (PSD) is used to describe the energy distribution of EEG signals across different frequency bands. It is estimated using the Welch method, which analyzes the EEG signals... Divided into There are overlapping sub-segments, each sub-segment having a length of [number]. And apply a window function to each segment. (For example, the Hanning window) to reduce spectral leakage; the energy normalization factor of the window function is defined as:

[0094]

[0095] For the Sub-segment ( ), its periodic diagram The calculation formula is as follows:

[0096]

[0097] in Sampling frequency, As a frequency variable, subsequently, by applying all The signal is obtained by averaging the periodicity of each segment. In frequency Power spectral density estimation on:

[0098]

[0099] S103. Data Mapping: Data mapping is based on preprocessed features to construct a two-dimensional EEG graph, so as to simultaneously preserve the time domain, frequency domain and spatial domain information of the EEG signal.

[0100] Furthermore, Figure 3 This is a schematic diagram of data mapping and EEG provided in an embodiment of the present invention, such as... Figure 3 As shown, the data mapping includes the following:

[0101] Based on the differential entropy (DE) and power spectral density (PSD) extracted in step S102, they are cross-stacked along the channel dimension to enhance the diversity and discriminative power of the features. Specifically, if the EEG signal contains C channels and each channel corresponds to F frequency bands (such as θ, α, β, γ), then a feature matrix of shape C×F can be constructed.

[0102] Subsequently, based on the spatial topology of the EEG electrodes, the feature matrix is ​​mapped into a two-dimensional EEG electrode map, making it spatially correspond to the actual electrode locations. This preserves the time, frequency, and spatial information of the EEG signal. During this mapping process, each data point corresponds to a 0.5s feature value at a specific electrode location to ensure that the feature distribution conforms to the EEG electrode layout, thereby enhancing the model's ability to model the spatial relationships of brain regions. Furthermore... Figure 3 The mapped two-dimensional EEG map is also shown, which serves as input features for further training and inference of the CNN-Transformer model.

[0103] S104, CNN-Transformer Model Training: The CNN-Transformer model is used to extract features from the input two-dimensional EEG images. Further, Figure 4 This is a schematic diagram of the CNN-Transformer hybrid model architecture provided in an embodiment of the present invention. This model is used for... Figure 3 The generated two-dimensional EEG images are used for feature extraction, and emotion recognition results are output. The processing flow includes the following core modules:

[0104] Location encoding module: Performs channel-level normalization on the input features and embeds learnable location encoding vectors to enhance the model's ability to perceive the spatial topology of brain electrodes;

[0105] Hemispherical attention module: It employs dual-scale depthwise separable convolution to extract local spatial-frequency domain features from the left and right hemispheres respectively; furthermore, it models frequency domain dependencies through a cross-hemispheric attention mechanism to enhance the interaction and complementarity of features from the left and right hemispheres.

[0106] Multi-scale Transformer module: It integrates multi-scale temporal convolution and self-attention mechanism to capture local details and global temporal patterns of EEG signals respectively; further, it introduces adaptive feature fusion weights to dynamically aggregate multi-scale temporal features and achieve highly robust emotional state modeling.

[0107] Furthermore, the location encoding module aims to enhance the CNN-Transformer model's ability to perceive the spatial topology of brain electrodes and optimize feature extraction performance.

[0108] In embodiments of the present invention, the input two-dimensional EEG map is first normalized at the channel level to eliminate feature scale differences between different electrode channels and ensure the stability of the data distribution. The normalized feature matrix is ​​then embedded with a learnable location encoding vector to introduce spatial location information into the feature representation, enabling the model to capture the relative relationships between different electrodes.

[0109] Specifically, for input two-dimensional EEG maps Global normalization is performed, and the calculation formula is as follows:

[0110]

[0111] in, , The normalized data matrix, Representative Channel Features throughout the entire 2D space, This represents the mean value across all spatial locations within the channel. This represents the standard deviation of all spatial locations within the channel;

[0112] Normalized features It is composed of all time points Normalized characteristic matrix The constructed three-dimensional tensor, with respect to the normalized features Embedded learnable positional encoding pos, denoted as , , Indicates in By embedding location coding information on top of this, the perception of spatial structure can be greatly enhanced.

[0113] Furthermore, in order to fully explore the spatial and frequency characteristics of EEG signals in the left and right hemispheres, this invention proposes a hemispherical attention module. Figure 5 This is a schematic diagram of the hemispherical attention module provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the module first uses different receptive fields. of Convolution kernel pair A dual-scale convolutional transformation is performed to capture local high-frequency and global low-frequency information. Then, QKV is extracted through a 1D convolution and a depthwise convolution. The calculation formula is as follows:

[0114]

[0115] Among them, the dimensions of Q, K, and V are... Consistent, This indicates batch normalization, SELU indicates activation function, Reshape indicates data reshaping, and Split indicates splitting.

[0116] The obtained Q, K, V are then adjusted for dimensionality using Reshape(·) and normalized using Norm(·), and then subjected to matrix multiplication. and learnable temperature coefficient After calculating the attention weights, the softmax(·) activation function is applied to normalize them into a probability distribution. Finally, the attention weights are compared with... Multiplication yields the output:

[0117]

[0118] Output After adjusting the dimensions using Reshape(·), and then processing by the Dropout layer and linear projection layer, the enhanced feature representation is obtained through residual connections and BatchNorm(·) normalization. :

[0119]

[0120] Since the feature extraction process for the left and right hemispheres is exactly the same, the above operations can be performed separately for each hemisphere to extract their respective enhanced features. and This process involves concatenating and fusing enhanced features to obtain a final representation with stronger expressive power. ;

[0121] Furthermore, to fully utilize the temporal characteristics of EEG signals, this invention proposes a multi-scale Transformer module to enhance the model's ability to model short-term, medium-term, and long-term dependencies. Figure 6 This is a schematic diagram of the multi-scale Transformer module provided in an embodiment of the present invention, as shown below. Figure 6 As shown, the module combines multi-scale convolution and self-attention mechanism to capture both local details and global temporal patterns simultaneously.

[0122] First, in order to capture global time dependencies, this invention performs a process on the input signal. Perform a dimensional transformation using permute(·) to obtain To adapt it to the multi-head self-attention (MHA) mechanism:

[0123]

[0124] Subsequently, the time-domain weighted feature representation is calculated using the multi-head attention (MHA) mechanism:

[0125]

[0126] in, This is the weighted self-attention output; to further compress the time dimension and extract key features, adaptive pooling is used to obtain preliminary time-dependent features:

[0127]

[0128] After obtaining preliminary time-dependent features, this invention employs multi-scale 1D convolution to simultaneously capture short-term, medium-term, and long-term time-dependent features in order to further extract information at different time scales. Let the convolution kernel size be... ,in (Corresponding to short-term, medium-term, and long-term respectively), the feature extraction process is as follows:

[0129]

[0130] in, and These represent short-term, medium-term, and long-term time-dependent characteristics, respectively. This represents convolution operations at different scales. This process ensures that features at different time scales have consistent dimensionality.

[0131] To adaptively adjust the importance of features at different time scales, this invention also proposes a lightweight weight generation module, the calculation method of which is as follows:

[0132]

[0133] in, The weight values ​​are calculated by a lightweight weight generation module and are used to dynamically adjust the importance of features at different time scales. Perform feature transformation. Enhance nonlinear expressive power, is a learnable weight matrix, and b is a learnable bias term. Normalized feature distribution, This is applied to weight normalization to ensure that its value range is within a certain range. Between, for At each time scale, all weight tensors are stacked using the stack(·) operation and normalized to a global weight tensor:

[0134]

[0135] Based on weight tensor This invention performs weighted summation of features at different scales to obtain the final global fusion features. :

[0136]

[0137]

[0138] in, This indicates the removal of redundant dimensions. This indicates element-wise multiplication.

[0139] To further enhance the representation capability of time-domain features, this invention incorporates time-domain globally dependent features. With self-attention output Perform element-wise multiplication and further fuse the results through convolution:

[0140]

[0141] Furthermore, after feature fusion, the input is processed through a feedforward network (FFN). First, the input... The model is first transformed into hidden layer dimensions through the first linear transformation layer (linear1), then non-linearly transformed using the ReLU activation function, and then subjected to Dropout to enhance its generalization ability. Next, the output passes through the second linear transformation layer (linear2) to compress the dimensions, followed by another Dropout operation. Finally, it undergoes LayerNorm normalization to obtain the final output containing rich contextual information. The formula is as follows:

[0142]

[0143] The final result It incorporates both global temporal dependency features and retains multi-scale local temporal dependency information, providing a more discriminative temporal domain feature representation for subsequent emotion classification tasks.

[0144] Furthermore, during model training, the loss function calculation module performs error calculation on the output features and adjusts the model parameters based on the calculation results to optimize the feature extraction effect. The training process of the CNN-Transformer model can be iteratively optimized based on supervised learning, enabling the model to effectively learn the emotion-related features of EEG signals, providing highly robust feature inputs for subsequent emotion recognition.

[0145] S105. Emotion Recognition: Generate emotion recognition results based on the output of the CNN-Transformer model.

[0146] Emotion recognition results can be represented using two methods: discrete emotion classification and dimensional emotion modeling. Discrete emotion classification is based on basic emotion theories in psychology, mapping EEG signals into categories such as pleasure, anger, sadness, fear, surprise, disgust, and neutrality. Dimensional emotion modeling, on the other hand, uses continuous variables such as valence, arousal, and dominance to quantitatively describe emotional states.

[0147] Through the above processing, the CNN-Transformer model can efficiently extract deep spatiotemporal features of EEG signals, achieve accurate emotion recognition, and select different emotion classification methods according to application needs to adapt to various human-computer interaction and intelligent emotion computing tasks.

[0148] This invention provides a CNN-Transformer-based EEG emotion recognition method that combines self-attention, cross-attention, and multi-scale feature extraction techniques to efficiently learn general emotion-related EEG features. First, EEG signals are acquired using an EEG cap, and bandpass filtering, differential entropy (DE), and power spectral density (PSD) feature extraction are performed to retain key information. Then, the features are cross-stacked and mapped to a two-dimensional EEG image, which is then input into the CNN-Transformer model for feature learning.

[0149] General emotion-related EEG features refer to high-dimensional feature representations learned through a CNN-Transformer model. These features effectively preserve the temporal, frequency, and spatial information of EEG signals and highlight patterns closely related to emotional states. Specifically, the location encoding module enhances the spatial structure perception ability of EEG using spatial normalization and learnable location vectors, the hemispherical attention module extracts cross-frequency dependencies between the left and right hemispheres through dual-scale deep convolution, and the multi-scale Transformer module fuses EEG signals from different time scales to model short-term and long-term dynamic features.

[0150] In some embodiments, to further extract target features relevant to a specific emotion recognition task, this invention employs a multilayer perceptron (MLP) as a decoder to extract task-specific features from general EEG representations and optimizes the extraction based on a cross-entropy loss function to obtain the final emotion recognition result. This process can be represented as:

[0151]

[0152] in, This represents the general EEG feature representation extracted by the CNN-Transformer model. As a decoder, features are mapped to specific emotion categories and optimized based on minimizing cross-entropy loss. This mechanism can adaptively adjust the feature learning process, improving the accuracy and generalization ability of emotion recognition.

[0153] Furthermore, in order to verify the generalization ability of the model and ensure the stability of emotion recognition, this invention uses 10-fold cross-validation for experimental evaluation.

[0154] Specifically, let the dataset be... It contains a sample set ,in To determine the total number of samples, firstly, Divide into ten equal-sized subsets Each subset contains One sample.

[0155] In each training iteration, select one subset. The remaining nine subsets serve as the test set. As the training set, that is:

[0156]

[0157] This process is repeated ten times, with each subset serving as a test set once. Finally, the average performance metric of all test results is calculated to evaluate the model's generalization ability.

[0158] Through the above methods, the present invention can effectively improve the robustness of EEG emotion recognition, ensure the stable performance of the model under different experimental conditions, and is applicable to emotion computing scenarios involving multiple subjects and multiple tasks.

[0159] Table 1 shows the comparative experimental results of an EEG emotion recognition method based on CNN-Transformer provided in this embodiment of the invention on the DREAMER dataset.

[0160] Table 1 - Comparative experimental results on the DREAMER dataset

[0161] ;

[0162] The EEG emotion recognition method based on CNN-Transformer provided in this embodiment of the invention was trained and tested on the publicly available emotion recognition dataset DREAMER, and the classification accuracy of other existing algorithms on this dataset was compared with that of other algorithms using 10-fold cross-validation. The results show that the classification accuracy achieved by our method can be improved by about 1%-9% compared with other algorithms.

[0163] like Figure 9 and Figure 10 As shown, this embodiment of the invention performs classification tests on the arousal and valence dimensions of each subject on the DREAMER dataset. Figure 9 This demonstrates the trend of classification accuracy across different subjects. Figure 10 The corresponding F1 score is then given.

[0164] As can be seen from the data in the figure, most subjects maintained a high level of accuracy in both dimensions, and their F1 scores were also relatively stable, indicating that the method proposed in this invention can effectively capture the emotional characteristics contained in EEG signals in multi-subject situations.

[0165] Table 2 shows the comparative experimental results of an EEG emotion recognition method based on CNN-Transformer provided in this embodiment of the invention on the DEAP dataset.

[0166] Table 2 - Comparative Experiment Results on the DEAP Dataset

[0167] ;

[0168] The EEG emotion recognition method based on CNN-Transformer provided in this embodiment of the invention was trained and tested on the publicly available emotion recognition dataset DEAP, and the classification accuracy of other existing algorithms on this dataset was compared with that of other algorithms using 10-fold cross-validation. The results show that the classification accuracy achieved by this algorithm (ourmethod) can be improved by about 2%-8% compared with other algorithms.

[0169] like Figure 9 and Figure 10 As shown, this embodiment of the invention performs classification tests on the DEAP dataset for each subject's arousal and valence dimensions. Figure 9 This demonstrates the trend of classification accuracy across different subjects. Figure 10 The corresponding F1 score is then given.

[0170] As can be seen from the data in the figure, most subjects maintained an accuracy rate of over 90% in both dimensions, and their F1 scores were also relatively stable. This result further indicates that the method can fully explore the differentiated features of EEG signals in the dimensions of arousal and valence, and maintain stable recognition performance in various emotional states, providing strong data support for emotion computing and intelligent human-computer interaction.

[0171] Figure 11 This is a block diagram of an EEG emotion recognition device based on CNN-Transformer provided in an embodiment of the present invention. Please refer to... Figure 11 The CNN-Transformer's EEG emotion recognition device includes:

[0172] The data acquisition unit 201, used to acquire EEG signals, mainly consists of an EEG cap and a signal acquisition circuit. This unit has an adaptive impedance matching function to adapt to different users' scalp contact conditions, ensuring the stability and consistency of signal acquisition.

[0173] The preprocessing unit 202 performs sliding window segmentation on the EEG signal and extracts the first feature, which includes differential entropy (DE) and power spectral density (PSD), to characterize the time-frequency characteristics of the EEG signal.

[0174] Feature mapping unit 203 is used to map the first feature output by the preprocessing unit into a two-dimensional EEG map to preserve the time-domain, frequency-domain, and spatial-domain information of the signal. Specifically, this unit employs:

[0175] Two-dimensional topology mapping: Based on the spatial distribution rules of EEG electrodes, the extracted first feature is projected onto a two-dimensional mesh structure to preserve the topological information between electrodes;

[0176] Channel-level feature stacking: The first feature is stacked crosswise along the channel dimension to enhance the ability to model spatiotemporal relationships.

[0177] The model training unit 204 receives the two-dimensional EEG images generated by the feature mapping unit and performs feature extraction and model training using a CNN-Transformer to extract secondary features. During training, supervised learning is used to optimize model parameters to improve the accuracy and robustness of emotion recognition.

[0178] The second feature includes: deep spatiotemporal features, cross-channel dependency features, and multi-scale dynamic features.

[0179] The deep spatiotemporal feature fusion of EEG spatiotemporal features extracted by CNN and Transformer captures the temporal dependencies of EEG signals. Cross-channel dependency features enhance information interaction between different brain regions through the hemispherical attention module. At the same time, cross-channel frequency features are extracted to make full use of the spectral information of different brain regions and improve the ability to identify emotional states. Multi-scale dynamic features integrate EEG signal features at different time scales to improve the model’s temporal modeling ability.

[0180] The emotion recognition unit 205 performs emotion recognition on the extracted second features based on the trained CNN-Transformer model. Its output third features include, but are not limited to:

[0181] Discrete emotion features: Based on discrete emotion theory, the second feature is classified to identify specific emotion categories, such as pleasure, anger, sadness, etc.

[0182] Dimensional Emotional Characteristics: Based on dimensional emotion theory, the state of the second characteristic in emotional dimensions such as valence, arousal, and dominance is quantified.

[0183] The communication unit 206 is used to transmit emotion recognition results to external terminals or cloud platforms, enabling remote data interaction. This unit supports multiple communication protocols and is adaptable to different terminal devices to ensure the visualization and further analysis of the recognition results.

[0184] Figure 12 This is a system block diagram of an EEG emotion recognition system based on CNN-Transformer provided by the present invention. The system includes an EEG emotion recognition device, a user interface, a data storage unit, and a cloud collaboration module.

[0185] The core of this system is an EEG emotion recognition device 301 based on CNN-Transformer, which is responsible for EEG signal acquisition, preprocessing, feature extraction, model training, and emotion recognition. Detailed structure and functions have been described in the preceding embodiments and will not be repeated here.

[0186] The user interface 302 is used to intuitively display EEG signal waveforms, emotion recognition results, and signal quality feedback. This interface can use visual charts, real-time line graphs, or heatmaps to display the dynamic changes in EEG signals and use different colors or values ​​to indicate emotional states. Furthermore, the user interface supports user-defined parameter adjustments, such as signal filtering options and feature extraction parameters, to optimize individualized recognition results.

[0187] Data storage unit 303 is used to store EEG-related data, including raw EEG signals, preprocessed feature data, and historical emotion recognition records. This unit supports both local and remote storage, ensuring data security and traceability. Furthermore, this unit can be equipped with data encryption and access control mechanisms to comply with healthcare data privacy protection standards.

[0188] The cloud collaboration module 304 is used to expand the system's computing power and data sharing capabilities. Specifically, this module can upload emotion recognition results to a cloud server for large-scale data analysis and individualized model optimization. Furthermore, the cloud collaboration module supports remote model updates; the server can dynamically adjust the parameters of the CNN-Transformer model based on different user data and distribute the optimized model to local devices, improving the system's adaptability and generalization capabilities.

[0189] Through the collaborative work of the above components, this system can achieve end-to-end automated processing of EEG emotion recognition, support users to view recognition results in real time, and ensure long-term data storage and remote access.

[0190] Although the present invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the disclosed embodiments can be combined with each other in any manner. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A CNN-Transformer-based EEG emotion recognition method, characterized in that, The steps are as follows: S1. Segment the original EEG signal and extract multiple frequency band features within each segment; S2. Map the multiple frequency band features into a two-dimensional brain electrode map to simultaneously retain time domain, frequency domain, and spatial domain information; S3. The two-dimensional EEG is processed using the CNN-Transformer model to extract comprehensive spatiotemporal features; S4. Generate emotion recognition results based on the spatiotemporal features; In step S3, the CNN-Transformer model includes: a position encoding module, a hemispherical attention module, and a multi-scale Transformer module; The location encoding module inputs a two-dimensional EEG image. Global normalization is performed, and the calculation formula is as follows: ; in, , The normalized data matrix, Representative Channel Features across the entire 2D space This represents the mean value across all spatial locations within the channel. This represents the standard deviation of all spatial locations within the channel; Normalized features X ′ is composed of all time points t Normalized characteristic matrix X t The three-dimensional tensor formed by ′, with respect to the normalized features X 'Embedded learnable positional encoding pos, represented as X f Indicates in X Based on the existing structure, location coding information is embedded to enhance the perception of spatial structure; The hemispherical attention module uses dual-scale deep convolution to extract spatial and frequency domain features from the left and right hemispheres, and combines the attention mechanism to model the frequency domain dependency. First, different receptive fields are used... of Convolution kernel pair A dual-scale convolutional transformation is performed to capture local high-frequency and global low-frequency information. Then, a 1D convolution and a depthwise convolution are used to extract the query, key, and value vectors. The calculation formula is as follows: ; Among them, the dimensions of Q, K, and V are... This indicates batch normalization, SELU indicates activation function, Reshape indicates data reshaping, and Split indicates splitting. The obtained Q, K, V are then adjusted for dimensionality using Reshape(·) and normalized using Norm(·), and then subjected to matrix multiplication. and learnable temperature coefficient After calculating the attention weights, the softmax(·) activation function is applied to normalize them into a probability distribution. Finally, the attention weights are compared with... V Multiplication yields the output: ; Output After adjusting the dimensions using Reshape(·), and then processing by the Dropout layer and linear projection layer, the enhanced feature representation is obtained through residual connections and BatchNorm(·) normalization. : ; Since the feature extraction process is exactly the same for both hemispheres, the above operations are performed separately for each hemisphere to extract their respective enhanced features. and This process involves concatenating and fusing enhanced features to obtain a final representation with stronger expressive power. ; The multi-scale Transformer module achieves dynamic capture of short-term, medium-term, and long-term temporal features by fusing multi-scale convolution and self-attention mechanisms. First, in order to capture global temporal dependencies, the input signal... Perform a dimensional transformation using permute(·) to obtain To adapt it to the multi-head self-attention mechanism MHA: ; Subsequently, the time-domain weighted feature representation is computed using the multi-head attention mechanism (MHA): ; in, This is the weighted self-attention output; to further compress the time dimension and extract key features, adaptive pooling is used to obtain preliminary time-dependent features: ; After obtaining the initial time-dependency features, to further extract information at different time scales, multi-scale 1D convolution is employed to simultaneously capture short-term, medium-term, and long-term time-dependency features. The convolution kernel size is set to... ,in These correspond to short-term, medium-term, and long-term, respectively. The feature extraction process is as follows: ; in, , and These represent short-term, medium-term, and long-term time-dependent characteristics, respectively. Convolution operations representing different scales ensure that features at different time scales have consistent dimensions.

2. The EEG emotion recognition method based on CNN-Transformer according to claim 1, characterized in that, In step S1, the original EEG signal is segmented using a sliding window technique. The sliding window has a duration of 0.5 seconds and a sampling frequency of 128 Hz. At the same time, the differential entropy characteristics and power spectral density characteristics of the four frequency bands θ, α, β, and γ are calculated.

3. The EEG emotion recognition method based on CNN-Transformer according to claim 2, characterized in that, The differential entropy is used to quantify the complexity and randomness of the EEG signal, and the formula is as follows: ; in, Indicates EEG signal The differential entropy, yes The probability density function, when Follows a Gaussian distribution N(μ,σ) 2 )hour, = exp(- The above formula simplifies to: ; Where μ is the mean, σ 2 Let be the variance, and e be the natural logarithm.

4. The EEG emotion recognition method based on CNN-Transformer according to claim 2, characterized in that, The power spectral density is used to describe the energy distribution of EEG signals across different frequency bands. It is estimated using the Welch method, which analyzes the EEG signals. Divided into There are overlapping sub-segments, each sub-segment having a length of [number]. And apply a window function to each segment. To reduce spectrum leakage; Window function The energy normalization factor is defined as: ; For the Sub-segment Its periodic diagram The calculation formula is as follows: ; in Sampling frequency, As a frequency variable, subsequently, by applying all The signal is obtained by averaging the periodicity of each segment. In frequency Power spectral density estimation on: .

5. The EEG emotion recognition method based on CNN-Transformer according to claim 1, characterized in that, In step S2, mapping the multiple frequency band features into a two-dimensional brain electrode map to simultaneously retain time domain, frequency domain, and spatial domain information specifically includes: mapping differential entropy features and power spectral density features to a two-dimensional brain electrode map after cross-stacking along the channel dimension, wherein each data point on the two-dimensional brain electrode map corresponds to a feature value at a specific electrode location.

6. The EEG emotion recognition method based on CNN-Transformer according to claim 1, characterized in that, To adaptively adjust the importance of features at different time scales, a lightweight weight generation module is proposed. The calculation method of the lightweight weight generation module is as follows: ; in, The weight values ​​are calculated by a lightweight weight generation module and are used to dynamically adjust the importance of features at different time scales. Perform feature transformation. Enhance nonlinear expressive power, W is a learnable weight matrix, and b is a learnable bias term. Normalized feature distribution, This is applied to weight normalization to ensure that its value range is within a certain range. Between, for At each time scale, all weight tensors are stacked using the stack(·) operation and normalized to a global weight tensor: ; Based on weight tensor The features at different scales are weighted and summed to obtain the global fused features. : ; ; in, This represents element-wise multiplication. To further enhance the ability to represent time-domain features, the features will be... With self-attention output Perform element-wise multiplication and further fuse the results through convolution: ; After feature fusion, the input is processed further through a feedforward network (FFN). First, the input... The model is first transformed into hidden layer dimensions through the first linear transformation layer (linear1), then non-linearly transformed using the ReLU activation function, and then subjected to Dropout to enhance its generalization ability. Next, the output passes through the second linear transformation layer (linear2) to compress the dimensions, followed by another Dropout operation. Finally, it undergoes LayerNorm normalization to obtain the final output containing rich contextual information. The formula is as follows: 。 7. An apparatus for implementing the EEG emotion recognition method based on CNN-Transformer as described in any one of claims 1-6, characterized in that, The device comprises: The data acquisition unit is used to acquire raw EEG signals in real time. This unit uses an EEG cap and signal acquisition circuit, and has an adaptive impedance matching function to adapt to different users' scalp contact conditions. The preprocessing unit, connected to the data acquisition unit, performs sliding window segmentation processing on the raw EEG signal and extracts the differential entropy features and power spectral density features of the four frequency bands θ, α, β, and γ. The feature mapping unit is used to cross-stack the frequency band features output by the preprocessing unit and map them into a two-dimensional brain electrode map, where each data point corresponds to the feature value of the corresponding electrode position, thereby preserving time domain, frequency domain and spatial domain information. The model training unit is used to input the two-dimensional EEG generated by the feature mapping unit into the CNN-Transformer-based deep learning model for training. This unit integrates position encoding, hemispherical attention and multi-scale Transformer functions, and is equipped with a fully connected layer and loss function calculation module to achieve efficient extraction and representation of emotion-related features. The emotion recognition unit is used to perform emotion recognition on real-time acquired EEG signals based on the trained CNN-Transformer model, and to post-process the model output to generate multi-dimensional emotion classification results. The communication unit is used to transmit emotion recognition results to external terminals or cloud platforms to achieve real-time data interaction.

8. An EEG emotion recognition system based on CNN-Transformer, characterized in that, Includes the device as described in claim 7, as well as a user interface, a data storage unit, and a cloud collaboration module; The user interface is used to display EEG signal waveforms, emotion recognition results, and signal quality feedback in real time. The data storage unit is used to store the raw EEG signal, processed feature data, and historical identification records. The cloud-based collaboration module is used to upload emotion recognition results to a cloud platform for data analysis and supports remote model updates.

Citation Information

Patent Citations

  • Electroencephalogram emotion recognition method and system based on multiple tasks and attention mechanism

    CN118576206A

  • Classification method for electroencephalogram emotion recognition through multi-scale spatial-temporal feature extraction based on CNN and Transform

    CN118797496A

Cited By

  • An electroencephalogram emotion recognition method based on a physical-function coupling dynamic graph network

    CN122515782A