Electroencephalogram emotion recognition method and related device

Through a multi-level spatial aggregation network, the spatial characteristics of EEG signals are extracted step by step, and the problem of insufficient spatial dependence in EEG emotional recognition is solved, and higher recognition accuracy and generalization ability are achieved.

CN120257039AActive Publication Date: 2025-07-04WUYI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510140317.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-07-04
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The existing convolutional neural networks and recurrent neural networks have shortcomings in capturing the long-distance spatial dependence of EEG signals, which is difficult to fully reflect the complex dependence relationships across multiple electrodes and brain regions, resulting in inaccurate EEG emotional recognition.

Method used

A multi-level spatial aggregation network is adopted, including the whole-brain channel learning layer, the intra-brain channel learning layer and the inter-brain learning layer. The spatial characteristics of the EEG signal are extracted step by step through the Transformer encoder, and combined with the self-attention mechanism and the importance of learning the brain region, an EEG emotion recognition model is constructed.

Benefits of technology

The accuracy of EEG emotional recognition is improved, and by extracting discriminant features related to emotions in a multi-level and gradual manner, the model's ability to capture spatial dependencies within and outside the brain region is enhanced, and information loss and noise interference are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257039A_ABST
    Figure CN120257039A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an electroencephalogram emotion recognition method and a related device. The method comprises the following steps: acquiring electroencephalogram data of a subject; an electroencephalogram emotion recognition model is constructed based on a multi-level space aggregation network, the multi-level space aggregation network comprises a whole-brain channel learning layer, an intra-brain channel learning layer, an inter-brain learning layer and a classification layer, and the whole-brain channel learning layer is used for extracting whole-brain channel-level spatial features from electroencephalogram data; the brain region inner channel learning layer is used for extracting brain region level spatial features from the channel level spatial features, the brain region inner channel learning layer is used for extracting whole brain feature vectors from the brain region level spatial features, and the classification layer is used for identifying emotions; and inputting the electroencephalogram data into an electroencephalogram emotion recognition model to obtain an electroencephalogram emotion recognition result of the subject. Based on this, the embodiment of the invention can realize gradual extraction of emotion-related discriminative features from three levels of whole brain channels, intra-brain channels and brain intervals, thereby improving the accuracy of brain electrical emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of electroencephalogram signal processing, and in particular, to an electroencephalogram emotion recognition method and related device. Background Art

[0002] Affective brain-computer interfaces (aBCIs) are systems that enhance cognition, communication, and decision-making by monitoring and regulating human emotional states. aBCIs technology is particularly important for understanding and modeling emotional dynamics, and emotions are an important part of individual and collective behavior in social systems. The data of aBCIs has multiple modalities, including stereoelectroencephalography (SEEG), functional magnetic resonance imaging (fMRI), and electroencephalography (EEG). Among them, EEG is the most favored in research and practical applications due to its non-invasiveness, convenience, and low cost. Emotion recognition plays a key role in aBCIs. Researchers mainly study emotion recognition models based on EEG. Currently, many emotion recognition methods mainly use supervised machine learning techniques and rely on extracting sufficient discriminative EEG features to achieve emotion recognition. EEG data can be used to achieve emotion classification by extracting time-domain features, frequency-domain features, and brain connection features of electroencephalogram signals. In frequency-domain features, differential entropy (DE) is a widely used feature extraction method, and DE features are manual features that can efficiently identify emotions. DE features are usually independently extracted from five frequency bands (delta, theta, alpha, beta, gamma) of EEG signals, and emotion classification is performed through machine learning models (such as support vector machines, extreme learning machines, etc.). However, traditional machine learning models show limitations in processing complex electroencephalogram signals. In recent years, deep learning models have significantly improved emotion recognition performance. Convolutional Neural Networks (CNNs) and Recurrent Neural Network (RNNs) are widely used as mainstream methods to extract spatio-temporal features of EEG signals. However, CNNs and RNNs are insufficient in capturing long-distance spatial dependencies of EEG signals and are difficult to fully reflect the complex dependencies across multiple electrodes and brain regions, resulting in inaccurate electroencephalogram emotion recognition of subjects. Therefore, how to improve the accuracy of electroencephalogram emotion recognition of subjects has become an urgent technical problem to be solved. Summary of the Invention

[0003] Embodiments of the present invention provide an electroencephalogram (EEG) emotion recognition method and related devices, which can progressively extract discriminative features related to emotions from three different levels: whole-brain channels, intra-brain region channels, and inter-brain regions, thereby improving the accuracy of EEG emotion recognition for subjects.

[0004] In a first aspect, embodiments of the present invention provide an EEG emotion recognition method, including:

[0005] Obtain EEG data of a subject;

[0006] Construct an EEG emotion recognition model based on a multi-level spatial aggregation network, where the multi-level spatial aggregation network includes a whole-brain channel learning layer, an intra-brain region channel learning layer, an inter-brain region learning layer, and a classification layer. The whole-brain channel learning layer is used to extract channel-level spatial features of the whole brain from the EEG data, the intra-brain region channel learning layer is used to extract intra-brain region-level spatial features from the channel-level spatial features, the inter-brain region learning layer is used to extract a whole-brain feature vector from the intra-brain region-level spatial features, and the classification layer is used to recognize emotions;

[0007] Input the EEG data into the EEG emotion recognition model to obtain the EEG emotion recognition result of the subject.

[0008] In some embodiments, the step of inputting the EEG data into the EEG emotion recognition model to obtain the EEG emotion recognition result of the subject includes:

[0009] Input the EEG data into the whole-brain channel learning layer to obtain channel-level spatial features;

[0010] Input the channel-level spatial features into the intra-brain region channel learning layer to obtain intra-brain region-level spatial features;

[0011] Input the intra-brain region-level spatial features into the inter-brain region learning layer to obtain a whole-brain feature vector;

[0012] Input the whole-brain feature vector into the classification layer to obtain the EEG emotion recognition result.

[0013] In some embodiments, the whole-brain channel learning layer includes multiple Transformer encoders. The step of inputting the EEG data into the whole-brain channel learning layer to obtain channel-level spatial features includes:

[0014] Convert the EEG data into multiple embedding sequences through a linear embedding layer and positional encoding;

[0015] Input the multiple embedding sequences into a multi-layer Transformer encoder to obtain the channel-level spatial features, where the multi-layer Transformer encoder includes a multi-head self-attention layer and a multi-layer perceptron.

[0016] In some embodiments, the step of inputting the channel-level spatial features into the intra-brain region channel learning layer to obtain the brain-region-level spatial features includes:

[0017] Divide the channel-level spatial features into M feature subsets of different sizes, where M is the number of brain regions;

[0018] Input each of the feature subsets into a corresponding Transformer encoder to learn the spatial information from different brain regions in parallel, and each Transformer encoder corresponds to a different brain region;

[0019] Perform average pooling on the output of each Transformer encoder along the channel dimension to obtain multiple regional features;

[0020] Concatenate the multiple regional features to obtain the brain-region-level spatial features.

[0021] In some embodiments, the step of inputting the brain-region-level spatial features into the inter-brain region learning layer to obtain the whole-brain feature vector includes:

[0022] Perform average pooling on the last dimension of the brain-region-level spatial features to obtain brain-region feature statistics;

[0023] Input the brain-region feature statistics into two layers of linear layers to learn the importance of different brain regions, and obtain the corresponding brain-region importance weights;

[0024] Perform weighted sum and average pooling processing according to the brain-region-level spatial features and the brain-region importance weights to obtain the whole-brain feature vector.

[0025] In some embodiments, in the whole-brain channel learning layer, the input electroencephalogram (EEG) data is defined as where e represents the number of channels, f represents the number of frequency bands, and t represents the number of sliding windows; the EEG data x is obtained by separately extracting differential entropy features for the original EEG data of each channel using t sliding windows on each frequency band; by changing the dimension of the samples, each channel is regarded as an independent embedding block, and where Through a linear embedding layer and positional encoding the EEG data is converted into an embedding sequence The formula is as follows:

[0026]

[0027] Introduce a Transformer encoder with L layers in the whole-brain channel learning layer, denoted as E main , and the output of the L-th layer of the Transformer encoder is where, is the channel-level spatial feature extracted by the whole-brain channel learning layer, and the calculation process is expressed by the following formula:

[0028]

[0029] In formulas (2) and (3), and represent the outputs of the multi-head self-attention layer and the multi-layer perceptron respectively.

[0030] In some embodiments, in the channel learning layer in the brain region, the brain region is divided into M partition regions based on e electrodes, and each partition region contains an unequal number of electrodes; the channel-level spatial feature is divided into feature subsets of different sizes where M represents the number of brain regions; by grouping channels into partition regions, the features within each partition region are input into the corresponding Transformer encoder to learn the spatial information from different brain regions in parallel; each Transformer encoder corresponds to a specific partition region to learn the channel-level spatial information within the brain region; subsequently, average pooling is performed on the outputs of each Transformer encoder along the channel dimension to obtain the corresponding regional feature Finally, the outputs of each Transformer encoder are concatenated to obtain the brain-region-level spatial feature z is defined as:

[0031]

[0032] z = [z 1 ; z 2 ;...; z M , (5)

[0033] where, represents the average pooling function.

[0034] In some embodiments, in the inter-brain region learning layer, after obtaining the brain region-level spatial feature z, first, average pooling is performed on the last dimension of the brain region-level spatial feature z to obtain a vector The calculation process is expressed as:

[0035]

[0036] Next, two simple linear layers are used to learn the importance of different brain regions:

[0037]

[0038] where δ represents the ReLU activation function, and r is the reduction ratio for dimensionality reduction, σ is the sigmoid activation function; represents the learned importance of brain regions; the output of the inter-brain region learning layer is expressed as:

[0039]

[0040] where, represents expanding s into a matrix with the same shape as z through the broadcasting mechanism, represents element-wise multiplication, and the broadcasting mechanism is used to weight the features of different brain regions; subsequently, an average pooling operation is applied to the weighted features

[0041] Second, the embodiments of the present invention also provide an electroencephalogram emotion recognition device, and the device includes:

[0042] An acquisition module, configured to acquire electroencephalogram data of a subject;

[0043] A construction module, configured to construct an electroencephalogram emotion recognition model based on a multi-level spatial aggregation network, where the multi-level spatial aggregation network includes a whole-brain channel learning layer, an intra-brain region channel learning layer, an inter-brain region learning layer, and a classification layer. The whole-brain channel learning layer extracts channel-level spatial features of the whole brain from the electroencephalogram data, the intra-brain region channel learning layer extracts brain region-level spatial features from the channel-level spatial features, the inter-brain region learning layer extracts a whole-brain feature vector from the brain region-level spatial features, and the classification layer is used to recognize emotions;

[0044] A recognition module, configured to input the electroencephalogram data into the electroencephalogram emotion recognition model to obtain an electroencephalogram emotion recognition result of the subject.

[0045] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions for executing the electroencephalogram (EEG) emotion recognition method as described in the first aspect.

[0046] According to the EEG emotion recognition method and related device provided by the embodiments of the present invention, the EEG emotion recognition method includes: obtaining EEG data of a subject; constructing an EEG emotion recognition model based on a multi-level spatial aggregation network, where the multi-level spatial aggregation network includes a whole-brain channel learning layer, an intra-brain region channel learning layer, an inter-brain region learning layer, and a classification layer. The whole-brain channel learning layer extracts channel-level spatial features of the whole brain from the EEG data, the intra-brain region channel learning layer is used to extract intra-brain region-level spatial features from the channel-level spatial features, the inter-brain region learning layer extracts a whole-brain feature vector from the intra-brain region-level spatial features, and the classification layer is used to recognize emotions; inputting the EEG data into the EEG emotion recognition model to obtain the EEG emotion recognition result of the subject. Based on this, the embodiments of the present invention realize progressive extraction of discriminative features related to emotions from three different levels: the whole-brain channels, the intra-brain region channels, and the inter-brain regions, thereby improving the accuracy of EEG emotion recognition for the subject. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a flowchart of an EEG emotion recognition method provided by an embodiment of the present invention;

[0048] Figure 2 is an overall block diagram of a multi-level spatial aggregation network provided by an embodiment of the present invention;

[0049] Figure 3 is a schematic diagram of an EEG emotion recognition device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0051] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the flowchart in the specification and claims and the following drawings. Terms such as "first" and "second" in the specification, claims, and the following drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0052] In the embodiments of the present invention, words such as "furthermore", "exemplarily", or "optionally" are used to represent examples, illustrations, or explanations, and should not be construed as being more preferred or having more advantages than other embodiments or design solutions. The use of words such as "furthermore", "exemplarily", or "optionally" is intended to present relevant concepts in a specific manner.

[0053] For the convenience of describing the working principle of the embodiments of the present invention later, the following first gives an introduction to the relevant technical scenarios.

[0054] Affective brain-computer interfaces (aBCIs) are systems that enhance cognition, communication, and decision-making by monitoring and regulating human emotional states. aBCIs technology is particularly important for understanding and modeling emotional dynamics, which are an important part of individual and collective behavior in social systems. The data of aBCIs has multiple modalities, including stereoelectroencephalography (SEEG), functional magnetic resonance imaging (fMRI), and electroencephalography (EEG). Among them, EEG is the most favored in research and practical applications due to its non-invasiveness, convenience, and low cost. Emotion recognition plays a key role in aBCIs. Researchers mainly study emotion recognition models based on EEG. Currently, many emotion recognition methods mainly use supervised machine learning techniques and rely on extracting sufficient discriminative EEG features to achieve emotion recognition. EEG data can be used for emotion classification by extracting time-domain features, frequency-domain features, and brain connectivity features of electroencephalogram signals. In frequency-domain features, differential entropy (DE) is a widely used feature extraction method, and DE features are handcrafted features that can efficiently identify emotions. DE features are usually independently extracted from five frequency bands (delta, theta, alpha, beta, gamma) of EEG signals and used for emotion classification through machine learning models (such as support vector machines, extreme learning machines, etc.). However, traditional machine learning models show limitations in processing complex electroencephalogram signals. In recent years, deep learning models have significantly improved emotion recognition performance. Convolutional Neural Networks (CNNs) and Recurrent Neural Network (RNN) as mainstream methods are widely used to extract spatio-temporal features of EEG signals. However, CNNs and RNNs are insufficient in capturing long-distance spatial dependencies of EEG signals and are difficult to comprehensively reflect the complex dependencies across multiple electrodes and brain regions, resulting in inaccurate electroencephalogram emotion recognition of subjects. Therefore, how to improve the accuracy of electroencephalogram emotion recognition of subjects has become a technical problem to be solved urgently.

[0055] Based on this, the present invention provides an electroencephalogram (EEG) emotion recognition method and related device. The EEG emotion recognition method includes: obtaining the EEG data of a subject; constructing an EEG emotion recognition model based on a multi-level spatial aggregation network, where the multi-level spatial aggregation network includes a whole-brain channel learning layer, an intra-brain-region channel learning layer, an inter-brain-region learning layer, and a classification layer. The whole-brain channel learning layer is used to extract the whole-brain channel-level spatial features from the EEG data, the intra-brain-region channel learning layer is used to extract the intra-brain-region channel-level spatial features from the channel-level spatial features, the inter-brain-region learning layer is used to extract the whole-brain feature vector from the intra-brain-region level spatial features, and the classification layer is used to recognize emotions; inputting the EEG data into the EEG emotion recognition model to obtain the EEG emotion recognition result of the subject. Based on this, the embodiments of the present invention realize progressive extraction of discriminative features related to emotions from three different levels: the whole-brain channels, the intra-brain-region channels, and the inter-brain-region, so as to improve the accuracy of EEG emotion recognition for the subject.

[0056] It should be noted that in each specific embodiment of the present invention, when it comes to relevant processing that needs to be performed according to data related to the subject's identity or characteristics, such as subject information, subject behavior data, subject historical data, and subject location information, the permission or consent of the subject will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present invention need to obtain the personal information of the subject, the separate permission or separate consent of the subject will be obtained by means of form collection or filling out an informed consent form. After clearly obtaining the separate permission or separate consent of the subject, the necessary subject-related data for the normal operation of the embodiments of the present invention will be obtained.

[0057] The following further elaborates on the embodiments of the present invention with reference to the accompanying drawings.

[0058] As Figure 1 shown, Figure 1 is a flowchart of an EEG emotion recognition method provided by an embodiment of the present invention. The EEG emotion recognition method may include, but is not limited to, steps S101 to S103.

[0059] Step S101: Obtain the EEG data of the subject;

[0060] Step S102: Construct an EEG emotion recognition model based on a multi-level spatial aggregation network, where the multi-level spatial aggregation network includes a whole-brain channel learning layer, an intra-brain-region channel learning layer, an inter-brain-region learning layer, and a classification layer. The whole-brain channel learning layer is used to extract the whole-brain channel-level spatial features from the EEG data, the intra-brain-region channel learning layer is used to extract the intra-brain-region channel-level spatial features from the channel-level spatial features, the inter-brain-region learning layer is used to extract the whole-brain feature vector from the intra-brain-region level spatial features, and the classification layer is used to recognize emotions;

[0061] Step S103: Input the EEG data into the EEG emotion recognition model to obtain the EEG emotion recognition result of the subject.

[0062] It can be understood that the present invention uses two publicly available emotion EEG datasets, SEED and SEED-IV, from Shanghai Jiao Tong University as experimental data. The SEED dataset contains EEG signal data of 15 subjects; the SEED-IV dataset contains EEG signals of 15 subjects. Among them, each dataset contains data of 3 Sessions. In each Session, the SEED dataset has 15 movie clips of stimulating emotions, and the SEED-IV dataset has 24 movie clips of stimulating emotions.

[0063] It can be understood that as Figure 2 shown, the multi-level spatial aggregation network HSA-Former of the present invention mainly consists of a Whole-Brain Channel Learning Layer, an Intra-Brain-Region Channel Learning Layer, an Inter-Brain-Region Learning Layer, and a Classification Layer. Among them, the Whole-Brain Channel Learning Layer is used to learn the channel-level spatial information of the whole brain, the Intra-Brain-Region Channel Learning Layer is used to learn the channel-level spatial information within each brain region in parallel, the Inter-Brain-Region Learning Layer is used to learn the brain-region-level spatial information, and the Classification Layer is used to identify emotions. In the model architecture of the present invention, HSA-Former can extract the spatial relationship of EEG signals from multiple levels and multiple perspectives, providing more discriminative features for emotion recognition.

[0064] It can be understood that in the Whole-Brain Channel Learning Layer, the input EEG data is defined as where w represents the number of channels, f represents the number of frequency bands, and t represents the number of sliding windows; the EEG data x is obtained by separately extracting differential entropy features for the original EEG data of each channel using t sliding windows on each frequency band. To learn the channel-level spatial information, each channel is regarded as an independent embedding patch by changing the dimension of the sample, and where through the linear embedding layer and the positional encoding the EEG data is converted into an embedding sequence The specific formula is as follows:

[0065]

[0066] To extract the channel-level spatial features of the whole brain, the present invention introduces a Transformer encoder with L layers in the whole-brain channel learning layer, denoted as E main . The output of the L-th layer of this encoder is where is the channel-level spatial feature extracted by the whole-brain channel learning layer. The specific calculation process is represented by the following equations:

[0067]

[0068] In formulas (2) and (3), and represent the outputs of the multi-head self-attention layer (MSA) and the multi-layer perceptron (MLP), respectively.

[0069] It can be understood that since there are significant differences in the emotional activation levels of each channel, directly performing average pooling on the channel-level spatial feature may lose important information. From a physiological perspective, the human brain consists of multiple functional regions, and the responses to different emotions can be observed through the activation of different brain regions. The activation differences between brain region functions can provide effective features for distinguishing emotions. Therefore, it is crucial to consider both the local connectivity between channels and the activation differences of brain functional regions for emotion recognition. Based on this observation, the present invention proposes a method of aggregating pooling according to the hierarchical structure of brain regions.

[0070] (1) The present invention divides the brain into M regions based on e electrodes, and each brain region contains a different number of electrodes. According to the brain region partition of the electrodes in Figure 2 . Therefore, the channel-level spatial feature can be further divided into feature subsets of different sizes where M represents the number of brain regions. On this basis, the present invention further explores the influence of brain regions on spatial features by grouping channels into brain regions.

[0071] (2) The present invention explores the EEG features within specific brain regions to retain these key information. Specifically, the present invention inputs the features within each partition region into the corresponding Transformer encoder to parallelly learn the channel-level spatial information from different brain regions. Each Transformer encoder Corresponding to a specific partition region, it can learn the internal channel spatial information of the brain region. Subsequently, the present invention performs average pooling on the outputs of each Transformer encoder along the channel dimension respectively to obtain the brain region features corresponding to the output Finally, the outputs of each Transformer encoder are concatenated to obtain the brain region level spatial features z is defined as:

[0072]

[0073] z = [z 1 ; z 2 ;...; z M , (5)

[0074] wherein, represents the average pooling function. By this method, the present invention can suppress noise interference while reducing the loss of important dynamic information, and improve the generalization ability of the model.

[0075] It can be understood that after obtaining the brain region level spatial feature z, in order to learn the importance of M brain region groups, the present invention utilizes a self-attention module in the inter-brain region learning layer. The self-attention module generates attention weights to further explore the importance of brain regions under different emotional states. Specifically, the present invention first performs average pooling on the last dimension of the brain region representation to obtain a vector This calculation process can be expressed as:

[0076]

[0077] Then, the present invention uses two simple linear layers to learn the importance of different brain regions.

[0078]

[0079] wherein, δ represents the ReLU activation function, and r is the reduction ratio for dimensionality reduction, and σ is the sigmoid activation function. represents the learned importance of brain regions.

[0080] The output of the inter-brain region learning layer module can be expressed as:

[0081]

[0082] wherein, represents expanding s into a matrix with the same shape as z through the broadcasting mechanism, denotes the element - wise product, which is used to weight the features of different brain regions. Subsequently, an average pooling operation is applied to the weighted features

[0083] Finally, the present invention uses a simple linear layer with a softmax function to predict the emotion labels.

[0084]

[0085] Among them, and are the parameters of the fully - connected layer, and C is the number of emotion label categories.

[0086] It can be understood that in order to fully explore the spatial features between electrode channels, within brain regions, and between brain regions, the present invention proposes a multi - level spatial aggregation neural network model HSA - Former based on Transformer. The present invention has the following characteristics: 1) The model is a multi - level spatial information learning architecture that progressively extracts different spatial features from electrodes, intra - brain regions to inter - brain regions. This architecture enhances the model's ability to capture local and global spatial relationships, thus fully exploring useful spatial information. 2) The model can parallel - learn the internal features of different brain regions from different channels. This structure allows simultaneous learning of different spatial information and enhances the ability to capture complex spatial dependency relationships within brain regions. 3) The model is an effective aggregation method that alleviates the problem of loss of effective information caused by direct average pooling in the ViT structure. This method sequentially aggregates spatial features at the levels from channel - level to brain - region - level through the self - attention mechanism.

[0087] The following further illustrates the electroencephalogram emotion recognition method of the present invention in combination with specific embodiments.

[0088] Step1: Definition of the emotion recognition problem. In terms of the form of data, after extracting differential entropy features from the original electroencephalogram data, the present invention obtains the feature where e represents the number of channels, f represents the number of frequency bands, t represents the number of sliding windows; and the one - hot label corresponding to x is C represents the number of categories.

[0089] Step2: Construct a multi - level aggregation neural network (HSA - Former) based on Transformer, as shown in the process Figure 1As shown in the figure. HSA-Former includes a Whole-Brain Channel Learning Layer to learn the channel-level spatial information of the whole brain, an Intra-Brain-Region Channel Learning Layer to learn the channel-level spatial information within each brain region in parallel, an Inter-Brain-Region Learning Layer to learn the brain-region level spatial information, and a Classification Layer to perform related emotion prediction.

[0090] Step3: In the Whole-Brain Channel Learning Layer, by changing the dimension of the sample x, each channel is regarded as an independent embedding patch, and we get where Through the linear embedding layer and positional encoding the electroencephalogram data is converted into an embedding sequence The specific formula is as follows:

[0091]

[0092] To extract the channel-level channel features, the present invention introduces a Transformer encoder with L layers in the Whole-Brain Channel Learning Layer, denoted as E main . The output of the L-th layer of this encoder is where, is the channel-level spatial feature extracted by the Whole-Brain Channel Learning Layer. The specific calculation process is represented by the following equations:

[0093]

[0094] In formulas (2) and (3), and represent the outputs of the Multi-Head Self-Attention (MSA) and the Multilayer Perceptron (MLP) respectively.

[0095] Step4: In the Intra-Brain-Region Channel Learning Layer, the present invention divides the brain into M regions based on e electrodes, and each brain region contains an unequal number of electrodes. According to Figure 2 the brain region partition of the electrodes. Therefore, the channel-level spatial feature can be further divided into feature subsets of different sizes where M represents the number of brain regions. Next, the present invention inputs the features within each partition region into the corresponding Transformer encoder to learn the channel-level spatial information from within different brain regions in parallel. The present invention performs average pooling on the outputs of each Transformer encoder along the channel dimension to obtain the output Finally, the outputs of each Transformer encoder are concatenated to obtain the brain-region-level spatial features z is defined as:

[0096]

[0097] z = [z 1 ; z 2 ;...; z M , (5)

[0098] where, represents the average pooling function. Through this method, the present invention can suppress noise interference while reducing the loss of important dynamic information and improve the generalization ability of the model.

[0099] Step5: In the inter-brain-region learning layer, the present invention performs average pooling on the last dimension of the brain-region-level feature z to obtain the brain-region feature statistic This calculation process can be expressed as:

[0100]

[0101] Next, the present invention uses two simple linear layers to learn the importance of different brain regions.

[0102]

[0103] where, δ represents the ReLU activation function, and are two fully connected layers without bias, r is the reduction ratio for dimensionality reduction, and σ is the sigmoid activation function. represents the learned importance of brain regions.

[0104] The output of the inter-brain-region learning layer module can be expressed as:

[0105]

[0106] where, represents expanding s into a matrix with the same shape as z through the broadcasting mechanism, denotes the element-wise product, which is used to weight the features of different brain regions. Subsequently, an average pooling operation is applied to the weighted features

[0107] Step6: In the classification layer, the present invention uses a simple linear layer with a Softmax function to predict the emotion labels.

[0108]

[0109] wherein, and are the parameters of the fully connected layer, and C is the number of categories of emotion labels.

[0110] It should be noted that the present invention verifies the effectiveness of different components in HSA-Former. The present invention conducts ablation experiments on the SEED and SEED-IV datasets. The Whole-Brain Channel Learning Layer (WBCLL), Intra-Brain-Region Channel Learning Layer (IBRCLL), and Inter-Brain-Region Learning Layer (IBRLL) are specifically as follows:

[0111] WBCLL architecture: This model only has the whole-brain channel learning layer.

[0112] WBCLL + IBRCLL architecture: This model has the whole-brain channel learning layer and the intra-brain-region channel learning layer.

[0113] WBCLL + IBRLL architecture: This model has the whole-brain channel learning layer and the inter-brain-region learning layer.

[0114] WBCLL + IBRCLL + IBRLL architecture: This model is the HSA-Former of the present invention.

[0115] As shown in Table 1, each component of HSA-Former is important and indispensable. The performance of WBCLL + IBRCLL is better than that of a single WBCLL, which indicates that IBRCLL can effectively learn the channel-level spatial information within the brain region. At the same time, the performance of WBCLL + IBRLL is better than that of a single WBCLL, which indicates that IBRLL can effectively learn the spatial information between brain regions. When HSA-Former can integrate the IBRCLL module and the IBRLL module, the model performance reaches the optimal.

[0116] Table 1 Performance comparison of different models for evaluating the effectiveness of each component in HSA-Former on the SEED and SEED-IV datasets

[0117]

[0118] The present invention compares the performance of the proposed method with various methods such as SVM, Bi-HDM, and SOGNN, as shown in Table 2.

[0119] Table 2 Emotion recognition performance of the leave-one-subject method on the SEED and SEED-IV datasets

[0120]

[0121] It can be seen from the data comparison in Table 2 that HSA-Former is significantly better than other methods.

[0122] Based on this, an electroencephalogram (EEG) emotion recognition method proposed by the present invention uses a multi-level spatial aggregation network to progressively extract discriminative features related to emotions from three different levels: whole-brain channels, intra-brain region channels, and inter-brain regions. As a useful aggregation framework, compared with the feature loss caused by average pooling, the multi-level dimensionality reduction of the multi-level spatial aggregation network can reduce the dimension from the physiological mechanism of emotions and at the same time reduce the loss of feature information caused by dimensionality reduction.

[0123] In addition, as Figure 3 shown, an embodiment of the present invention also discloses an EEG emotion recognition device, which includes:

[0124] An acquisition module 110, configured to acquire EEG data of a subject;

[0125] A construction module 120, configured to construct an EEG emotion recognition model based on the multi-level spatial aggregation network, where the multi-level spatial aggregation network includes a whole-brain channel learning layer, an intra-brain region channel learning layer, an inter-brain region learning layer, and a classification layer. The whole-brain channel learning layer is configured to extract channel-level spatial features of the whole brain from the EEG data. The intra-brain region channel learning layer is configured to extract intra-brain region-level spatial features from the channel-level spatial features. The inter-brain region learning layer is configured to extract a whole-brain feature vector from the intra-brain region-level spatial features. The classification layer is configured to recognize emotions;

[0126] A recognition module 130, configured to input the EEG data into the EEG emotion recognition model to obtain an EEG emotion recognition result of the subject.

[0127] The EEG emotion recognition device of the embodiment of the present invention is used to execute the EEG emotion recognition method in the above embodiment, and its specific processing process is the same as that of the EEG emotion recognition method in the above embodiment, and will not be described in detail here.

[0128] In addition, an embodiment of the present invention also discloses a computer-readable storage medium, which stores computer-executable instructions for executing the EEG emotion recognition method in any of the previous embodiments.

[0129] The system architecture and application scenarios described in the embodiments of the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. As can be known to those skilled in the art, with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0130] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0131] In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0132] As used in this specification, the terms "component", "module", "system", etc. are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, an execution thread, a program, or a computer. By way of illustration, both an application running on a computing device and the computing device can be components. One or more components can reside within a process or execution thread, and a component can be located on one computer or distributed between two or more computers. In addition, these components can execute from various computer-readable media having various data structures stored thereon. A component can communicate, for example, by way of signals with one or more data packets (e.g., data from two components interacting with another component from a local system, a distributed system, or a network, such as the Internet interacting with other systems via signals) through local or remote processes.

Claims

1. A method for electroencephalogram emotion recognition, characterized in that, Including: Obtaining the electroencephalogram (EEG) data of the subject; Constructing an EEG emotion recognition model based on a multi-level spatial aggregation network, where the multi-level spatial aggregation network includes a whole-brain channel learning layer, an intra-brain region channel learning layer, an inter-brain region learning layer, and a classification layer. The whole-brain channel learning layer is used to extract the channel-level spatial features of the whole brain from the EEG data. The intra-brain region channel learning layer is used to extract the intra-brain region-level spatial features from the channel-level spatial features. The inter-brain region learning layer is used to extract the whole-brain feature vector from the intra-brain region-level spatial features. The classification layer is used to recognize emotions; Inputting the EEG data into the EEG emotion recognition model to obtain the EEG emotion recognition result of the subject.

2. The EEG emotion recognition method according to claim 1, characterized in that The step of inputting the EEG data into the EEG emotion recognition model to obtain the EEG emotion recognition result of the subject includes: Inputting the EEG data into the whole-brain channel learning layer to obtain the channel-level spatial features; Inputting the channel-level spatial features into the intra-brain region channel learning layer to obtain the intra-brain region-level spatial features; Inputting the intra-brain region-level spatial features into the inter-brain region learning layer to obtain the whole-brain feature vector; Inputting the whole-brain feature vector into the classification layer to obtain the EEG emotion recognition result.

3. The EEG emotion recognition method according to claim 2, characterized in that, The whole-brain channel learning layer includes multiple layers of Transformer encoders. The step of inputting the EEG data into the whole-brain channel learning layer to obtain the channel-level spatial features includes: Converting the EEG data into multiple embedding sequences through a linear embedding layer and positional encoding; Inputting the multiple embedding sequences into the multiple layers of Transformer encoders to obtain the channel-level spatial features, where the multiple layers of Transformer encoders include a multi-head self-attention layer and a multi-layer perceptron.

4. The electroencephalogram emotion recognition method according to claim 2, characterized in that The step of inputting the channel-level spatial features into the intra-brain region channel learning layer to obtain the intra-brain region-level spatial features includes: Dividing the channel-level spatial features into M feature subsets of different sizes, where M is the number of brain regions; Inputting each feature subset into the corresponding Transformer encoder to parallelly learn the spatial information from different intra-brain regions, and each Transformer encoder corresponds to a different brain region; Performing average pooling on the output of each Transformer encoder along the channel dimension to obtain multiple regional features; Concatenating the multiple regional features to obtain the intra-brain region-level spatial features.

5. The electroencephalogram emotion recognition method according to claim 2, wherein The step of inputting the intra-brain region-level spatial features into the inter-brain region learning layer to obtain the whole-brain feature vector includes: Performing average pooling on the last dimension of the intra-brain region-level spatial features to obtain the intra-brain region feature statistic; Inputting the intra-brain region feature statistic into two layers of linear layers to learn the importance of different brain regions to obtain the corresponding intra-brain region importance weights; Performing weighted sum and average pooling processing according to the intra-brain region-level spatial features and the intra-brain region importance weights to obtain the whole-brain feature vector.

6. The electroencephalogram emotion recognition method according to claim 1, wherein In the whole-brain channel learning layer, the input EEG data is defined as where e represents the number of channels, f represents the number of frequency bands, and t represents the number of sliding windows; the EEG data x is obtained by separately extracting differential entropy features for the original EEG data of each channel using t sliding windows on each frequency band; by changing the dimension of the sample, each channel is regarded as an independent embedding block, obtaining where Through the linear embedding layer and positional encoding the EEG data is converted into an embedding sequence The formula is as follows: Introduce a Transformer encoder with L layers into the whole-brain channel learning layer, denoted as E main , the output of the L-th layer of the Transformer encoder is where is the channel-level spatial feature extracted by the whole-brain channel learning layer, and the calculation process is expressed by the following formula: In formulas (2) and (3), and represent the outputs of the multi-head self-attention layer and the multi-layer perceptron, respectively.

7. The EEG emotion recognition method according to claim 6, characterized in that, In the channel learning layer within the brain region, the brain region is divided into M partition regions based on e electrodes, and each partition region contains a different number of electrodes; the channel-level spatial features are divided into feature subsets of different sizes where M represents the number of brain regions; by grouping channels into partition regions, the within-region features of each partition are input into the corresponding Transformer encoder to learn the spatial information from different brain regions in parallel; each Transformer encoder corresponds to a specific partition region to learn the internal information of each brain region; subsequently, average pooling is performed on the output of each Transformer encoder along the channel dimension to obtain the region features corresponding to each brain region Finally, the outputs of each Transformer encoder are concatenated to obtain the brain-region-level spatial features z is defined as: z = [z 1 ; z 2 ;...; z M , (5) Among them, represents the average pooling function.

8. The EEG emotion recognition method according to claim 7, characterized in that, In the inter-brain learning layer, first, the obtained brain-region-level spatial feature z is processed, and average pooling is performed on the last dimension of the brain-region-level spatial feature z to obtain the brain-region feature statistic The calculation process is expressed as: Next, use two simple linear layers to learn the importance of different brain regions: where δ represents the ReLU activation function, and r is the reduction ratio for dimensionality reduction, and σ is the sigmoid activation function; represents the learned importance of brain regions; The output of the inter-brain learning layer is expressed as: Among them, means expanding s into a matrix with the same shape as z through the production of a broadcaster, means element-wise multiplication, and the broadcasting mechanism is used to weight the features of different brain regions; subsequently, an average pooling operation is applied to the weighted features 9. An electroencephalogram emotion recognition device, characterized in that, The device includes: An acquisition module for acquiring the EEG data of the subject; A building module for constructing an electroencephalogram (EEG) emotion recognition model based on a multi-level spatial aggregation network, wherein the multi-level spatial aggregation network includes a whole-brain channel learning layer, an intra-brain region channel learning layer, an inter-brain region learning layer, and a classification layer. The whole-brain channel learning layer is used to extract channel-level spatial features of the whole brain from the EEG data. The intra-brain region channel learning layer is used to extract intra-brain region-level spatial features from the channel-level spatial features. The inter-brain region learning layer is used to extract a whole-brain feature vector from the intra-brain region-level spatial features. The classification layer is used to recognize emotions; An identification module for inputting the EEG data into the EEG emotion recognition model to obtain the EEG emotion recognition result of the subject.

10. A computer-readable storage medium storing computer-executable instructions for executing the EEG emotion recognition method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • A P300 detection method based on CNN-LSTM network

    CN109389059A

  • Driving fatigue identification method based on CNN-LSTM deep learning model

    CN109820525A

  • Emotion recognition method of multi-channel electroencephalogram (EEG) data and electronic device

    CN111134666A

  • Brain electrical emotion recognition method based on attention mechanism space-time position coding

    CN115736923A

  • Electroencephalogram signal emotion recognition method based on CNT

    CN117860250A